Run a vLLM model¶
Before you start¶
Choose a model architecture and artifact supported by the selected vLLM runtime and compute target. A Hugging Face search result alone is not proof of compatibility. You need Operator/Administrator access, memory headroom and download space.
Create¶
- Open Models → Create, select a local model and vLLM.
- Choose an eligible CPU, NVIDIA, AMD or Intel target as offered by the dashboard.
- Search Hugging Face or enter
hf://<publisher>/<repository>and select the artifact. - Review the context length, maximum sequences, KV precision and resource budget. Begin with one sequence and a moderate context for a new model/hardware combination.
- Review engine-specific deployment options only when needed. An AMD vision attention backend must exist in the selected runtime image; the selector does not install missing libraries.
- Create the model and inspect its status and Logs.
Verify and tune¶
After Ready, send a request through API Access. Increase context or concurrency gradually and test your actual workload. The memory estimator is planning assistance, not protection from every GPU allocator or attention-kernel failure.
Explicit CPU offloading, where offered, uses RAM on the same node. It can slow responses and needs enough host memory in addition to the GPU budget.
Ordinary vLLM uses KubeAI. Duplex audio is a separate vLLM-Omni profile,
not a switch that makes every chat model implement /v1/realtime.
See model lifecycle, runtime contracts and offloading diagnostics.