Skip to content

Engine and hardware compatibility

This page describes the repository capability catalog, not successful tests on every device. Availability also requires live driver/resource readiness, free slots and usable memory telemetry.

Source: compute-target catalog.

Ordinary engine targets

Compute target Architecture Configured engines
CPU amd64, arm64 vLLM, Ollama
NVIDIA GPU amd64, arm64 vLLM, Ollama, FreeToken
AMD GPU (ROCm) amd64 vLLM, Ollama
Intel GPU (XPU) amd64 vLLM

FreeToken admission policy

  • Upstream version: 0.1.3.
  • OS / host architecture: linux / amd64.
  • Vendor: nvidia; NVIDIA driver major at least 580; CUDA major 13.
  • Admitted compute capabilities: 8.6, 8.9, 12.0.
  • Allocation contract: whole-gpus-single-node; whole homogeneous GPUs, not MIG or time-slicing slots.
  • Memory strategies: auto, fused, offload, cpu, hybrid; start with auto.
  • Precision choices: auto, float16, bfloat16, float32.
  • Model IDs are checked against the catalog policy; a family name is not blanket approval.

See FreeToken configuration and the user guide.

Experimental Realtime profiles

These are runtime selections, not hardware or model certifications. An empty image needs an explicitly compatible runtime.

Profile Compute target Requested device counts Bundled image
qwen3-omni nvidia-gpu 1, 2 Pinned
qwen3-omni-rocm amd-gpu 1, 2 Not supplied
qwen3-omni-xpu intel-gpu 1, 2 Not supplied
qwen3-omni-cpu cpu 1 Not supplied

A CPU profile has no physical GPU despite its stage/device-count convention. See Realtime and GPU operator support for additional boundaries.