Realtime validation¶
ROCm build checks¶
The ROCm image workflow checks real installed Omni imports during the Docker build, before exporting the large image or a cache. A hash-bound repair skips the pinned upstream's CUDA-only shutdown patch on HIP/CPU builds; it does not mock a GPU or suppress runtime errors. The generated one-/two-device stage-contract tests still run after the build. Registry caches are branch-specific, contain final-image layers only, and are written only after successful checks. None of these checks replaces GPU or audio acceptance below.
The final package/license inventory scans the tested container's read-only filesystem with installed-package catalogers and no network. This avoids Syft's Docker-daemon TAR export (and a second full-size image copy). Reports retain the tested image ID, source revision and SBOM checksum; SPDX/CycloneDX conversion and the existing license review remain required before publication.
Protocol limits and acceptance¶
The selected Qwen plugin supports audio/text turns, server VAD or explicit commits, concurrent input and interruptible audio output. It does not provide model-native continuous input KV processing: each committed turn creates an inference request. Realtime tool calling and reference-voice audio are not implemented by this plugin. OpenAI compatibility is not full feature parity; named voices, session resumption across proxies and every Playground option must be checked with the selected runtime.
Local tests cover configuration persistence, open experimental selection, scheduling resource bounds, stage budgets, readiness, lifecycle, slot accounting, log ownership, catalog publication and the dashboard form. They do not download weights or run CUDA/ROCm. Before declaring the installation accepted:
- Confirm the image pulls and all three stages load on the intended GPU(s).
- In sharing mode, verify the assigned NVIDIA slot or AMD claim/CDI binding, concurrent inference with another model, and manual memory budgets. Confirm exhausted slots block new starts and Stop releases only this model's slot.
- In the LiteLLM Realtime Playground, verify microphone input, intelligible audio output, a second turn with context, and barge-in/server VAD.
- Verify the authenticated Gateway WebSocket path as well as localhost testing; confirm an unauthorized API key cannot open the session.
- Stop/start/restart the model, confirm route withdrawal/recovery and GPU release, and check that another ordinary Ollama/vLLM model still answers requests.
- Record the source digest, GPU/driver, settings, logs and outcomes. A Ready
/healthalone is not proof of successful speech inference.
Primary sources: