vLLM-Omni unifies speech, diffusion, world-model and robot-policy serving under one orchestrator
Synopsis
The vLLM-Omni team presents an orchestrator-centered multi-stage serving runtime that splits omni, TTS, diffusion, world-model and VLA workloads into stages bound to individual engines and resource budgets, moves hidden states, codec codes, latents and paged KV through OmniConnector, and adds session control for long-lived interaction; it reports H100/H200 measurements for Qwen3-Omni speech serving, a TTS portfolio, image and video diffusion, and MiniCPM-o duplex health metrics.
Figure 9 : Qwen3-Omni Fixed-Len-2500/900 on 2 × 2{\times} H100: async-chunk versus no-async-chunk for mean TTFT, E2EL, and audio RTF.
arXivInterpretation
The system separates stages from an orchestrator: a stage binds one model component, one execution engine and an explicit resource budget, while the orchestrator alone owns request admission, cross-stage advancement, client-visible progress and completion, and explicitly does not schedule tokens inside a stage, manage paged KV, denoise diffusion steps or copy large tensors. Prior LLM stacks such as vLLM and SGLang target a single autoregressive decode loop, and diffusion stacks target DiT denoising; neither offers a shared control plane across AR, speech codecs, diffusion and action generators. This report makes orchestration, streaming progress, replica pools, intermediate-data movement and sessions explicit runtime concerns. Architectural description with module figures and section-by-section development rather than a controlled experiment; concrete payload examples come from the Qwen3-Omni thinker-to-talker-to-code2wav pipeline and diffusion DiT/VAE edges.
Intermediate data is managed as typed payloads: hidden states and embeddings, codec codes, paged KV blocks, and latents or multimodal tensors are assembled by stage runners and bridge processors, while OmniConnector exposes only a semantics-agnostic put/get/cleanup/health transport and the control plane carries lightweight metadata. KV is separated from the generic tensor path into a dedicated manager that owns extraction, keying and injection, so prefill-decode and AR-to-DiT edges reuse one transport instead of proliferating per-artifact networking stacks. Illustrated by the Qwen3-Omni thinker-to-talker edge carrying selected-layer hidden states, prompt embeddings and TTS special embeddings, and the talker-to-code2wav edge publishing codes.audio with left-context size and finished flags; the KV path references Mooncake-style backends.
async_chunk pre-warms downstream stages and delivers chunks on the data plane so waveform decoding overlaps codec prediction, lowering time to first audio packet; the mode is selected per edge and a batch full-payload path remains available. Streaming input and output are modeled as request progress under one request identity rather than an out-of-band channel; diffusion workloads can also accept a prompt update mid-generation, and long-lived interaction adds session identity, retention and admission on the same control plane. On Fixed-Len, disabling async-chunk produces multi-second TTFT and mean RTF approaching one at high concurrency, while the async recipe shows markedly lower TTFT and E2EL at the same load, indicating the overlap matters more as load grows.
The evaluation covers Qwen3-Omni speech serving, a TTS portfolio, image and video diffusion, Cosmos3/MiniMax video cells and MiniCPM-o duplex health metrics, comparing each model's default deployment with opt-in optimized profiles. It reports system-level readings for MRv2, a fused single-GPU Qwen3-TTS profile, Cache-DiT, Ulysses2+CFG2 and USP2+HSDP under one runtime rather than isolated single-model benchmarks. Speech tables are measured on H200; image, video and world-model tables come from the public H100 multimodal nightly CI at freeze day 2026-09-14; the report states that all reported cells complete successfully.
Perspective
The runtime targets engineering teams that need online multimodal generation serving: omni chat and speech, TTS, image and video diffusion, world-model rollouts, and OpenPI robot policies. The intended setting is one loaded multi-stage pipeline per process, with stages co-located or disaggregated across processes, devices and hosts, and per-edge choice between chunked and full-payload transfer. For long-lived interaction, the session layer adds identity, retention and admission on the same control plane, so duplex speech, world models and robot loops do not need a second runtime. Diffusion paged KV is now available on selected DiT pipelines under scheduler-managed block allocation, but coverage and cache-aware admission remain uneven across model families.
The report excludes cross-framework overlays and engineering-health aggregates from this revision, so it is hard to place these readings relative to other serving stacks. Several numeric values and concurrency levels appear as placeholders in the text and tables, so exact figures require checking the original tables. On Qwen3-Omni, MRv2 needs a one-line attribute guard to start, and on the text-only route its thinker exited and failed requests in two runs, so only the default deployment is reported. MiniCPM-o's turn-mode MRv2 profile is faster but appends unrelated speech after the target sentence in 13% of outputs, so it is not reported as a result. Diffusion paged KV coverage, cache-aware admission, broader session workloads, closed-loop robot evaluation and cross-hardware measurements are all listed as future work.
