Skip to main content
Back to timeline
arXivSource publication:

Realtime-Venus splits real-time dialogue from background work into a dual-loop runtime with two 9B frontends, leading online models on six of eight video benchmarks

Synopsis

The work presents Realtime-Venus, a proactive full-duplex interaction system with two separately trained 9B models, Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction, which align continuous perception, conversational control, native speech generation, and private delegation requests on a shared causal timeline while Realtime-Venus-Harness executes background capabilities asynchronously in a dual-loop runtime, achieving the highest scores among compared online models on six of eight video benchmarks and leading or matching on several audio understanding, spoken question answering, and full-duplex overlap metrics.

AI-generated editorial illustration: Realtime-Venus: A full-duplex interaction system with asynchronous delegation

Interpretation

It introduces two separately trained 9B full-duplex frontends: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction, each serving as a complete conversational frontend that integrates continuous perception, conversational control, and native speech generation in a single autoregressive interaction loop. Relative to prior work such as Moshi, Qwen2.5-Omni/Qwen3-Omni, and MiniCPM-o 4.5, the authors place conversational control, response generation, and private delegation within a shared policy and shared timeline, and state that Realtime-Venus-Omni is the first full-duplex omni model to support asynchronous backend invocation for reasoning and tool execution while maintaining video interaction. The paper describes the architecture (SigLIP2, Whisper-Medium, a Qwen3-8B backbone, S3 speech tokens, and a streaming flow-matching decoder), the one-second chunk schedule, and the <|listen|>/<|speak|>/<|turn_eos|> control tokens, and reports the highest scores among compared online models on six of eight video benchmarks.

It designs Realtime-Venus-Harness as a framework for asynchronous capability execution and result delivery, binding each delegation request to its originating session, fixing an evidence snapshot at the request boundary, routing the task to a registered capability, and returning results to the originating session through a private background channel. Relative to ReAct-style interleaving of reasoning and actions and to routes such as DuplexSLA, MoshiRAG, DuplexOmni, and JoyAI-VL-Interaction that connect tools or retrieval to dialogue, the framework separates the execution boundary from the delivery timing: background work uses evidence fixed at the request boundary, while the frontend interprets the returned result using the conversational state at reintegration and decides when to speak the prepared reply. The paper details the capture, dispatch, and return stages, the work-item lifecycle (Queued, Running, Completed, Failed, Delivering, Delivered), freshness checks and playback acknowledgment, plus capability contracts and a normalized result format.

It builds a data pipeline coupling duplex and delegation behaviors, combining scenario planning, paralinguistic and acoustic realization, and temporal alignment into unified trajectories that distinguish backchannels, other-directed speech, background speech, and interruptions on one timeline. Relative to treating overlapping speech as a uniform stop signal or relying on binary VAD reactions, the pipeline annotates events by conversational function and links actions such as continuing to speak, awaiting input, yielding the floor, stopping queued output, and revising a response with delegation decisions such as retaining, canceling, revising, retrying, or replacing work, all tied to the event's effect on user intent. The paper reports a post-training corpus of over 2.8 million samples across nine data categories, with the video group about 70% and the audio group about 30%, offline understanding about 56%, proactive duplex about 37%, and delegation about 6%, and notes that supervision covers only response spans and that only the Thinker is updated.

In evaluation, Realtime-Venus-Omni scores 70.2% on StreamingBench, 64.7% on OVO-Bench, and 81.3% on Daily-Omni, the highest among compared online models on six of eight video benchmarks, while Realtime-Venus-Audio reaches 78.0% on MMAU, 63.2% on MMAU-Pro, 83.8% on Llama Questions, and 67.8% on Speech CMMLU, and on Full-Duplex-Bench v1.5 responds to 75% of user interruptions with continuation rates of 97%, 88%, and 86% under backchannels, other-directed speech, and background speech. Relative to MiniCPM-o 4.5, Realtime-Venus-Omni improves StreamingBench and OVO-Bench by 2.3 and 4.0 percentage points and OmniPro probe-mode accuracy by 3.4 points, and it exceeds Gemini 3.1 Live and GPT-4o on all three continuation metrics of Full-Duplex-Bench v1.5. Results are given in tables that separate online from offline baselines and report the four overlap scenarios individually; the paper also notes lower scores than MiniCPM-o 4.5 on ProactiveVideoQA and WorldSense and the absence of component ablations.

Perspective

The results target digital and physical interaction settings that require continuous perception and timely responses, apply to streaming audio-visual input organized in one-second chunks, and cover conversational tasks that need external reasoning or tool execution; the dual-loop runtime and registered capability contracts let one delegation interface serve both the audio and omni frontends, and the training-free long-video memory module targets hour-scale sessions. For users, this means complex requests can be handed to background execution while foreground dialogue continues, with the frontend choosing when to speak the prepared reply.

The paper flags several points to watch: lower scores than MiniCPM-o 4.5 on ProactiveVideoQA and WorldSense indicate that streaming comprehension, proactive response quality, and offline video understanding need separate assessment; it is not the top score on MMAR and MMSU, suggesting room in contextual acoustic reasoning and fine-grained spoken-language understanding; its interruption-response rate is below Joy-Duplex, GPT-4o, and Gemini 3.1 Live, showing that strong continuation does not imply strong interruption handling; on tool use its argument accuracy and Pass@1 trail GPT-Realtime, and because the metrics use different success criteria their difference cannot be read as an intermediate-stage failure rate; the delegation benchmark is internally constructed, and the two frontends show different patterns of missed versus unnecessary delegation; memory gains depend on the backbone and evaluation setting; and the paper states that results do not isolate the contribution of individual training components, which requires controlled ablations, with future work on finer-grained chunks, longer context, and multi-step concurrent tool execution.

Sources