OneStreamer unifies perception, memory, and timely response in streaming video with proactive generation, a 4B model topping eight benchmarks among compared methods
Synopsis
OneStreamer introduces a 4B streaming video LLM that jointly learns query-independent evidence recording and task response through a shared proactive generation process: Proactive Hierarchical Caption Memory produces time-grounded local-detail captions and summaries of completed events, Proactive State Transition Learning keeps supervision at all output anchors while supervising only 27.5% of annotated state tokens, and a streaming data synthesis pipeline yields OneStreamer-1M with over one million records, giving the best results among compared methods across eight streaming video understanding benchmarks.
Interpretation
PHCM turns observed content into two kinds of time-aligned textual records: </Observe> produces dense local-detail captions and </Summary> produces sparser semantic summaries of completed events; at inference these complement a Recent-FIFO visual window, supplying distant factual context without retrieving or revisiting historical visual features. Prior memory mechanisms either compress or retrieve distant visual evidence so that retained historical visual tokens compete with recent observations for limited context, or rely on a strict recent window that discards distant events; this work separates query-independent memory writing from later task-conditioned use and carries past events in compact text. Ablations with the same checkpoint and Recent-16 window show that retaining generated records raises OVOBench-Backward ASI from 63.5 to 71.6 and EPM from 62.0 to 62.6; PHCM also lifts OVOBench Real-Time from 80.9 to 81.4 and StreamingBench Real-Time from 86.3 to 86.9, and exceeds Full on ASI (71.6 vs. 67.6) while its EPM is slightly lower than Full (62.6 vs. 63.0).
PSTL preserves supervision at all output anchors and uniformly subsamples larger transition groups to a shared quota, selecting representative state-change and state-persistence tokens so that repeated waiting states no longer dominate the state objective. Dense state supervision overemphasizes waiting because repeated silence tokens outnumber output-initiating tokens, while supervising only state changes omits direct supervision of state persistence; PSTL covers both, and unselected state tokens remain in the causal sequence and are excluded only from the state loss. Under the same training data, input sequences, and text supervision, PSTL supervises only 27.5% of annotated state tokens yet achieves the best results on three benchmarks: ProactiveVQA 48.7, OmniMMI 36.6, and OVO-Timing 41.6; a supervision-matched random sparse baseline reaches only 26.6, 27.4, and 1.4, indicating the gains are not due to sparsity alone.
The team built a reusable streaming data synthesis pipeline that converts offline video resources into streaming training sequences: the caption branch releases local descriptions only after their supporting intervals are observed and segment summaries only after events complete, while the QA branch localizes a coarse evidence interval, verifies answerability, and then calibrates response time with sliding windows. Existing streaming QA data often provide poorly calibrated response-time supervision, with some targets assigned before sufficient evidence is available and others assigned to the end of a grounding interval even when evidence suffices much earlier; this work separates evidence localization from response-time calibration. Calibration advances response time by 13.72 s on average relative to the original annotation or the end of the evidence interval, with 86.62% of records receiving an earlier response time; combined with cleaned open-source data this yields OneStreamer-1M, whose listed records total 1,160,004, decontaminated against every evaluated benchmark at the source-video level.
The 4B OneStreamer achieves the best results among compared methods across eight streaming video understanding benchmarks, spanning perception and memory as well as proactive response. Relative to the size-matched Qwen3-VL base model and the 11B MOSS-VL-Realtime, it improves real-time perception, long-range online understanding, and proactive response, supporting proactive generation as a shared learning interface connecting perception, memory formation, and timely response. The four perception and memory benchmarks are OVOBench 72.1, StreamingBench Real-Time 86.9, OVBench 66.8, and ODVBench 71.3, exceeding the size-matched Qwen3-VL base by 13.3, 5.1, 11.4, and 13.7 points; the four proactive-response benchmarks are ProactiveVQA 48.7, OmniMMI 36.6, OVO-Timing 41.6, and ViSpeak 2.87, with the first three exceeding the base by 14.4, 7.2, and 12.2 points.
Perspective
This work targets settings that must observe an open video stream continuously and respond when evidence becomes sufficient, such as live-stream assistants, wearable agents, and security monitoring; its conclusions apply to a configuration initialized from Qwen3-VL-4B-Instruct, trained with single-stage supervised fine-tuning on OneStreamer-1M plus offline data, with the vision encoder frozen and the projector and LLM optimized. PHCM's gains were measured with a Recent-16 visual window and a hierarchical caption memory built from 128 frames at 4 fps, and the efficiency numbers come from a 360 s sample and an online throughput test on a single H200. The data synthesis pipeline relies on Gemini and Seed models for multi-granularity annotation and QA generation and on Qwen models for evidence-interval verification and response-time calibration, so reproducibility depends on those external models.
The paper's own failure case shows the model can misread object-state changes and, even when later frames make the state clear, does not revise its earlier description, indicating that state-change interpretation and error correction remain open questions. Response-time calibration is described by the authors as not always identifying the earliest answerable moment, trading off timeliness against evidence sufficiency. In ablations, PHCM's EPM is slightly below the full-visual-history configuration, suggesting the two representations emphasize different metrics. In addition, the loaded text includes the main body, appendices, and per-benchmark result tables but not the image content of Figures 3 and 5 through 12, so understanding of the architecture diagram and qualitative cases rests on the prose descriptions.
