OneStreamer unifies perception, memory, and timely response in streaming video through a shared proactive generation interface, with a 4B model achieving the best results among compared methods across eight streaming video understanding benchmarks
Related research and updatesSynopsis
OneStreamer jointly learns query-independent evidence recording and task response through a shared proactive generation process: its Proactive Hierarchical Caption Memory produces time-grounded local-detail captions and summaries of completed events, Proactive State Transition Learning outperforms dense state supervision while supervising only 27.5% of annotated state tokens, a streaming data synthesis pipeline yields the OneStreamer-1M dataset with over one million records, and the 4B model achieves the best results among compared methods across all eight evaluated streaming video understanding benchmarks, with ablations showing that retaining generated captions improves historical QA without degrading real-time perception.
Figure 1 : Evaluation results of OneStreamer. Blue bars highlight OneStreamer, with light purple portions indicating the Qwen3-VL baseline. OneStreamer achieves state-of-the-art results across all eight benchmarks, with an average relative improvement of 25.0% over the Qwen3-VL baseline and 6.1% over the strongest competing method on each benchmark.
arXivInterpretation
OneStreamer is introduced to jointly learn query-independent evidence recording and task response through a shared proactive generation process, so a streaming video LLM retains evidence before its relevance is known and responds when sufficient evidence becomes available. Prior streaming video understanding typically separates perception from memory; this work unifies them in one generation interface, supervising the interpretation of observed video prefixes with streaming caption targets during training. Method description at the abstract level plus ablation results: retaining generated captions improves historical QA without degrading real-time perception.
Proactive Hierarchical Caption Memory (PHCM) produces time-grounded local-detail captions and summaries of completed events; at inference, model-generated records complement a recent visual window, providing reusable factual context without revisiting historical visual features. Memory shifts from repeated access to historical visual features to model-generated textual records, forming reusable factual memory while preserving real-time perception. The abstract reports PHCM's composition and inference behavior, and an ablation indicates the effect of retaining generated captions on historical QA.
Proactive State Transition Learning (PSTL) reduces the dominance of repeated waiting states by preserving supervision at all output anchors and selecting representative state-change and state-persistence tokens. Compared with dense state supervision, PSTL performs better while supervising only 27.5% of annotated state tokens, indicating that selectivity of supervision matters more than coverage density. Comparison reported in the abstract: PSTL outperforms dense state supervision while supervising only 27.5% of annotated state tokens.
A streaming data synthesis pipeline aligns output content and timing with available evidence, and combining the resulting streaming captions and QA with cleaned open-source data yields OneStreamer-1M, a broad-coverage streaming video interaction dataset with over one million records spanning diverse tasks. The dataset treats content-and-timing alignment as a synthesis objective, providing broad-coverage training resources for streaming video interaction. The abstract reports over one million records spanning diverse tasks; the 4B model achieves the best results among compared methods across all eight evaluated benchmarks.
Perspective
This work targets streaming video interaction settings where evidence must be recorded before its relevance is known and responses are made when sufficient evidence becomes available, such as streaming video understanding and historical QA. Beneficiaries include researchers and engineers working on streaming video LLMs, real-time multimodal interaction, and long-horizon memory mechanisms. The applicable setting is: at inference the model complements a recent visual window with self-generated records without revisiting historical visual features, and training uses streaming caption targets and state transition supervision. OneStreamer-1M and the streaming data synthesis pipeline provide resources for subsequent training and evaluation on broader tasks.
The visible text is an abstract and does not include per-benchmark metrics, detailed model-scale comparisons, full ablation settings, or specific filtering criteria for dataset synthesis, so the magnitude of advantages and stability across tasks cannot be judged. Readers may watch for: the reusability of PHCM-generated records over longer videos and more complex event hierarchies, the robustness of PSTL's representative-token selection when state patterns change, and how OneStreamer-1M and the open-source data cleaning pipeline transfer to downstream tasks.
