Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

OneStreamer unifies perception, memory, and timely response in streaming video with proactive generation, a 4B model topping eight benchmarks among compared methods

OneStreamer introduces a 4B streaming video LLM that jointly learns query-independent evidence recording and task response through a shared proactive generation process: Proactive Hierarchical Caption Memory produces time-grounded local-detail captions and summaries of completed events, Proactive State Transition Learning keeps supervision at all output anchors while supervising only 27.5% of annotated state tokens, and a streaming data synthesis pipeline yields OneStreamer-1M with over one million records, giving the best results among compared methods across eight streaming video understanding benchmarks.
arXiv

OneStreamer unifies perception, memory, and timely response in streaming video through a shared proactive generation interface, with a 4B model achieving the best results among compared methods across eight streaming video understanding benchmarks

OneStreamer jointly learns query-independent evidence recording and task response through a shared proactive generation process: its Proactive Hierarchical Caption Memory produces time-grounded local-detail captions and summaries of completed events, Proactive State Transition Learning outperforms dense state supervision while supervising only 27.5% of annotated state tokens, a streaming data synthesis pipeline yields the OneStreamer-1M dataset with over one million records, and the 4B model achieves the best results among compared methods across all eight evaluated streaming video understanding benchmarks, with ablations showing that retaining generated captions improves historical QA without degrading real-time perception.