Public articles linked to the same research event.
arXiv OneStreamer introduces a 4B streaming video LLM that jointly learns query-independent evidence recording and task response through a shared proactive generation process: Proactive Hierarchical Caption Memory produces time-grounded local-detail captions and summaries of completed events, Proactive State Transition Learning keeps supervision at all output anchors while supervising only 27.5% of annotated state tokens, and a streaming data synthesis pipeline yields OneStreamer-1M with over one million records, giving the best results among compared methods across eight streaming video understanding benchmarks.
OneStreamer introduces a 4B streaming video LLM that jointly learns query-independent evidence recording and task response through a shared proactive generation process: Proactive Hierarchical Caption Memory produces time-grounded local-detail captions and summaries of completed events, Proactive State Transition Learning keeps supervision at all output anchors while supervising only 27.5% of annotated state tokens, and a streaming data synthesis pipeline yields OneStreamer-1M with over one million records, giving the best results among compared methods across eight streaming video understanding benchmarks.
OneStreamer introduces a 4B streaming video LLM that jointly learns query-independent evidence recording and task response through a shared proactive generation process: Proactive Hierarchical Caption Memory produces time-grounded local-detail captions and summaries of completed events, Proactive State Transition Learning keeps supervision at all output anchors while supervising only 27.5% of annotated state tokens, and a streaming data synthesis pipeline yields OneStreamer-1M with over one million records, giving the best results among compared methods across eight streaming video understanding benchmarks.
OneStreamer introduces a 4B streaming video LLM that jointly learns query-independent evidence recording and task response through a shared proactive generation process: Proactive Hierarchical Caption Memory produces time-grounded local-detail captions and summaries of completed events, Proactive State Transition Learning keeps supervision at all output anchors while supervising only 27.5% of annotated state tokens, and a streaming data synthesis pipeline yields OneStreamer-1M with over one million records, giving the best results among compared methods across eight streaming video understanding benchmarks.
arXiv OneStreamer jointly learns query-independent evidence recording and task response through a shared proactive generation process: its Proactive Hierarchical Caption Memory produces time-grounded local-detail captions and summaries of completed events, Proactive State Transition Learning outperforms dense state supervision while supervising only 27.5% of annotated state tokens, a streaming data synthesis pipeline yields the OneStreamer-1M dataset with over one million records, and the 4B model achieves the best results among compared methods across all eight evaluated streaming video understanding benchmarks, with ablations showing that retaining generated captions improves historical QA without degrading real-time perception.
OneStreamer jointly learns query-independent evidence recording and task response through a shared proactive generation process: its Proactive Hierarchical Caption Memory produces time-grounded local-detail captions and summaries of completed events, Proactive State Transition Learning outperforms dense state supervision while supervising only 27.5% of annotated state tokens, a streaming data synthesis pipeline yields the OneStreamer-1M dataset with over one million records, and the 4B model achieves the best results among compared methods across all eight evaluated streaming video understanding benchmarks, with ablations showing that retaining generated captions improves historical QA without degrading real-time perception.
OneStreamer jointly learns query-independent evidence recording and task response through a shared proactive generation process: its Proactive Hierarchical Caption Memory produces time-grounded local-detail captions and summaries of completed events, Proactive State Transition Learning outperforms dense state supervision while supervising only 27.5% of annotated state tokens, a streaming data synthesis pipeline yields the OneStreamer-1M dataset with over one million records, and the 4B model achieves the best results among compared methods across all eight evaluated streaming video understanding benchmarks, with ablations showing that retaining generated captions improves historical QA without degrading real-time perception.
OneStreamer jointly learns query-independent evidence recording and task response through a shared proactive generation process: its Proactive Hierarchical Caption Memory produces time-grounded local-detail captions and summaries of completed events, Proactive State Transition Learning outperforms dense state supervision while supervising only 27.5% of annotated state tokens, a streaming data synthesis pipeline yields the OneStreamer-1M dataset with over one million records, and the 4B model achieves the best results among compared methods across all eight evaluated streaming video understanding benchmarks, with ablations showing that retaining generated captions improves historical QA without degrading real-time perception.