Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Foresight lets a frozen Qwen3-VL-8B plan its own future perception in streaming video, reaching 23.0 joint F1 on OmniPro Online and beating the strongest trained baseline by 9.5%

The authors introduce Foresight, a training-free dual-stream architecture with two Siamese LLMs sharing weights, input encoders, and a KV cache, where one copy continuously ingests the stream while the other runs ahead over the same causal state to write a plan deciding when to reason next, what to check then, and how densely to sample; with a frozen Qwen3-VL-8B backbone it reaches 23.0 mean joint F1 on OmniPro Online (versus 13.5 for the strongest trained baseline, MiniCPM-o 4.5), improves the backbone by 6.7 points on StreamingBench with a best overall score, and by 15.4 points on OVO-Bench, with the largest gain of 18.7 points on Forward Active Responding where evidence arrives later in the stream.