Skip to main content
Back to timeline
arXivSource publication:

Foresight lets a frozen Qwen3-VL-8B plan its own future perception in streaming video, reaching 23.0 joint F1 on OmniPro Online and beating the strongest trained baseline by 9.5%

Related research and updates

Synopsis

The authors introduce Foresight, a training-free dual-stream architecture with two Siamese LLMs sharing weights, input encoders, and a KV cache, where one copy continuously ingests the stream while the other runs ahead over the same causal state to write a plan deciding when to reason next, what to check then, and how densely to sample; with a frozen Qwen3-VL-8B backbone it reaches 23.0 mean joint F1 on OmniPro Online (versus 13.5 for the strongest trained baseline, MiniCPM-o 4.5), improves the backbone by 6.7 points on StreamingBench with a best overall score, and by 15.4 points on OVO-Bench, with the largest gain of 18.7 points on Forward Active Responding where evidence arrives later in the stream.

AI-generated editorial illustration: Foresight: planning future perception in streaming VLMs without retraining

Interpretation

The paper reframes anticipation from a prediction target into a means of controlling computation: at each step the model identifies what remains uncertain, decides when to look again, and sets how densely to sample until then, turning streaming inference into a closed loop rather than a fixed sequence of computations. Earlier proactive streaming methods such as StreamAgent do feed predictions back into computation, but their planning runs on a fixed schedule; in Foresight a single plan generated by the VLM carries when the next reasoning step happens, the question that step should resolve, and the sampling density in between. Evaluated with a frozen Qwen3-VL-8B backbone on three streaming benchmarks (OmniPro Online, StreamingBench, OVO-Bench); Appendix F reports a lead-time experiment in which, two seconds before an event, the pooled score closes about three quarters of the gap between the blind baseline and the oracle, falling to about a tenth by 30 seconds.

A dual-stream Siamese architecture lets planning run concurrently with perception: the Ingest copy is the only writer to the shared KV cache, while the Think copy reads a read-only snapshot and writes plans, so planning never blocks ingestion. Other asynchronous systems use different LLMs or encoders and must duplicate video encoding and processing; Foresight uses asynchronous twin VLMs over a shared KV cache and separates a transient evidence state from a persistent control update. Algorithm 1 presents the concurrent perception and planning loops; ablations show that ignoring all plan fields drops joint F1 from 23.0 to 0.7 and removing the Answer field drops it to 0.3, indicating the control interface carries much of the effect.

Online reconfiguration is achieved with low overhead: the plan decodes only the values that require reasoning, boolean decisions are read as a yes/no logit-ratio probe, later steps decode only a diff against the current plan, and a Live Executor applies the result to the input gate, the encoder, and the cache. Compared with controllers that rebuild context at every check, the shared KV cache prefills each visual token exactly once, and the diff plus probe keep each check to a constant few tokens. The cost model in Appendix G gives order-of-magnitude differences among design choices (for example, dropping the shared KV cache raises total compute by about an order of magnitude) and notes these factors come from the model rather than measurement; measured real-time factors are 0.053–0.084 for QueryStream and 2.13–2.70 for StreamAgent.

This inference-time controller improves both proactive responding and general streaming understanding, with the largest gains where evidence arrives later. Foresight reaches 23.0 mean joint F1 on OmniPro Online against 13.5 for the strongest trained baseline, MiniCPM-o 4.5; it attains the best overall score of 66.00 on StreamingBench and 62.16 overall on OVO-Bench, with an 18.7-point gain on Forward Active Responding. All Foresight numbers are computed on an 80% held-out split of each benchmark, with thresholds fitted only on a 20% calibration split; the authors treat StreamingBench's Proactive Output as a diagnostic because its confidence intervals are wide.

Perspective

The result targets streaming video understanding that requires proactive, asynchronous responses, such as real-time human-AI interaction where the system must respond to events without being explicitly prompted. The method wraps a frozen Qwen3-VL-8B backbone and uses Qwen3-VL's native video-input pipeline, grouping consecutive admitted frames into two-frame clips as video inputs; because the backbone takes video only, the OmniPro evaluation uses the subset that does not need audio. No weights are updated, so the same controller can transfer directly to stronger backbones. Plan fields persist in a standing controller state, and omitted fields retain their previous values, so frame rate and the open question stay active across later steps.

Anticipation reliability is bounded by the underlying VLM: the authors state that Foresight's anticipation and temporal event localization depend on the backbone's capacity, and reliable operation requires a VLM with strong future anticipation and temporal grounding. After an initial false trigger, the model can occasionally keep reporting an event that later observations do not support; giving the Think LLM a history of what has been observed and reported reduces this substantially but does not eliminate it. The controller also remains sensitive to prompt wording. Appendix F shows current VLM foresight fades within seconds, so the controller must spend some checks on moments where nothing happens, and reaching the ideal rate of one check per event would need a model whose anticipation lasts much longer. The cost curves in Appendix G use representative token counts and assumed check-rate distributions, which the authors note measured benchmark check rates could replace. StreamingBench's Proactive Output has few questions per model and wide confidence intervals, so the authors treat it as a diagnostic rather than a headline result. This summary is based on the paper's full text and does not include every per-item numeric detail from figures and appendices.

Sources