Skip to main content
Back to timeline
arXivSource publication:

CoEvoWhen lets a frozen VLM coevolve policies and tools for ultra-long video, raising ExtremeWhenBench mIoU by 74.9% while cutting visual tokens by 11.4%

Synopsis

CoEvoWhen introduces a policy-tool coevolution framework that jointly evolves high-level policies and executable media tools from a VLM's agentic reasoning trajectories into a reusable skill without updating model parameters, improving ultra-long video temporal grounding accuracy and reducing inference visual token cost across five benchmarks and three VLMs, with the evolved skill transferring to general long-video QA without additional task-specific evolution.

AI-generated editorial illustration: CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding

Interpretation

The framework represents high-level policies and executable media tools together as one external skill and iteratively improves it from feedback formed by VLM execution trajectories paired with ground-truth intervals, keeping VLM parameters frozen throughout. Existing agentic methods rely largely on predefined policies and tool capabilities, whereas this work brings tool implementations themselves into the evolution loop, starting from a minimal base skill and having an external skill updater distill transferable experience and write code to upgrade or create tools. The paper provides a formal problem formulation and skill-update procedure, and states that evolution makes a single pass over 100 queries, forming a feedback batch after every four queries while VLM weights stay fixed.

The evolved skill lets the frozen VLM autonomously orchestrate image-based and video-based observations, yielding consistent gains on three ultra-long video temporal grounding benchmarks. Relative to the base skill, VUE-LVTR held-out IoU AUC rises from 0.4107 to 0.5137, ExtremeWhenBench mIoU and Recall@0.5 improve by 74.9% and 88.4%, and CoMET-Bench gains 0.0280 mIoU and 10.93 points of Rejection-F1. Results use each benchmark's official metrics, with ExtremeWhenBench evaluated on all 2,273 official test queries, CoMET-Bench retaining 1,599 queries with videos of at least 30 minutes, and VUE-LVTR evaluated on 307 queries disjoint from the evolution set.

Accuracy gains come together with lower inference visual cost, accompanied by a reallocation of the visual budget between image and video observations. Average visual tokens per query on the VUE-LVTR held-out set drop from 202.6k to 141.9k, with reductions of 11.4% on ExtremeWhenBench and 18.9% on CoMET-Bench; on ExtremeWhenBench image tokens rise from 179.3k to 197.5k while video tokens fall from 47.8k to 3.7k. Token cost is computed by summing the image and video tokens actually received by the VLM across rounds and averaging over queries, and model calls and local tool calls also decrease.

The evolved skill transfers across VLMs and across tasks: evolving separately on different backbones improves each, and applying the skill directly to long-video QA raises accuracy. On ExtremeWhenBench, mIoU improves by 99.2% for Qwen3.5-9B and 24.7% for Qwen3.6-27B; applying the grounding-evolved skill directly to LVBench and LSDBench raises Qwen3.5-27B overall accuracy by 8.65 and 7.98 points without task-specific evolution. Cross-VLM results compare base and evolved skills on three models evolved separately, cross-task results come from directly transferring the frozen skill, and ablations show joint evolution outperforms policy-only or tool-only evolution.

Perspective

The results target ultra-long video (tens of minutes to hours) temporal grounding and long-video QA, in a setting where a frozen VLM acts as the main agent acquiring visual evidence through tools. Evolution is performed on 100 labeled queries from a single data source, and evaluation covers three grounding benchmarks (VUE-LVTR, ExtremeWhenBench, CoMET-Bench) and two QA benchmarks (LVBench, LSDBench). What can be reused directly is the evolved skill (policy documents and callable tools) rather than model weights; cross-VLM use requires evolving separately on the corresponding backbone.

Evolution currently relies on labeled feedback from a single data source, so larger evolution scales, broader data sources, and evolution without labeled feedback remain open questions. All visual observations are currently performed directly by the VLM as the main agent, leaving open when it should observe directly versus delegate observation to subagents and how they should coordinate. The framework currently handles only the visual modality, and extension to omni-modal settings merits investigation. On CoMET-Bench, F1@0.5 remains lower for queries with more events, so precise multi-event localization is still difficult; Qwen3.6-27B shows a slight decrease on the Environment category and the 8-15 event group, indicating that the distribution of evolution gains depends on both model and query type.

Sources