VideoLoop uses dual-loop bounded working memory to ease semantic thrashing in long-video agents, lifting Gemini 3.1 Pro by 3.2 to 4.5 points on VideoMME (long) and two other benchmarks
Synopsis
The work formalizes a failure mode it calls semantic thrashing, in which append-only working memory keeps accumulating noise and dilutes key evidence, and proposes VideoLoop: an outer multimodal agent explores the video inside a sandboxed filesystem while an inner memory orchestrator retrieves relevant artifacts from an unbounded filesystem and rewrites a bounded working memory after every step; across VideoMME (long), VideoMMMU, and LongVideoBench (long), VideoLoop improves four LVLM backbones in a plug-and-play manner, averaging a 4.2-point gain over baseline on VideoMME (long) and reaching 88.3%, 88.8%, and 80.9% with Gemini 3.1 Pro.
Interpretation
The paper gives a structural argument that append-only working memory can incorporate newly observed target evidence but cannot remove accumulated irrelevant evidence or prevent ordered context growth without a rewrite operator. Most prior video agents use a single reasoning loop with append-only memory and treat memory degradation as a tuning issue; this work frames semantic thrashing by analogy to OS thrashing and decomposes the gap between memory and the optimal evidence set into missing target evidence and redundant noise, arguing the problem is structural to append-only updates. The argument rests on formal derivations (Eq. 2 through Eq. 4) and a conceptual diagnostic; the authors explicitly call it a conceptual diagnostic rather than a theorem-like reduction, and it motivates rather than proves the design.
VideoLoop uses a decoupled dual-loop design: the outer loop has a policy model explore the video and write observations and intermediate analysis to a persistent filesystem, while the inner loop retrieves question-relevant artifacts after each step and rewrites a bounded working memory. Unlike systems built on predefined pipelines or prebuilt indices, VideoLoop runs inside a coding sandbox and writes code to direct its own exploration with only four tool primitives (Analyze, Transcribe, Execute, Answer), requiring no upfront video preprocessing or fixed schema; compared with memory mechanisms such as VideoMem, WorldMM, and MemGPT, it assigns video exploration and working-memory maintenance to a dedicated inner agent. Method details are complete: a 32K-token working-memory budget, at most 6 selected key frames, 6 to 50 outer reasoning iterations, a six-section memory document, and a three-tier filesystem; the rewrite benefit is stated as a condition in Eq. 7 and Eq. 8, which the authors note is a condition on rewrite quality, not an unconditional guarantee.
On three long-video benchmarks, VideoLoop improves multiple LVLM backbones in a plug-and-play manner and substantially mitigates memory degradation on the hardest questions. Relative to native Gemini 3.1 Pro, it gains 4.5 points on VideoMME (long), 4.2 on VideoMMMU, and 3.2 on LongVideoBench (long); margins over the strongest prior agentic methods are 7.1, 10.4, and 4.5 points. Ablations show the append-only agent at 81.9%, dual-loop rewriting at 83.3%, and the full filesystem design at 85.8%. Evaluation covers 900 VideoMME (long) questions, 900 VideoMMMU questions (300 per cognitive track), and 564 LongVideoBench (long) questions; difficulty quartiles were sorted by Claude Opus 4.8 with human review. Memory quality is probed by a blind judge that reads only a frozen context snapshot without video or tools: on Q4 it answers 81.1% for VideoLoop versus 60.9% for the append-only agent.
Memory orchestration also improves token efficiency and exploration behavior: the filesystem externalizes intermediate evidence and reuses it selectively instead of repeatedly carrying the full history in context. With Gemini 3 Flash, the append-only baseline uses 614.9K tokens per question at 81.9% accuracy; dual-loop only uses 647.0K at 83.3%; full VideoLoop uses 618.2K at 85.8%, with input tokens dropping from 584.6K to 559.0K. Behavior analysis shows VideoLoop's viewed-frame distribution concentrates in the low-cost regime, uses fewer iterations per question, and spreads frames more evenly across the timeline. The token and accuracy comparison is a controlled comparison on the same benchmark; the behavior analysis rests on observed trends in viewed-frame distribution and iteration counts, which is observational evidence.
Perspective
The result targets multimodal agents that must gather evidence over long videos across many steps, and it fits deployments that can provide a sandboxed filesystem and tool calls; the method acts on existing LVLM backbones in a plug-and-play way without retraining, so it is most directly usable by teams that want better long-video question answering without changing model weights. The authors generalize the conclusion into a principle for long-horizon agents, that memory should be curated rather than accumulated, which offers a transferable design idea for dialogue agents, document agents, and other systems that must retain evidence across many steps. The evaluation setting is VideoMME (long), VideoMMMU, and LongVideoBench (long), with a 32K-token working-memory budget, at most 6 key frames, and 6 to 50 outer iterations, and these configurations define the scope in which the results apply.
The rewrite benefit in Eq. 8 is a condition rather than an unconditional guarantee, so the rewrite quality of the inner orchestrator remains a key variable; the paper uses a blind judge reading only a context snapshot as a retrievability proxy, and how that proxy relates to final answer accuracy deserves observation on more tasks. Difficulty quartiles were sorted by Claude Opus 4.8 with human review, so whether the Q1-to-Q4 gaps hold under other stratification schemes remains to be tested. The token-efficiency conclusion comes from a single comparison on Gemini 3 Flash, and the cost structure under other backbones and longer videos remains an open question. In addition, the readable text here is the full paper plus the homepage abstract without raw figure or table data, so specific numbers for viewed-frame distribution and iteration counts can only be described as reported in the prose.
