Google proposes a multi-agent video co-director framework, reaching a peak quality score of 81.4 on GenAD-Bench and generating ten-minute long videos
Synopsis
A Google research team introduces a unified multi-agent framework (comprising AI video co-director, CANVAS, A²RD, and VQQA) that frames long-form video generation as a global optimization and world-state tracking problem, reporting a peak quality score of 81.4 on GenAD-Bench and improvements in multi-shot narrative consistency, character persistence, and long-duration stability on ST-Bench, HardContinuityBench, VBench-Long, LVBench-C, T2V-CompBench, VBench2, and VBench-I2V, alongside a demonstration of ten-minute video generation.
Interpretation
AI video co-director formalizes video storytelling as a global optimization problem, using hierarchical parameterization and a multi-armed bandit (MAB) to search creative configurations across creative strategy, narrative mode, and aesthetic archetype, with a multimodal LLM Judge feeding a factored reward signal back for iterative refinement. Compared with existing agentic pipelines that rely on rigid linear prompt chains driven by independent, handcrafted prompting, this delegates the choice of creative direction to global search with closed-loop feedback, so the whole pipeline operates under a unified vision. The text reports a peak quality score of 81.4 on GenAD-Bench and enhanced story consistency on ViStoryBench; the authors also state the architecture is model-agnostic and can sit on top of any foundation generative model, while noting that detailed architectures, training configurations, and baseline comparisons are in the individual papers.
CANVAS maintains structured representations of characters, locations, and object states plus a persistent visual memory, retrieving or initializing visual anchors as needed to explicitly plan visual continuity in multi-shot narratives. Compared with direct generation using the base model (Gemini-3.1-Pro) or an alternative multi-agent framework (AutoStudio), the text uses a museum heist sequence to show prop inconsistency and background drift in the former and character drift (the thief's cap disappears) plus background inconsistency in the latter, whereas CANVAS's persistent visual memory keeps characters, spatial geometry, and object states coherent across the narrative. Evidence comes from the comparison figure shown in the text (covering consecutive and non-consecutive transitions) and from reported continuity gains across scene reappearances on ST-Bench and HardContinuityBench.
A²RD is an agentic autoregressive architecture with segment-by-segment generation plus a multimodal video memory, running a retrieve-synthesize-refine-update loop per segment and adaptively switching between extrapolation (to advance the plot) and interpolation (to anchor segments to existing entities and environments). Compared with standard video generators that suffer severe visual decay over long durations (characters mutate, locations morph), A²RD continuously queries its multimodal video memory to maintain character identity, costume details, and structural geometry. The text demonstrates this with a ten-minute movie and reports improved character and environment consistency over continuous multi-minute runs on VBench-Long and LVBench-C by minimizing layout drift.
VQQA dynamically generates visual questions tailored to the specific prompt, uses VLM critiques as semantic gradients, iteratively refines the text prompt through a natural language interface, and employs a Global Selection mechanism that scores every candidate along the optimization trajectory against the original unedited prompt. Compared with test-time optimization methods that are typically computationally expensive or require white-box access to model internals, VQQA works as a black-box prompt optimizer that does not edit pixels directly but corrects high-level compositional defects such as attribute binding errors or inconsistent character attributes. The text gives two examples: rendering a realistic mylar balloon texture onto a strict cuboid geometry, and keeping the violinist and pianist consistently anchored to their respective instruments across cuts; it also reports notable absolute quality gains on T2V-CompBench, VBench2, and VBench-I2V.
Perspective
This work targets video creation scenarios that need multi-shot, minute-level to ten-minute-level coherent narratives, aimed at creators and production workflows that want to drive generation from high-level creative specifications rather than hand-maintaining shot-by-shot consistency; it is positioned as an orchestration layer that can sit on top of any foundation generative model and natively inherits safety mechanisms such as SynthID watermarking, with additional safety classifiers applicable across the final video in production to guard against unintended contextual interactions between individually safe clips. The authors say the next step is exploring how to integrate human-in-the-loop workflows, with the goal of keeping creators in control of creative direction and narrative design.
The text reports peak scores and gain directions without full numeric tables, sample sizes, or statistical details for each benchmark, so the magnitude and stability of the gains still need confirmation in the individual papers; the comparison figures and ten-minute video demonstration are largely qualitative, and readers may want to know more about behavior across different genres, different foundation models, and longer durations; in addition, human-in-the-loop workflows are still being explored, and how they would combine with the existing automated orchestration remains an open question.
