VideoGen-Agent uses multitask reinforcement learning to teach a video-generation agent to call retrieval, simulation, and verification tools, lifting its base generator from 56.5 to 75.6 on VABench and to 86.1 with upgraded generation tools
Synopsis
The work introduces VideoGen-Agent, a multimodal agent trained with supervised fine-tuning followed by multitask agentic reinforcement learning to coordinate augmentation, generation, and verification tools across six video-generation tasks (Procedural Knowledge, Single-Entity Identity, Multi-Entity Identity, Physics Simulation, Compositional Scene, and Multi-Shot), and builds VABench, a held-out benchmark of 600 prompts; on VABench it improves the base T2V generator Seedance 1.0 from 56.5 to 75.6, and further to 86.1 when compatible upgraded generation tools unseen during training are used, with human raters preferring that configuration in 100 pairwise comparisons.
Interpretation
VideoGen-Agent uses one shared multimodal policy to coordinate augmentation, generation, and verification tools across six video-generation tasks, selecting tools per prompt and intermediate observation through a reasoning-action-observation loop. Prior agentic video-generation approaches often target specific objectives, whereas this work unifies tool selection and use in a single policy spanning Procedural Knowledge, Single-Entity Identity, Multi-Entity Identity, Physics Simulation, Compositional Scene, and Multi-Shot. The paper provides a mapping from the six task categories to default tool subsets and typical workflows, such as text retrieval plus T2V for Procedural Knowledge, code plus M2V for Physics Simulation, and T2V plus ELF plus I2V for sequential Multi-Shot generation.
Training proceeds in two stages: supervised fine-tuning on 16K teacher-distilled trajectories, then GRPO-based multitask agentic reinforcement learning on the remaining 8K prompts, with a hybrid reward combining format, VLM, and tool-use rewards and with task-level advantage normalization. The work extends agentic reinforcement learning from question answering and information seeking to video generation, where the reward evaluates both tool-call validity and generated video quality, and task-level normalization balances learning signals across categories. The paper gives the hybrid reward equation, group-relative and task-level advantage normalization equations, and the token-level GRPO objective, plus ablations: removing the VLM reward lowers the overall score to 73.3 and removing the tool reward to 71.2, both below the full configuration's 75.6.
On VABench, VideoGen-Agent with Toolset 1 reaches 75.6, exceeding the base T2V generator Seedance 1.0 at 56.5 and the strongest standalone baseline Seedance 2.0 at 73.2; with Toolset 2 it reaches 86.1 and achieves the highest score in every category. The work introduces VABench, a held-out benchmark of 600 prompts with 100 per category, evaluated with category-specific VLM rubrics; tool upgrades improve performance without additional agent training, indicating complementary benefits from learned tool use and advances in generation models. The main results table lists per-category and overall scores for each baseline; the tool upgrade raises Multi-Entity Identity from 65.3 to 86.7, the largest increase; human raters prefer the Toolset 2 configuration in 100 pairwise comparisons.
Ablations show prompt rewriting and zero-shot tool use yield limited gains (57.3 and 59.5), SFT raises the score to 69.2, and RL adds further to 75.6; single-task RL reaches 76.3 overall, slightly above multitask RL's 75.6, while multitask RL covers all six tasks with one agent. The work separates and quantifies the contributions of training stages, reward components, and single-task versus multitask settings, showing that tool access and prompt elaboration alone do not account for the main improvement, whereas trajectory supervision and reward optimization do. The ablation table lists per-category scores for each variant; the paper notes the zero-shot agent frequently fails to invoke tools correctly despite workflow guidance; single-task RL scores higher on Procedural Knowledge, Multi-Entity Identity, and Compositional Scene, while multitask RL scores higher on the remaining categories.
Perspective
The result targets video-generation settings that need external knowledge and tools, covering six tasks: Procedural Knowledge, Single-Entity Identity, Multi-Entity Identity, Physics Simulation, Compositional Scene, and Multi-Shot. For teams wanting to upgrade generation tools without redesigning workflows, the paper shows the same trained policy can be used with compatible new generation tools, and it demonstrates through continued training that an action-to-video pipeline can be added for robotic manipulation. Evaluation centers on VABench's 600 prompts with category-specific VLM scoring, supplemented by a human preference comparison on 100 video pairs.
The paper states that VideoGen-Agent is currently limited by the quality and latency of its generation tools, the coverage of its verification feedback, and its six-task training setting. For Physics Simulation, the VLM assesses plausibility of contact, rebound, motion direction, and induced rotation, and does not verify exact masses, velocities, or accelerations from appearance alone, with quantitative checks provided by separate code. For Multi-Shot, requested durations are checked against tool support and timestamps adjusted before generation. Reward-model agreement with human preferences is lowest for Procedural Knowledge, which the paper notes remains a challenge for the judge. In addition, some numbers appear as placeholders in the loaded text, such as the improvement points and human preference rate in the main results paragraph, so those specific figures cannot be given in this summary.
