Skip to main content
Back to timeline
arXivSource publication:

EVO-WAM lets world action models self-train on their own generated video, lifting RoboTwin unseen-task success from 26.9% to 68.0%

Synopsis

The work proposes EVO-WAM, a framework that enables complete autoregressive rollouts without external execution feedback via state prediction and anchored multi-frame context, then selects task-completing prefixes with a vision-language model and verifies video-action consistency with an inverse dynamics model for iterative self-training; on seven unseen RoboTwin 2.0 tasks it raises average success from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, and on three real-world long-horizon composite tasks it raises Cosmos3 from 20.0% to 76.7%.

AI-generated editorial illustration: EVO-WAM: Evolving World Action Models through Video-Action Verification

Interpretation

EVO-WAM improves world action models on unseen tasks using only their own generated video-action trajectories, without additional expert demonstrations and without executing candidate actions in an external environment during self-evolution. Prior methods that generate experience for policy improvement still rely on a separate world model or task demonstrations, whereas this work uses the WAM's own joint video-action predictions as the supervision source. Validated on two backbones, Cosmos3 and DreamZero, across seven unseen RoboTwin 2.0 tasks with 100 Clean and 100 Randomized trials per task; real-world evaluation uses a Franka robot with ten trials per task.

Two-stage verification, where a vision-language model selects task-completing prefixes and an inverse dynamics model checks video-action consistency, is key to substantial self-training gains. The ablation shows VLM-only visual completion reaches 43.7% in Round 4 versus 68.0% with IDM added; VLM + Simulator reaches 72.7% but requires a simulator of the target task, while the IDM approach transfers to the real world. Controlled comparison from the same Cosmos3 30K initialization, 2,800 candidates per round, four rounds of 1,000 updates, and a 1:1 recorded/generated sampling ratio; verifier reliability analysis shows IDM reduces false acceptances from 63 to 15, raising precision from 68.0 to 86.0 while recall falls from 64.4 to 44.2.

Self-training gains are largest in the first round, later rounds bring smaller gains and occasional regressions, but accumulating verified prefixes is more stable than using only the latest round's data. Cosmos3 on RoboTwin rises from 26.9% to 58.3% in Round 1 and reaches 68.0% by Round 4; latest-round-only data falls back to 53.6% in Round 4 versus 68.0% with accumulated data. Four-round self-training sequence with an accumulated-versus-latest data comparison; on the real robot, Rounds 2 and 4 both reach 76.7% while Round 3 reaches 73.3%.

Improvements generalize to new scenes while largely preserving seen-task performance. On 100 newly sampled scenes per task-condition, EVO-WAMCosmos3 reaches 70.4% versus 24.9% for the baseline; on the 43 seen tasks it reaches 84.8% versus 85.8% for the 34K Cosmos3 baseline. New-scene evaluation uses 100 Clean and 100 Randomized trials per task, 1,400 in total; seen-task evaluation uses 50 Clean and 50 Randomized trials per task.

Perspective

The result is aimed at robot-learning researchers and engineering teams who want to adapt world action models to unseen tasks without collecting new demonstrations, in settings where a WAM backbone can generate video-action rollouts and an inverse dynamics model can be trained or obtained; the simulation side is validated on seven unseen RoboTwin 2.0 tasks, and the real side on a Franka robot across three long-horizon composite tasks. Because the method does not require executing candidate actions during self-training, it also suits real environments where trial-and-error is hard to run safely.

Later rounds bring smaller gains and occasional regressions, showing that additional self-training does not always improve performance, so when to stop and how to select rounds remain open questions; DreamZero's limited improvement on empty-cup placement and block stacking is attributed by the authors to constraints on the useful behaviors available in its generated candidates. The verifier reliability analysis shows IDM raises precision but lowers recall, and those labels describe complete action-tape replay rather than separate execution of selected prefixes; simulator failures may also reflect physics artifacts. Real-robot evaluation uses only ten trials per task across a limited number of layouts. These are scope and open questions rather than defects.

Sources