Skip to main content
Back to timeline
arXivSource publication:

Tsinghua's Leap Lab finds in a matched comparison that world action models generalize from a single inference-time forward pass, not from denoising a clean future

Synopsis

Under matched backbone, data, and training budget, the study compares explicit and latent world action models and finds that latent models match explicit ones in distribution (96.85 vs 97.75) yet fall behind on all three generalization axes—environmental perturbation, data efficiency, and task generalization—while leaving future video tokens at pure noise and running a single forward pass recovers most of the gap, motivating Simple-WAM, which leads explicit models in generalization at latent-level inference speed.

AI-generated editorial illustration: What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling

Interpretation

The work proposes a three-axis generalization protocol that decomposes world action model generalization into environmental perturbation, data efficiency, and task generalization, and compares the explicit and latent inference paradigms under one matched setting. Earlier debate over whether the future must be generated at inference drew on each system's own backbone, data, and budget, and was largely collected in distribution; this work places the three dimensions side by side in a controlled setting where only whether the action expert reads future-token representations varies. The comparison uses a Wan2.2-5B video DiT with a B-scale action expert, with the two paradigms differing only in a structured attention mask; the perturbation axis uses the seven LIBERO-Plus factors, the data-efficiency axis reduces demonstrations per task from 40–50 to 10, and the task-generalization axis runs four-fold cross-validation over the four LIBERO suites.

Latent world action models match explicit ones in distribution but degrade consistently across all three generalization axes, indicating that inference-time future modeling is a requirement rather than only a training objective. Latent work had reported near-explicit success rates at a fraction of the cost; this study shows that in-distribution parity does not imply comparable generalization and quantifies the gap on three axes. Table results show 96.85 vs 97.75 in distribution, 53.75 vs 67.72 under perturbation, and 88.50 vs 96.95 at 10-shot; the largest separation is task generalization, at 2.10 vs 6.25 without video and 5.90 vs 69.90 with action-free video.

Almost all of the generalization gain comes from the first denoising step: leaving future video tokens at pure Gaussian noise and running one forward pass recovers most of the gap. This challenges the intuition that a more clearly denoised future yields a stronger policy and explains why the explicit paradigm's cost is not necessary—the benefit comes from preparing the future, not generating it. Holding the explicit-trained model fixed and varying only the number of video-expert runs K, K=1 with pure-noise future tokens still beats the latent paradigm by 26.0, 8.9, 8.0, and 67.7 points across four settings; the nine subsequent passes that carry the entire denoising cost add at most 0.6, 0.6, 0.6, and 0.6 points.

The resulting Simple-WAM conditions on a fully noised future in one inference pass and uses a single scalar to shift training-time flow-time sampling toward that point, obtaining explicit-paradigm generalization at latent-paradigm speed. Relative to the explicit paradigm it drops iterative denoising; relative to the latent one it keeps future-token conditioning; it adds no module or loss term, changing only inference forward passes and the training-time flow-time distribution. It averages 79.5 over the seven LIBERO-Plus factors, against 67.7 for the explicit paradigm and 53.8 for the latent one; at 10-shot it reaches 97.2 on LIBERO and 37.2 on RoboTwin; with action-free video it reaches 73.6 on LIBERO and 44.5 on RoboTwin; per-chunk latency is 112 ms versus 289 ms explicit and 98 ms latent. Ablations replacing noise tokens with learnable queries or zeros drop the three-axis average from 94.2 to 52.7 and 54.9.

Perspective

The result targets research and engineering settings built on video-generation-initialized models doing language-conditioned manipulation with action chunks: on LIBERO, LIBERO-Plus, RoboTwin 2.0 simulation, and an AgileX Aloha dual-arm platform, Simple-WAM lets a policy retain explicit-paradigm generalization under environmental perturbation, reduced demonstrations, and held-out tasks while bringing per-chunk latency close to the latent paradigm. For teams needing high-frequency closed-loop control that also want to acquire new skills from action-free video, it offers a route that adds no module or loss term; for evaluators, the three-axis protocol can serve as a shared framework for comparing different world action model implementations.

Conclusions rest on a B-scale backbone without embodied pretraining, and the authors explicitly leave open whether they hold at larger scale or with such pretraining. Within the three-axis protocol, task generalization without video sits near the floor for both paradigms (6.25 and 2.10), so that axis discriminates mainly in the with-action-free-video setting. Real-world evaluation covers four tasks with 30 trials each and scores the fraction of sub-goals completed rather than binary success, with per-task standard deviations in the appendix. In addition, this reading covers the paper body and appendix tables; the specific curves of Figures 1 through 6 are not expanded in the text, so checking the point-by-point K sweep or the real-world task illustrations still requires the original figures.

Sources