InternW0-Δ pretrains a world action model on 20K+ hours of heterogeneous data, reaching 92.8% on LIBERO-Plus and 71.9% on RoboTwin 2.0 Clean2Random
Synopsis
InternW0-Δ couples a pretrained video expert and an action expert through a directed Mixture-of-Transformers, uses a frozen VLM for scene semantics, learns future-relevant scene changes via Causal Imprint from training-only future supervision, and distills 4D geometric and motion priors from a Track4World teacher at training time only; pretrained on a corpus of over 20K hours unifying robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data under a canonical state-action representation, it reaches 92.8% on LIBERO-Plus, 71.9% on RoboTwin 2.0 Clean2Random, an overall score of 66.0 on EBench, and an average success rate of 23.91% on RoboDojo, with deployment on four real-robot platforms (two gripper-based, two dexterous-hand).
Interpretation
It introduces a directed World-Action Mixture-of-Transformers in which a pretrained video expert and an action expert keep modality-specific parameters and exchange information only through masked joint attention, a frozen VLM supplies scene semantics to the action expert, and training-time future observations serve only as supervision rather than entering the forward path of action prediction. Where earlier WAMs require test-time future-video generation or future-frame inputs, this design confines future information to training supervision so that inference predicts actions directly without sampling future video. The paper describes the attention mask, the training objectives (video and action flow matching, two Causal Imprint losses, and a distillation MSE), and the inference procedure; ablations on LIBERO-Plus show success rising from a 49.59% baseline to 76.45% as sparse memory, VLM, Causal Imprint, and representation alignment are added.
Causal Imprint uses learnable tokens to encode future-relevant scene changes from anchor, recent, and current observations, supervised by adjacent clean-latent differences and by representation alignment with future video features. It converts future prediction into an imprint on representations available at inference time, giving the action expert predictive representations without access to realized future frames. In ablations, adding Causal Imprint alone raises LIBERO-Plus success from 69.08% to 70.80%, and adding representation alignment raises it further to 76.45%; the paper states future observations are used only to construct supervision targets.
It builds and unifies a heterogeneous corpus of over 20K hours: robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data are mapped into an 80-dimensional canonical state-action space and filtered for signal consistency, static boundary trimming, visual quality, action magnitude, instruction correctness, and video-instruction consistency. The paper describes this as the largest open-source corpus of its kind and reports before-and-after filtering hours and episode counts together with Ego2Robot conversion coverage statistics. After filtering, robot data totals 11,302.20 hours across 1,247,656 episodes; EgoDex and EgoVerse retain 4,061.35 hours; Hy-UMI-10K retains 2,075.42 hours; Ego2Robot conversion covers 1,101.55 hours and yields 5,633.77 robot-hours.
It validates that the same pretrained checkpoint can be post-trained to different embodiments and control interfaces across four simulation benchmarks and four real-robot platforms, and reports training and deployment infrastructure optimizations. One checkpoint is adapted separately to gripper and dexterous-hand interfaces, with reported round-trip latency under asynchronous execution and training-throughput speedups. LIBERO-Plus 92.8%, RoboTwin 2.0 Clean2Clean 90.0% and Clean2Random 71.9%, EBench overall score 66.0, RoboDojo average success rate 23.91%; the dexterous-hand deployment averages 152.8 ms controller-observed round-trip latency on a single RTX 5090; end-to-end training throughput speedups are reported on LIBERO and RoboTwin.
Perspective
The work targets generalist robot manipulation and applies to policy learning that must transfer across embodiments and task distributions: the pretraining corpus spans robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data, and post-training adapts the same checkpoint to LIBERO-Plus, RoboTwin 2.0, EBench, RoboDojo, and four real-robot platforms (AC-One, Arx5, Franka+XHand, TianJi Marvin+Wuji Hand). For teams wishing to reuse its data recipe, canonical state-action representation, or deployment optimizations, the paper commits to releasing training code, model weights, infrastructure, data-processing pipeline, and processed data where licenses permit, with versioned indices recording filtering decisions.
The contribution of egocentric human demonstrations is not yet systematically studied; the paper states it has not examined different egocentric data sources, conversion strategies, scaling ratios, or supervision forms. Agent-assisted control is only a preliminary exploration, without systematic study of invocation policies, long-horizon planning, or hierarchical decision making. Data-source ablation results are specific to the current data scale and to evaluation settings dominated by gripper-based manipulation, and the authors caution against concluding that egocentric data does not improve learned representations. Some real-robot tasks use limited evaluation trials, so stability across tasks and embodiments needs further verification. In addition, this reading is of the full text, but equations and tables appear as placeholders in the text, so exact numerical details should be checked against the original tables.
