WorldLine pretrains on 10,000 hours of action-free robot video and grounds 2,000 hours of action trajectories across ten embodiments, lifting robot-mask IoU by 0.1626 on failed trajectories and predicting trajectory success at 74% mean accuracy
Synopsis
WorldLine introduces an action-driven visual simulator that decouples transferable robot-object dynamics learning from heterogeneous action grounding: it pretrains on more than 10,000 hours of action-free robot videos, grounds them with over 2,000 hours of action trajectories across more than ten embodiments through image-space action maps, adds multi-view, failure-enriched and relational-regularized training, and distills a robot-focused few-step causal rollout; across held-out and out-of-domain settings it maintains visual quality and robot-motion agreement, improves robot-mask IoU by 0.1626 over the strongest baseline on failed trajectories, predicts trajectory success with 74% mean accuracy on RoboTwin and AgiBot, and improves task success by up to 21.
Interpretation
WorldLine splits dynamics learning from action grounding: Stage I adapts a pretrained video model to more than 10,000 hours of action-free robot videos drawn from six collections and spanning more than ten embodiments and over 3,500 manipulation tasks, and Stage II learns how controls drive future motion from over 2,000 hours of action trajectories across more than ten embodiments. Existing action-conditioned simulators largely learn visual dynamics and action semantics jointly from the same action-labeled trajectories, coupling simulation coverage to scarce action supervision; WorldLine lets action-free videos and embodiment-specific action trajectories provide complementary supervision. The paper reports data scale and sources (AgiBotWorld, RoboCOIN, the RoboMIND series, Galaxea, among others), and a compute-matched ablation shows that removing action-free data degrades both visual quality and planning, which the authors attribute to complementary interaction coverage rather than extra optimization.
WorldLine uses image-space action maps as a shared cross-embodiment control interface: end-effector pose and gripper state are projected into each camera view to form nine-channel action maps that are spatially and temporally aligned with video latents and added to the DiT. Native action vectors are tied to embodiment-specific kinematics and coordinate systems and are hard to share, while latent actions need additional grounding; an image-space representation preserves action geometry so different control spaces can share one conditioning interface. In ablations, replacing image-space action maps with FiLM-injected native action vectors consistently worsens robot IoU, LPIPS and planning success under both policies; under matched RoboTwin training and a 32-rollout budget, WorldLine outperforms the projection-based GE-Sim-V2 and Masked Visual Actions.
WorldLine improves interaction-sensitive prediction through synchronized multi-view observations, roughly 200 hours of failure trajectories and relational regularization against a frozen V-JEPA2, and uses robot-focused four-step causal distillation to cut sampling from 35 steps to four. A single camera can miss gripper-object contact and success-heavy data underrepresent failure outcomes; relational regularization explicitly constrains temporal and cross-view relations, while distillation supports online updates and low-latency rollout while preserving action-critical motion. Ablations show head-view-only training reduces robot-mask overlap, success-only trajectories barely change one policy but lower LingBot-VLA success by 10.8 points, and removing relational regularization mainly affects visual quality and geometric consistency; distillation cuts 129-frame generation from 90 to 21 seconds and per-frame latency from 0.698 to 0.163 seconds.
Across action-conditioned generation, policy evaluation and embodied planning, WorldLine maintains visual quality and robot-motion agreement in held-out and out-of-domain settings, improves robot-mask IoU by 0.1626 over the strongest baseline on failed trajectories, predicts trajectory success with 74% mean accuracy on RoboTwin and AgiBot, and improves task success by up to 21.4 percentage points without RoboTwin training or adaptation. These results move a visual simulator from looking plausible toward being usable for ranking candidate actions and evaluating policies, and they hold across unseen environments and embodiments. Evaluation uses a scene-disjoint AgiBotWorld split and the fully out-of-domain DROID, which is excluded from training, adaptation and checkpoint selection; planning experiments fix policy checkpoints, share candidate trajectories, and use Qwen3VL-8B with hidden candidate identities, reporting standard deviation over five independently sampled candidate banks.
Perspective
The work targets robot-learning settings that need action outcomes predicted before physical execution: policy developers can use it to evaluate candidate behaviors, perform best-of-n trajectory selection, and obtain cross-embodiment rollouts without touching target-platform data. It applies where camera calibration and robot kinematics are available and where an initial observation plus an action sequence can be supplied; the paper explicitly excludes RoboTwin and DROID from training, adaptation and checkpoint selection, so its out-of-domain claims concern unseen environments, camera configurations and embodiments. The four-step causal variant targets low-latency online updates and long-horizon rollout, and the authors propose using it as an RL environment and as a test-time scaling tool.
Several points remain worth watching: planning gains are bounded by the diversity of trajectories proposed by the underlying policy and by the quality of the multimodal selector, and the authors note that Qwen3VL-8B is imperfect so candidates do not guarantee monotonic gains; few-step distillation does not impose uniform quality loss but changes which task-relevant cues survive in the rollout, being more useful when outcomes are visually distinguishable and less reliable when selection depends on finer differences, and whether spending saved runtime on more candidates can offset ranking errors remains open; failure cases show residual instability on out-of-domain DROID, including severe scene distortion and blur, action misalignment and inconsistent interaction outcomes; rare manipulation patterns, deformable objects, unusual camera configurations and substantially longer action horizons may need broader and more targeted supervision. In addition, this evidence bundle is a full-text parse, but figures and some appendix details appear as references rather than itemized values, so individual numbers should still be checked against the original.
