Aligning to Frozen World-Model Features Lifts a 0.8B Robot Policy to 97.9% on LIBERO and +2.3 Points on RoboCasa-GR1 at Zero Inference Cost
Synopsis
The work introduces a method that distills world-model representations into compact vision-language-action (VLA) policies: ordinary VLA training gains one cosine alignment term against cached features from a frozen world model, the teacher runs only once offline and is never loaded during training, the projector is discarded afterward, and the deployed policy is identical to the undistilled baseline; the 0.8B student reaches 97.9% average success across the four LIBERO suites (2.6 points above the identical undistilled student), improves RoboCasa-GR1 humanoid manipulation from 48.2% to 50.5%, and gains on both single-arm and bimanual real platforms, with the gain surviving changes of student scale, backbone, alignment layer, and teacher.
Interpretation
It proposes a zero-inference-cost representation distillation recipe: on top of the standard VLA action objective, a single cosine alignment term matches the student's per-view pooled image-token features to those of a frozen world model, with teacher features cached offline, no teacher weights loaded during training, the alignment projector discarded afterward, and a deployed graph identical to the undistilled baseline. Previously, using a world model as a policy required rolling the future online at high cost (the text reports DreamZero needing 3s per inference on an H100 and 45.9GB, versus this student policy at 32ms and 1.86GB on a consumer RTX 5090); this work separates what a world model knows from generating the future, inheriting only the internal features without the generative overhead. The method is validated on two simulation benchmarks (LIBERO, RoboCasa-GR1) and two real platforms (AgileX Nero single-arm, TRIP-Bag bimanual); the authors report the teacher cache takes about an hour on four GPUs for LIBERO and state that no teacher action targets are used because teacher and student do not share an action parameterization.
The distilled 0.8B student beats its same-scale undistilled control and approaches larger models: 97.9%±0.5 average on the four LIBERO suites versus 95.3%±0.7 for the identical undistilled student, a 2.6-point gain, and 48.2%±2.1 to 50.5%±2.3 on RoboCasa-GR1, a 2.3-point gain. The gain moves the 0.8B policy past QwenFAST, QwenPI, QwenOFT and both Isaac-GR00T releases on RoboCasa-GR1 and to within points of the 4B QwenGR00T (54.8%±2.0), roughly a five-fold parameter reduction; the authors note the same-scale comparison is the sharpest, with every other policy below 0.8B in the table sitting between 78.7% and 88.8%. Each checkpoint is evaluated four times, crossing two evaluation seeds with two GPUs (A100-80GB and RTX 5090), reporting the mean and standard deviation; the authors note LIBERO spread is about one to one and a half points while RoboCasa-GR1 can move by several points when the GPU changes, and that RoboCasa-GR1 ablations were scored only once on A100 due to limited compute.
The same pattern holds on real robots and the gain varies with task difficulty: on single-arm fruit pick-and-place the distilled 0.8B policy reaches 93.3%, matching the 4B control and 10 trials above the undistilled control; on the egg task the undistilled control falls to 46.7% while distillation recovers to 60.0%; on bimanual fruit handover the control is 40.0%, distillation 46.7%, and the 4B control 53.3%. The authors carry the fruit task from a single arm to the TRIP-Bag bimanual platform and rebuild it as a two-armed handover, changing embodiment, number of arms, cameras, and controller latencies at once, to ask whether the gain is a property of the method or of one particular setup; failures are described as placement errors rather than recognition errors, with an added grasp-slippage mode on eggs. Each policy, task, and platform is scored over 100 trials, with success requiring every object to end up inside the target receptacle; policies are fine-tuned per platform on roughly 30 minutes of teleoperated demonstration for single-arm tasks and about 1 hour for the bimanual one.
Ablations show the recipe is robust to its own design choices: it works across three students (0.8B, 1B with an InternVL backbone, 4B with a Qwen3-VL backbone; 50.3→52.8, 50.9→53.3, 56.8→58.4 respectively), every alignment layer from L8 to the final layer trains stably and lands within four points of the best, and three teachers (Cosmos3-Nano, Fast-WAM, V-JEPA2-AC) all lift the same student from 95.3% to between 96.5% and 97.9%. The authors attribute the transferred signal to world-model representations in general rather than to a particular teacher or a fragile alignment between two networks; the alignment-layer sweep runs only a quarter of the full schedule on RoboCasa-GR1, and the authors explicitly read only the ordering, not the absolute numbers. The teacher ablation is run on LIBERO to reuse existing pretrained teachers; the authors note the spread between the best and second-best layers is smaller than a quarter-length schedule can resolve, and keep the final layer because it is the representation the action head already consumes and the cheapest to implement.
Perspective
The result targets researchers and engineering teams who want to deploy compact robot policies on consumer hardware: the recipe applies when a frozen world model is available and teacher features can be precomputed offline, teacher and student may live in incompatible environments, and one cache can serve multiple students. The settings the authors demonstrate are the LIBERO and RoboCasa-GR1 simulations plus pick-and-place and handover tasks on the AgileX Nero single-arm and TRIP-Bag bimanual real platforms; the deployment-side benefit is explicitly limited to inheriting the representation without paying the generative cost, meaning latency and memory match the undistilled baseline.
The authors describe the gains as modest but consistent, and the method still trails several larger or data-engineering-heavy baselines on LIBERO and RoboCasa-GR1, so how much headroom this representational prior has on stronger baselines remains open. On evaluation, the authors note neither simulation benchmark is deterministic: policies sample actions from a flow-matching head, the renderer is not bit-identical across GPU models, and the same checkpoint can move by several points when the GPU changes; RoboCasa-GR1 ablations were scored only once on A100 due to limited compute, and the alignment-layer sweep ran only a quarter of the full schedule, all of which affect how small differences should be read. Real-robot failures are described as placement errors and egg grasp slippage, which the authors see as what a representational prior helps with least, leaving the boundary between this method's benefit and contact-level control to be clarified. In addition, the teacher cache takes about an hour on four GPUs for LIBERO, and the effect of the teacher read-out layer and pooling choice is not further swept in this text.
