Skip to main content
Back to timeline
arXivSource publication:

AnyStep-WAM distills frozen-teacher trajectories and schedules budgets by risk and benefit, cutting denoising steps by roughly half to 85% across three world-action models while holding success rates

Synopsis

The work introduces AnyStep World Action Model, a framework that performs budget-aligned flow-map distillation from frozen-teacher trajectory intervals and trains a lightweight risk-benefit scheduler to predict teacher-trajectory difficulty and budget-specific student fidelity from a single one-step preview, selecting the smallest denoising budget that meets a fidelity requirement; on RoboTwin 2.0 it reduces average denoising steps by 60.2%, 49.8%, and 85.28% on Motus, FastWAM, and LingBotVA while keeping average success within 0.24 percentage points of full-budget baselines, raises one-step success by 7.07, 12.08, and 8.94 percentage points respectively, and achieves 1.67-6.14x per-call speedups on six real-world manipulation tasks.

AI-generated editorial illustration: AnyStep-WAM: Budget-Aligned Distillation and Adaptive Inference for World Action Models

Interpretation

It proposes budget-aligned integral flow-map distillation: finite-interval displacements along a frozen teacher trajectory serve as supervision targets, with training intervals sampled to match candidate inference schedules, so one model can generate from one-step prediction to multi-step refinement. Earlier few-step and any-step methods such as MeanFlow and AnyFlow rely on derivatives or finite differences of the evolving student; this work instead uses integral targets tied to frozen-teacher endpoints and aligns training intervals with inference schedules. The paper reports that finite-difference targets show wider sample-level distributions and frequent gradient spikes, while frozen-teacher targets stay concentrated with more stable gradients; in ablations, removing interval supervision drops two-step success from 86.21% to 78.94% and adaptive success from 87.96% to 78.09%, raising the average budget from 3.98 to 5.27 steps.

It proposes risk-benefit adaptive inference: teacher-trajectory velocity variation serves as a difficulty proxy and offline-evaluated student-teacher agreement at each budget serves as a fidelity proxy, with a lightweight scheduler predicting both from a single one-step preview and choosing the smallest budget meeting a difficulty-dependent fidelity threshold. Existing adaptive methods mostly vary computation with the input but do not separately model teacher denoising difficulty and the student's fidelity at a given budget; this work frames budget selection as a risk-benefit decision. On RoboTwin, the scheduler improves average success over a training-free percentile gating baseline by 0.70, 1.43, and 1.88 percentage points on Motus, FastWAM, and LingBotVA while cutting average denoising steps by a further 21.8%, 24.5%, and 23.5%; shifting all thresholds by five percentile points keeps success within 0.36 percentage points of Full.

It proposes single-preview execution: the one-step preview both informs budget selection and is re-noised into the selected schedule, so no extra denoising pass, online teacher evaluation, or candidate rollout is needed. Compared with approaches that evaluate multiple candidate budgets online, this design reduces scheduling to lightweight prediction and avoids wasting the preview computation. The paper reports budget-selection overhead of only 2.881, 2.4996, and 2.9772 ms per prediction; on a single A100, average per-call latency drops from 1.932 to 0.859 s for Motus, 0.496 to 0.295 s for FastWAM, and 9.732 to 3.247 s for LingBotVA.

It validates across three WAM backbones and a real robot: distillation improves one-step success, and adaptive inference substantially cuts steps and latency while maintaining or improving average success. The paper states that to its knowledge prior work has not specifically addressed adaptive budget selection for flow-map-trained WAMs, and the validation covers both joint video-action and action-only inference forms. On RoboTwin, one-step success improves by 7.07, 12.08, and 8.94 percentage points over Base and exceeds Flash-WAM by 3.19, 7.58, and 1.30 percentage points; on the real robot, average success rises to 69.17%, 71.67%, and 70.00% while mean denoising steps fall by 59.4%, 63.9%, and 85.84%, with per-call latency dropping from 1.868 to 1.117 s, 0.280 to 0.122 s, and 4.654 to 0.758 s.

Perspective

The results target world-action models with flow-matching denoising branches, in settings such as robotic manipulation where action-chunk precision demands vary by stage; the paper validates on Motus, FastWAM, and LingBotVA, covering 50 bimanual RoboTwin 2.0 tasks and six real-world tasks (Bowl Pouring, Clean Table, Put Block, Put Cup, Stack Blocks, Stack Bowls) on a Unitree G1D dual-arm robot with Dex1-1 grippers. Deployers seeking lower inference latency can adopt its candidate budget set and shared LoRA adaptation; researchers on adaptive computation can transfer the separated risk and fidelity modeling to other iterative generative policies.

Both risk and fidelity are explicitly described as proxies rather than direct measures of task-failure probability or terminal action error, so how scheduler predictions relate to real task outcomes remains an open question. Real-world experiments use about 20 trials per task and 50 demonstrations per task, a limited sample; the paper also notes that a clean-only-trained scheduler comes close to full training, but the scope of that generalization needs more environments to confirm. The loaded text is the full paper, yet some tables and figures (such as Figure 3, Figure 5, Figure 6, and several appendix tables) appear by reference without their values expanded, so per-task detail still requires the original figures and tables.

Sources