Skip to main content
Back to timeline
arXivSource publication:

Rolling-WAM spreads joint denoising across replanning cycles, reaching 98.1% on LIBERO and 93.3% on RoboTwin 2.0 with a 4.5x steady-state replanning speedup over standard joint WAMs

Synopsis

Rolling-WAM introduces a rolling world action model that keeps a sliding window of video-action chunks at staggered noise levels, fully denoising only the imminent action chunk each replanning cycle while partially refining farther-future chunks, thereby distributing joint denoising computation across control cycles; it achieves competitive manipulation success on LIBERO, RoboTwin 2.0, and real-world Unitree G1 humanoid tasks while delivering a 4.5x steady-state replanning speedup over standard joint WAMs.

AI-generated editorial illustration: Rolling-WAM: World Action Models with Rolling Imagination

Interpretation

The paper locates the inference bottleneck of World Action Models (WAMs) in the fact that each replanning cycle runs joint video-action denoising over the entire prediction horizon from pure noise, which raises inference latency and limits closed-loop responsiveness. Prior work either drops future-video generation at deployment (Fast-WAM) or caches future context; Rolling-WAM instead keeps future imagination and reorganizes how denoising is distributed over time. The claim is supported by problem formulation (Sec. III-A) and latency measurements: on a single NVIDIA A100 under the RoboTwin 2.0 setting, Joint-WAM takes 978 ms and Fast-WAM 548 ms per update, versus 215 ms for Rolling-WAM.

The core method is rolling video-action denoising: the window is partitioned into temporally aligned chunks with progressively higher noise toward the future; each cycle fully denoises only the nearest chunk while retained chunks continue denoising, and the window advances with new camera observations and appends a new Gaussian-noise chunk. The authors extend rolling diffusion from sequence generation to joint video-action modeling, letting action tokens attend to partially denoised visual futures across the whole prediction window while restricting direct action-to-action attention to the same chunk. The method is specified through rolling and initialization noise schedules (Eqs. 2 and 4), the Euler update (Eq. 5), and a flow-matching training objective (Eqs. 6-8); ablations show that allowing cross-chunk action attention drops success to 76.3% on selected RoboTwin tasks and 97.9% on LIBERO, below the 78.2% and 98.1% obtained with within-chunk action attention.

On simulation benchmarks, Rolling-WAM reaches competitive success with only two denoising steps per replanning cycle: 98.1% average on LIBERO and 93.3% average on RoboTwin 2.0. On LIBERO, 98.1% is within 0.4 percentage points of the joint leaders LingBot-VA and Joint-WAM (both 98.5%) and above Motus (97.7%) and Fast-WAM (97.6%); on RoboTwin 2.0, the 93.3% average leads LingBot-VA (92.2%), Fast-WAM (91.8%), and Joint-WAM (90.6%). LIBERO covers the Spatial, Object, Goal, and Long suites with 50 rollouts per task; RoboTwin 2.0 covers 50 bimanual tasks with 100 rollouts per setting in both Clean and Randomized conditions, with per-task results reported.

On a real Unitree G1 humanoid, a single policy trained across three tasks reaches 85.0% average success, above Joint-WAM at 78.3% and Fast-WAM at 75.0%. The paper moves rolling denoising from simulation to real humanoid manipulation and reports the highest success on Bead Pouring (70%), matching the strongest baselines on Doll Placement (85%) and Plate Stacking (100%). Each of the three tasks (Doll Placement, Plate Stacking, Bead Pouring) is evaluated over 20 trials, with 50 demonstrations per task recorded at 10 Hz and actions executed at 10 Hz; the authors also report qualitative observations that pauses are more apparent with Joint-WAM and sometimes interrupt task progress.

Perspective

The work targets robotic manipulation settings that require repeated replanning in dynamic environments: language-conditioned simulated manipulation (LIBERO, RoboTwin 2.0) and real-world Unitree G1 humanoid tasks. It suits deployers who want to keep explicit future visual prediction without re-denoising the entire prediction horizon from scratch each cycle, and the paper notes the approach is orthogonal to other WAM efficiency techniques and can be combined to further reduce inference latency. The default configuration uses 16 denoising steps per chunk, a window of 5 chunks, and 16 actions per chunk, so a steady-state cycle needs only two steps; latency measurements were taken on a single A100 without torch.compile, TensorRT, or custom CUDA kernels.

Open questions the paper itself raises include that design choices such as window and chunk sizes warrant further exploration across task settings with different dynamics and control frequencies, and that despite observation feedback, retained predictions may lag behind rapid scene changes and misguide action generation, particularly with long windows, with adaptive window management and asynchronous execution offered as directions. Ablations also show larger windows do not consistently improve performance (78.2% at window 5 versus 69.5% at window 8 on selected RoboTwin tasks), which the authors conjecture reflects that more distant visual predictions are less constrained by the current observation. Alternative training noise mixtures also fail to improve both benchmarks at once: Random raises selected RoboTwin success to 78.5% but lowers LIBERO to 97.3%.

Sources