Async replanning stops restarting from independent noise: a shared noise trajectory follows real execution displacement, reaching 96.7% on real-robot cloth folding
Lead
Generative robot policies no longer restart from independent Gaussian noise at every asynchronous replanning step; adjacent action chunks now share an initial noise trajectory aligned to actual execution displacement and correlated along action time, reaching 96.7% success on real-robot cloth folding and 90.0% on object storage.
Story
In continuous asynchronous replanning, the stochastic starting point of consecutive inferences is now explicitly related, so a new action chunk continues from an execution-consistent generative state instead of restarting from independent noise. Earlier approaches concentrated constraints on the committed or imminently executed action prefix, while the remaining suffix beyond that prefix was still determined by a newly sampled stochastic source, leaving cross-inference stochastic-state dependency unmodeled. Flow inversion of a standard policy shows pronounced temporal locality in source latents within the same episode, with correlation decaying as window separation grows and near-zero correlation across episodes, and the correlation peak moving with actual execution displacement rather than staying at a fixed index.
The method introduces structured stochasticity at two levels: an inter-chunk shared autoregressive process that persists across policy calls and is re-indexed by actual execution displacement, and an intra-chunk autoregressive process that adds correlation along action time. Standard flow matching samples Gaussian sources independently at every action-time position, so different stochastic initializations can drive temporally overlapping inferences toward different, individually high-probability local modes. On the Kinetix mjc_swimmer task, adding inter-chunk noise to ttRTC raises success from 46.60 to 76.89 and adding intra-chunk noise further to 78.85, with the complete model at 79.26, so inter-chunk correlation supplies the dominant gain.
The aligned stochastic history is combined with committed action context and injected as a policy condition, so new chunks continue from execution-consistent generative states. Correlated noise alone does not tell the policy which historical stochastic state corresponds to the current window, so explicit alignment and conditioning are required. During training, delay and displacement are sampled from the evaluation support, and the entire history is dropped with some probability, in which case the current source is resampled independently and both the aligned history and its validity mask are zeroed to cover episode starts and invalid history.
What to watch
Next steps can test whether this source structure holds across more base policies and more delay ranges, especially by unifying the low-delay and high-delay protocols into a continuous delay sweep of a single model. Teams doing asynchronous deployment can attach the shared autoregressive source and aligned-history conditioning to their own flow-matching policies while keeping the original action-chunk execution logic, and observe behavior under their own delay distribution.
The low-delay and high-delay experiments each use matched training or fine-tuning protocols, so the two sets of numbers should not be read as a continuous delay curve of a single model. The two real-robot tasks use only 20 and 30 trials respectively, so the stability of the success-rate differences still needs more trials. The source structure changes the joint covariance of the full source, and the model must be trained to adapt to it, a cost worth watching when transferring to new policies.
