Foresight Without Trajectories: Guiding 3D Diffusion Policies with a Latent Movement Trend
Synopsis
The work introduces Movement Trend Guidance: a compact latent of interaction evolution is learned from a short history of point clouds and robot states, supervised during training only by sparse future gripper states, and at inference retained alone as future-oriented conditioning, consistently improving 3D diffusion policies on RoboTwin2.0, LIBERO-40, DexArt and five real-robot tasks while adding only 3.52% more parameters than DP3.
Interpretation
It formulates a notion of foresight that needs no explicit plan: the policy encodes not only what motion is feasible now but also where the interaction is heading. Unlike waypoint, keyframe or trajectory prediction, no intermediate control target is produced; future information only shapes the representation, preserving freedom to correct local motion from current geometry. Supported by design and ablation: explicit future-point conditioning scores below the full model on LIBERO-40, indicating that shaping a latent with future prediction beats conditioning directly on future points.
It offers a lightweight realization: future states supervise a compact latent that enters the standard global-conditioning pathway, with additional gated FiLM modulation confined to the UNet bottleneck. Relative to DP3 it adds 9.23M parameters (+3.52%, 271.67M total) and raises mean inference latency only from 50.28 ms to 50.90 ms (+1.23%), while preserving the original dense-action and receding-horizon formulation. Parameter and latency figures are reported in the text; injection comparisons show bottleneck-only gated FiLM at 62.8%, all-block gated FiLM at 55.8% and cross-attention at 52.9%.
It reports consistent gains across several simulation benchmarks and real-robot tasks. 50-task RoboTwin2.0 mixed training 62.8% vs 56.1%, LIBERO-40 71.93% vs 37.08%, DexArt average 59.25% vs 52.0%, and five SO101 real-robot tasks 72.0% vs 49.0%. Each benchmark follows its official protocol and demonstration budget (e.g. RoboTwin2.0 with 50 demonstrations per task and 100 evaluation episodes; LIBERO-40 with 50 demonstrations per task and three rollout seeds of 50 episodes each; real tasks with 50 demonstrations and 20 rollouts).
Ablations indicate the gain comes mainly from future-supervised representation learning rather than added representational capacity alone. A parameter-matched latent variant without the future loss stays only about one point above DP3, while adding future supervision yields a further large gain; movement-trend conditioning also lifts ACT, a non-diffusion backbone, from 24.00% to 51.67% mean success on six tasks. Ablations run under the fixed 50-task mixed-training protocol with the same 64-dimensional task embedding; the ACT result is explicitly limited by the authors to the evaluated six-task setting.
Perspective
The method targets 3D diffusion policies that observe point clouds and robot states and execute dense actions in a receding-horizon loop, suited to multi-stage manipulation such as bimanual handover, stacking and pick-and-transport; the authors report the largest gains on contact-rich and multi-stage tasks and provide reproducible code. For researchers and practitioners who want to add foresight without introducing an explicit planning interface, this design offers a lightweight path that can be layered onto existing policies.
Real-robot validation is limited to the SO101 platform and five tasks, and the authors explicitly list broader validation across objects, scenes and perturbations as future work; the failure analysis notes that when the arm occludes the object and the point cloud is incomplete, both the trend estimate and the baseline are affected, so gains are smaller on precise contact and placement tasks. DexArt also aggregates the best five evaluation checkpoints, which differs from the fixed epoch-1000 LIBERO reporting, so cross-benchmark comparisons should note this difference. A reader seeing only the abstract should return to the full text for per-task details and complete ablation numbers.
