PointZero: 3D Point Track Completion as a Pre-training Objective for Transferable 3D Dynamics
Synopsis
The work proposes 3D point track completion as a pre-training objective: given a single RGB-D observation and sparse partial 3D tracks, it predicts future 3D tracks of all observed points, thereby learning a transferable 3D dynamics prior without robot action labels; the authors contribute a 2.9 million synthetic frame dataset spanning deformable, articulated, and rigid objects and train a transformer, PointZero, which outperforms prior methods on the same data, outperforms baselines on the recent PGND 3D dynamics benchmark after fine-tuning, and outperforms or matches baselines on 6 of 7 simulated and real-world robot manipulation tasks, while also training PointZero from scratch to isolate the benefits of the architecture from those of the pre-training objective and dataset, and releasin
Figure 1 : PointZero proposes 3D point track completion as a pre-training objective for learning rich 3D dynamics priors. Given one RGB-D image and one or a few point trajectories (shown in orange), PointZero predicts the future 3D tracks of all observed scene points. PointZero is a robot-free pre-training objective for instilling 3D dynamics understanding in world models. Post-trained for dynamics modeling Zhang et al. (2025) and robot manipulation Hung et al. (2026b) , it outperforms application-specific baselines.
arXivInterpretation
It proposes and validates 3D point track completion as a pre-training objective that requires no robot action labels and yields a rich 3D dynamics prior. Existing methods typically require robot action labels to learn action-conditioned 3D dynamics, which excludes web video data from the training pool; this objective removes the dependence on action labels during pre-training. The abstract states the objective and its conclusion and notes transfer to two downstream applications; experimental details are not expanded in the provided text.
It contributes a diverse dataset of 2.9 million synthetic frames spanning deformable, articulated, and rigid objects and uses it to train PointZero. Relative to prior work, the dataset offers a broader source of diversity in object types and scale to support the proposed pre-training objective. The text explicitly states '2.9 million synthetic frames' and the three object categories, which is the authors' own description of data scale and coverage.
The transformer PointZero outperforms prior methods on the same data and, after fine-tuning, outperforms baselines on the recent PGND 3D dynamics benchmark. Comparison on the same data indicates a gain from the architecture itself; after fine-tuning to condition on end-effector pose, it achieves an advantage on a recent benchmark. The text states 'outperforms prior methods on the same data' and 'outperforms the baselines on the recent PGND 3D dynamics benchmark', but gives no specific metric values.
When fine-tuned to predict robot actions and 3D tracks, it outperforms or matches baselines on 6 of 7 simulated and real-world robot manipulation tasks; training from scratch further isolates the contributions of architecture versus pre-training objective and dataset. It grounds the pre-training objective in imitation learning as a concrete downstream task and adds a from-scratch comparison to separate the roles of architecture and pre-training/data. The text reports a count-based result of '6/7 simulated and real-world robot manipulation tasks' and a from-scratch control setting; per-task details are not given.
Perspective
The result targets 3D dynamics pre-training and downstream fine-tuning settings that take a single RGB-D observation and sparse partial 3D tracks as input, and it is relevant to researchers and engineering teams who want to draw on broader data sources without robot action labels; the authors release the dataset, checkpoints, and full training recipe, which supports reproduction and extension on the same data and benchmarks.
The provided text is the abstract and metadata, without figures, specific metric values, task lists, or ablation details, so the magnitude of gains per downstream task and the quantitative difference in the from-scratch comparison cannot be judged here; these are questions a reader should continue to watch when reading the original paper.
