TrackEverything pushes dense 3D point tracking past 1000 frames by de-duplicating scene representations, beating open-source all-frame dense trackers by over 20% APD on short clips
Synopsis
TrackEverything represents videos as persistent 3D scene tracks in world coordinates and, through voxelization-based de-duplication at sliding-window boundaries, an endpoint-then-trajectory decomposition, and 3D WAFT feature sampling, becomes the first 3D tracker to follow all visible points across videos exceeding 1000 frames within 40 GB of GPU memory, outperforming all open-source all-frame dense 3D trackers by more than 20% APD on TAPVid-3D short clips while remaining competitive with state-of-the-art sparse trackers on long sequences.
Interpretation
It introduces TrackEverything, a 3D point tracker that represents videos as persistent 3D scene tracks in world coordinates and, to the authors' knowledge, is the first 3D tracker able to track all visible points across videos exceeding 1000 frames within 40 GB of GPU memory. Sparse trackers previously handled long horizons but only for a user-specified set of query points, while dense trackers either tracked only points visible in the first frame or were confined to 48-64 frame clips; this work decouples model complexity from video duration so it scales with unique physical scene geometry instead. Evaluated on the official minival split of TAPVid-3D (50 clips each from ADT, DriveTrack, and PStudio), and reported on full-length 1000+ frame PointOdyssey videos with APD-P 31.8, APD-M 74.5, and 91.2% static/dynamic accuracy; no baseline can run in that setting, so no quantitative comparison is reported there.
It proposes a voxelization-based de-duplication mechanism at sliding-window boundaries: when points from different frames converge onto the same physical surface and occupy the same 3D coordinates, they are merged into a single canonical track, preventing repeated observations of the same surface from accumulating. Prior 3D tracking methods retained a full grid of points for every video frame, so memory scaled linearly with duration; this mechanism makes the representation scale with unique physical geometry rather than frame count. The voxel-size ablation shows that without voxelization the model quickly runs out of memory, that accuracy stays stable from a fine voxel size up to a certain value (31 APD-P), and that coarser voxels increase speed and reduce peak memory; the default voxel size retains peak accuracy while halving latency and memory.
It decomposes tracking into an endpoint refiner that predicts each point's destination at the window boundary and classifies it as static or dynamic, plus a lightweight trajectory refiner that decodes dense trajectories exclusively for dynamic points, and it proposes 3D WAFT, which replaces memory-prohibitive 4D correlation volumes with efficient feature sampling in the scene cloud. Because static points remain stationary in world coordinates, this factorization concentrates dense decoding on the small moving subset of the scene; 3D WAFT extends warp-aligned feature transforms to 3D point clouds, using template matching at projected locations instead of 4D correlations. Ablations show that forcing all points through the dynamic trajectory decoder leaves tracking accuracy identical (same APD-P) while increasing latency and peak GPU memory, that removing iterative trajectory refinement reduces APD-P from one value to another, and that removing 3D WAFT feature sampling drops APD-P as well; the classifier reliably segments moving objects on Dynamic Replica and PointOdyssey.
Complexity analysis shows TrackEverything keeps low latency and memory along both axes: it stays under 30 GB even at 900 frames, whereas Any4D exceeds memory at 96 frames, SpatialTracker-v2 and DeltaV2-dense at roughly 200 frames, DeltaV2-sparse at roughly 500 frames, and CoTracker3 at roughly 600 frames. Existing dense trackers did not scale to long videos in accuracy, latency, or memory, and sparse 3D trackers showed rapid growth in both latency and memory as either axis increased. Measured end-to-end inference time and peak GPU memory on a single L40S-46G using the PointOdyssey validation set, varying frame count at fixed query count and varying query count at 120 frames; the authors state the focus is on latency and memory trends rather than absolute accuracy.
Perspective
The result targets settings that need trajectories for all visible surfaces over long videos, such as video understanding, robotics, and dynamic world models; the method processes video in sliding windows and depends on camera poses and depth from sensors or feedforward reconstruction models such as VGGT, so it applies where geometry can be provided or estimated. The authors note that the modular design lets the model benefit from advances in geometry backbones and sensor hardware without retraining, and that users can protect important points from merging through the non-voxelized query-point branch. Where dense ground truth exists in simulation, the authors evaluate dense tracking on held-out splits of PointOdyssey and Dynamic Replica.
Cross-window voxelization merges tracks purely by spatial proximity at window boundaries and does so through irreversible mean-pooling, so tracking errors can permanently collapse distinct physical surfaces into a single canonical track; the authors list improving the de-duplication mechanism as an important direction. They also note that active memory still expands during perpetual open-world exploration, making spatial cache eviction for hour-long trajectories a natural next step, and that performance remains coupled to underlying pointmap fidelity. In addition, TAPVid-3D annotates only sparse query points, so dense-ground-truth evaluation is limited to simulated data; D4RT has no public code or weights, uses private training data, and is evaluated only on 48-frame clips, so its numbers are included for reference. This document is the full text, but tables appear as text and some values are missing in the parse, so a few specific ablation numbers cannot be restated here.
