Skip to main content
Back to timeline
arXivSource publication:

GEAR treats geometry as a memory address: sparse correspondence attention cuts ATE by 80.3% for minute-long camera-controlled video generation

Synopsis

The work introduces GEAR, which avoids fusing history into a persistent global 3D representation and instead uses per-frame geometry to build token-level history-to-target correspondences, treating geometry as an explicit address for visual memory; Geometric Correspondence Attention then sparsely injects only geometrically matched historical features, while an Invisible Octree rejects projectable but occluded correspondences, reporting state-of-the-art visual quality, camera control, and revisit consistency on DL3DV-Evaluation and WorldScore-Static and enabling minute-long generation.

AI-generated editorial illustration: Geometry as Address: Routing Attention to Visual Memory for Long-Horizon Camera-Controlled Video Generation

Interpretation

GEAR decouples geometry from visual memory: geometry only answers where to read from, while attention decides what to recover. Prior approaches either search historical context implicitly with dense attention or reconstruct history into persistent 3D memory and reproject it; GEAR keeps history as frame latents and uses per-frame geometry only to establish token-level correspondences, making geometry a transient address rather than persistent scene state. The paper supports this with design and ablations: removing GCA's residual injection increases temporal instability and deviation from camera motion, and replacing correspondence-restricted attention with dense attention over the full retrieved history degrades camera control and visual fidelity.

Geometric Correspondence Attention lets each target token access only geometrically matched historical tokens as keys and values, injected through a residual branch during denoising. Compared with UCM's warped positional encodings and Lyra 2.0's injected warped correspondence embeddings, GEAR explicitly restricts each target patch to attend to its matched historical patch set, giving direct access to historical visual features. GCA is inserted into every even-indexed DiT block with hidden dimension 640, totaling 262M parameters, only 1.9% of the backbone; the backbone stays frozen while rank-32 LoRA adapters and GCA are jointly trained for 10K iterations.

An Invisible Octree accumulates visibility evidence and rejects correspondences that are geometrically projectable but occluded in the target view. Cross-view projection alone can produce false correspondences; the octree stores no appearance and never serves as a rendering condition, providing only a coarse binary filter. Ablation shows that under large viewpoint changes, projectable yet occluded correspondences introduce local appearance errors and propagate structural inconsistencies through autoregressive history, which the module suppresses.

On DL3DV-Evaluation and WorldScore-Static, GEAR reports the best visual quality, camera-control accuracy, and revisit consistency. It reduces ATE by 80.3% versus the strongest baseline and achieves the best Content Alignment, Photometric Consistency, Style Consistency, and revisit performance on WorldScore (Revisit SSIM 0.6489, Revisit LPIPS 0.2019). On DL3DV-Evaluation: SSIM 0.3645, LPIPS 0.4459, FVD 837.59, TransErr 0.0116, RotErr 0.1228, ATE 0.0436; WorldScore-Static uses 50 randomly sampled scenes with closed-loop camera trajectories, with revisit metrics computed against the reference observation at the matched camera pose.

Perspective

The result targets long-horizon, camera-trajectory-controlled autoregressive video generation, especially closed-loop trajectories that revisit previously observed regions; the method relies on an external 3D model (Depth Anything 3) for camera poses and per-frame depth, and requires frame-aligned VAE encoding so that latent tokens map one-to-one to camera poses. For users, this means fine-grained historical memory access can be obtained by lightweight adaptation of a frozen pretrained video backbone, without building a persistent global 3D representation.

The paper states that it still relies on an external 3D model for depth estimation, adding computational overhead, and lists jointly learning geometry and memory addressing inside the generative model as future work. The correspondence-disturbance ablation keeps only 2-3 valid observations out of 9 and perturbs the rest by 64-128 pixels, so how far the robustness conclusion extends to stronger depth errors remains open. In addition, protocol differences such as camera-scale alignment, baseline depth substitution, and action discretization affect how cross-method comparisons should be read; this summary is based on the paper text and appendix and does not include the figures or video results themselves.

Sources