Skip to main content
Back to timeline
arXivSource publication:

VGGT-Diff routes VGGT-Ω geometry latents into Wan video diffusion, topping PSNR, LPIPS and DreamSim on 6,188 DL3DV targets

Synopsis

VGGT-Diff is a geometry-routed multi-view diffusion framework that uses a frozen VGGT-Ω to transform source-view features together with 3D points and confidence into query-aligned conditions via a confidence-aware Visual Geometry Router, injects them into a pretrained Wan2.1-I2V-14B video diffusion model, and regularizes predicted-clean residuals along reliable 3D tracks with Point-Track Residual Consistency; trained on only 980 DL3DV scenes, it ranks first in PSNR, LPIPS and DreamSim and second in SSIM on 6,188 DL3DV-Benchmark targets, and first in LPIPS and DreamSim with second-place PSNR on zero-shot Mip-NeRF 360.

AI-generated editorial illustration: VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis

Interpretation

The paper proposes a geometry-routed generation framework: a frozen VGGT-Ω associates each source feature with a 3D point and confidence, the Visual Geometry Router builds source-to-query correspondence through a depth-selected hard anchor (a hard z-buffer averaging features weighted by normalized VGGT-Ω confidence), and confidence-weighted front and secondary hypotheses serve as residual cues refined by a zero-initialized residual refiner near occlusion boundaries and projection errors. Prior generative NVS left source-to-query correspondence largely implicit, while reconstruction-based methods struggle to complete unobserved regions; this work explicitly converts geometry-foundation visual tokens, 3D locations and confidence into query-aligned diffusion conditions rather than treating recovered geometry as a complete scene representation or pre-rendered RGB. Ablations show the largest degradation when all VGGT-Ω-dependent components are removed, with PSNR falling and LPIPS rising the most; replacing routed features with point-rendered RGB costs PSNR and raises LPIPS, indicating appearance-bearing features provide richer evidence than incomplete RGB renderings; removing layered residual refinement causes only a smaller drop, indicating hard routing captures most of the benefit.

The paper proposes Point-Track Residual Consistency (PTRC), which reuses VGGT-Ω point predictions to establish tracks across every pair of query views, retaining only points that project inside both views, lie in front of both cameras and pass a per-view z-buffer visibility test, and aligns confidence-weighted Smooth-L1 predicted-clean residuals rather than raw velocities or VAE latents. Per-view flow-matching supervision does not enforce geometric coherence among remaining errors, while directly matching RGB, latents or velocities would be overly restrictive because viewpoint-dependent illumination and visibility legitimately alter these representations; PTRC couples only the error that should be removed, suppressing cross-view drift while retaining valid appearance changes. On 192 pose-stratified targets, PTRC improves PSNR and SSIM and reduces LPIPS relative to a control that applies the same point tracks, confidence weights and robust penalty directly to predicted velocities, with consistent gains across all six pose bins; removing PTRC is the single-component ablation with the largest SSIM loss.

The paper introduces geometry-condition regularization and inference-time Geometry-Prior CFG: during training the routed geometry condition is randomly retained, attenuated or removed while source appearance and Plücker rays stay unchanged, and at inference a reference branch that zeros only query-slot dense geometry smoothly strengthens the geometry prior over the first 20% of denoising steps. Geometry recovered from sparse observations is inevitably imperfect, and treating it as an infallible reconstruction invites over-reliance; this work treats geometry as uncertainty-aware evidence and lets the same dropout mechanism supply a matched guidance branch at inference. Removing condition regularization costs PSNR and SSIM and raises LPIPS; applying Geometry-Prior CFG to the same checkpoint further improves PSNR and SSIM while reducing LPIPS, which the authors read as regularization being the primary mechanism and CFG a complementary inference-time gain requiring no retraining.

The paper reports cross-dataset and cross-pose-difficulty evaluations and probes cross-view consistency with two frozen reconstructors. Single-target metrics cannot determine whether jointly generated views support a coherent scene, so the authors use Pi3, unused in training or conditioning, as an independent reconstructor alongside the VGGT-Ω primary probe. On 6,188 DL3DV targets the method ranks first in PSNR, LPIPS and DreamSim and second in SSIM while training on only 1K scenes; against FrameCrafter, which shares the Wan2.1-I2V-14B prior, it gains PSNR and SSIM; on 192 pose-stratified targets it leads five bins and trails SEVA only in interpolation-near; across 52 scenes and eight views it ranks first under both reconstructors, with Chamfer-L1 and F-score improvements over FrameCrafter, and the authors note these metrics measure relative consistency and reconstructability rather than absolute 3D accuracy.

Perspective

The result targets posed sparse source views to prescribed query cameras for static-scene novel view synthesis, trained on the 1K-scene split of DL3DV-10K (980 scenes after removing one unavailable scene and 19 overlapping with the 140-scene benchmark) and evaluated on 6,188 DL3DV-Benchmark targets plus zero-shot Mip-NeRF 360. The method freezes the VAE and VGGT-Ω and fine-tunes only the Wan DiT, expanded input projection, visual-feature adapter and Visual Geometry Router, so it can reuse an existing video diffusion prior directly; the authors note geometry routing and residual consistency can extend to more views, longer camera trajectories and dynamic non-rigid content, and point to multi-level feature fusion, parameter-efficient adaptation and richer visibility modeling as directions. Cross-view consistency metrics use target-RGB reconstructions as references and are meant for relative comparison rather than absolute 3D accuracy.

The loaded text is a fast parse, so the concrete values in Fig. 2, Fig. 4, Fig. 5, Fig. 6, Fig. 7, Fig. 8, Fig. 10 and Tables 2, 4, 5, 6, 7, 8, 9, 10, 12 and 13 are not fully present, and some gain magnitudes can only be taken from the prose. The authors report local regressions at the 42K, 60K and 88K checkpoints and a narrowing PSNR gain at 100K, leaving whether optimization is approaching saturation an open question; intermediate VGGT-Ω feature depths perform better while joint optimization degrades, which the authors attribute to flow-matching gradients altering the pretrained representation, an explanation still to be verified; cross-view consistency is referenced to reconstructions rather than absolute 3D accuracy; and the three-view and four-view robustness results come from a six-view-trained checkpoint without adaptation.

Sources