Skip to main content
Back to timeline
arXivSource publication:

SNAP constrains the decoder and predicts latent features so novel-view-synthesis pretraining yields transferable geometry

Related research and updates

Synopsis

The authors present SNAP, a self-supervised encoder-decoder transformer in which a pose-free encoder aggregates frozen DINOv3 patch features via full cross-view self-attention, while a PRoPE pose-conditioned decoder with a BlockNN mask reduces target-view self-attention to a pointwise residual MLP and is trained to reconstruct frozen DINOv3 latent features with a first-order spatial finite-difference penalty; after training on roughly 117,000 sequences from RealEstate10K, DL3DV, and Co3Dv2, the frozen encoder matches or exceeds self-supervised baselines on point correspondence, visual localization, relative pose, relative depth, and MimicGen robot manipulation, and degrades more gracefully than 2D semantic encoders under camera shifts.

Source-provided article image: Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis
Figure 1 ·

Figure 1 : Training and inference pathways. Top (Inference): A pose-free multi-view encoder processes reference frames via a frozen DINOv3 featurizer ( φ \varphi ) and cross-view self-attention to construct an implicit 3D scene latent Z Z . Bottom (Training): A pose-conditioned local decoder queries Z Z using PRoPE to predict the feature map of a novel target view. Supervision ( ℒ recon \mathcal{L}_{\mathrm{recon}} ) is applied in a semantic latent space against a frozen DINOv3 teacher, completely bypassing pixel-level rendering and forcing geometric structure into the encoder.

arXiv

Interpretation

SNAP reframes novel view synthesis from an image-reconstruction task into a geometric representation pretraining task: the encoder receives no camera parameters, the decoder receives target and context intrinsics and extrinsics, and PRoPE-rotated attention scores token pairs highly when their world-space rays are co-aligned, so the model predicts the target view's DINOv3 feature map in latent space. Earlier encoder-based NVS methods (LVSM, RayZer, LagerNVS and others) were evaluated mainly by image reconstruction quality and transferred poorly to geometric tasks; SNAP replaces RGB pixel supervision with frozen DINOv3 semantic latent features and confines pose conditioning to the decoder. The method is trained on roughly 117,000 sequences from RealEstate10K, DL3DV, and Co3Dv2, with a 16-layer multi-view encoder of embedding dimension 1024 and 8 blocks of alternating PRoPE cross-attention and masked self-attention, using GQA with 16 query heads and 8 KV groups, for 40,000 steps at effective batch size 256, with an EMA of decay 0.999 over all trainable parameters; evaluations use the EMA weights.

Larger decoder receptive fields lower reconstruction error but worsen downstream geometric transfer, whereas restricting target-view self-attention to a pointwise residual MLP (BlockNN window of 1) simultaneously improves point correspondence, visual localization, and pose estimation. The work turns the observation that reconstruction fidelity is not representation quality into an independently controllable mechanistic experiment: only the BlockNN mask window is varied, from 1 to spanning the full feature map, leaving the rest of the decoder architecture unchanged, thereby separating decoder expressivity from prediction space. The receptive-field sweep shows the reconstruction objective improving as receptive field grows while point correspondence error, visual localization, and pose estimation degrade; qualitatively, with window 1 similarity maps are sharply localized around the corresponding point and stable across viewpoints, while with full self-attention they do not always track the ground-truth correspondence.

The feature space of the prediction target determines representation quality: easier lower-layer targets closer to RGB do not yield the best spatial representations, and predicting higher-layer DINOv3 features improves downstream geometry; input layer 4 gives better in-distribution PCK and MRR than layer 11 but higher intrinsic dimensionality, so the authors choose layer 11 as input for a more compressed and invariant multi-view representation. Prior NVS models typically supervise reconstruction in pixel space, which can emphasize appearance, texture, and illumination; SNAP systematically sweeps DINOv3 input-layer and target-layer pairings and characterizes the trade-off with feature reconstruction error, PCK@0.1, patch retrieval MRR, and TwoNN intrinsic dimensionality. The sweep is run on RE10K and DL3DV validation subsets, reporting feature reconstruction error, nearest-neighbor point correspondence PCK@0.1, patch retrieval MRR, and TwoNN intrinsic dimensionality; Appendix G further uses mAP/MRR, cycle consistency, and variance ratio to show that mid-layer inputs with last-layer targets raise raw correspondence but lower cycle consistency.

The frozen SNAP encoder matches or exceeds self-supervised baselines across five tasks and, without any robotics training data, outperforms its own DINOv3 teacher and generally Adapt3R on MimicGen under camera shifts, tracking closely with the fully geometry-supervised VGGT. The work evaluates the NVS-pretrained representation uniformly on point correspondence, visual localization, relative pose, relative depth, and robot manipulation, with depth and manipulation evaluated on domains unseen during pretraining (Hypersim, 7Scenes, MimicGen) to test out-of-domain generalization. Point correspondence exceeds self-supervised baselines on RealEstate10K at all pixel thresholds and approaches VGGT on DL3DV; visual localization is at the self-supervised frontier on most metrics; relative pose shows high rotation accuracy among NVS-based methods; depth on 7Scenes reaches AbsRel 0.9257 versus 1.9474 for LVSM and 2.3858 for LagerNVS, though VGGT is better on Hypersim at 1.9138; robot manipulation is evaluated under camera shifts from 0 to some radians, where semantic encoders (DINOv3, ResNet) collapse rapidly beyond some radians.

Perspective

The results target researchers and engineering teams doing self-supervised geometric pretraining from posed multi-view images, in static, textured scenes and under a fixed compute and data budget for encoder-decoder NVS architectures; the method transfers directly to pretraining pipelines that use frozen visual teacher features as targets and has been validated on point correspondence, localization, pose, depth, and simulated robot manipulation. The authors note that current pretraining relies on posed video sequences and static textured scenes and that a fixed compute and data budget has yet to reveal the architecture's full scaling laws, with extension to dynamic video as the next step.

Readers should still watch: the receptive-field and prediction-space sweeps are reported mainly on RE10K and DL3DV validation subsets, with cross-dataset and cross-teacher robustness supplied in appendices; on depth, the explicitly 3D-supervised VGGT remains better on Hypersim, and LagerNVS and SNAP are not resolution-matched (518px versus 256px), a difference the authors note may affect the depth comparison; robot manipulation is evaluated in simulated MimicGen, leaving real-platform behavior open; dynamic video and larger-scale scaling laws remain open questions.

Sources