Public articles linked to the same research event.
arXiv The authors present SNAP, a self-supervised encoder-decoder transformer in which a pose-free encoder aggregates frozen DINOv3 patch features via full cross-view self-attention, while a PRoPE pose-conditioned decoder with a BlockNN mask reduces target-view self-attention to a pointwise residual MLP and is trained to reconstruct frozen DINOv3 latent features with a first-order spatial finite-difference penalty; after training on roughly 117,000 sequences from RealEstate10K, DL3DV, and Co3Dv2, the frozen encoder matches or exceeds self-supervised baselines on point correspondence, visual localization, relative pose, relative depth, and MimicGen robot manipulation, and degrades more gracefully than 2D semantic encoders under camera shifts.
The authors present SNAP, a self-supervised encoder-decoder transformer in which a pose-free encoder aggregates frozen DINOv3 patch features via full cross-view self-attention, while a PRoPE pose-conditioned decoder with a BlockNN mask reduces target-view self-attention to a pointwise residual MLP and is trained to reconstruct frozen DINOv3 latent features with a first-order spatial finite-difference penalty; after training on roughly 117,000 sequences from RealEstate10K, DL3DV, and Co3Dv2, the frozen encoder matches or exceeds self-supervised baselines on point correspondence, visual localization, relative pose, relative depth, and MimicGen robot manipulation, and degrades more gracefully than 2D semantic encoders under camera shifts.
The authors present SNAP, a self-supervised encoder-decoder transformer in which a pose-free encoder aggregates frozen DINOv3 patch features via full cross-view self-attention, while a PRoPE pose-conditioned decoder with a BlockNN mask reduces target-view self-attention to a pointwise residual MLP and is trained to reconstruct frozen DINOv3 latent features with a first-order spatial finite-difference penalty; after training on roughly 117,000 sequences from RealEstate10K, DL3DV, and Co3Dv2, the frozen encoder matches or exceeds self-supervised baselines on point correspondence, visual localization, relative pose, relative depth, and MimicGen robot manipulation, and degrades more gracefully than 2D semantic encoders under camera shifts.
The authors present SNAP, a self-supervised encoder-decoder transformer in which a pose-free encoder aggregates frozen DINOv3 patch features via full cross-view self-attention, while a PRoPE pose-conditioned decoder with a BlockNN mask reduces target-view self-attention to a pointwise residual MLP and is trained to reconstruct frozen DINOv3 latent features with a first-order spatial finite-difference penalty; after training on roughly 117,000 sequences from RealEstate10K, DL3DV, and Co3Dv2, the frozen encoder matches or exceeds self-supervised baselines on point correspondence, visual localization, relative pose, relative depth, and MimicGen robot manipulation, and degrades more gracefully than 2D semantic encoders under camera shifts.