GeoVerse injects video generative priors into geometric latent space, raising PSNR by 2.23 dB on DL3DV and cutting ATE by 32.4% on Mip-NeRF360
Synopsis
GeoVerse performs diffusion generation inside the geometric latent space of a pretrained 3D foundation model (DA3), injects frozen Wan2.2 VACE video features into the denoiser through a ControlNet-style adapter, and uses a continuously updated colored point-cloud spatial memory to reproject history as target-aligned guidance, synthesizing cross-view consistent novel views from sparse inputs; across DL3DV, RealEstate10K, Mip-NeRF 360, and ScanNetV2 it reports better visual quality and geometric consistency, with PSNR 2.23 dB higher than GLD on DL3DV and ATE 32.4% lower than GLD on Mip-NeRF 360, while reflow distillation brings inference from over 100 seconds to under 10 seconds.
Interpretation
GeoVerse places generation inside the multilevel geometric latent space of a frozen DA3 encoder and designates Level-1 as the synthesis boundary, so diffusion runs directly on geometric features that carry cross-view correspondence rather than generating frame by frame in image or video latent space. GLD had already repurposed DA3's geometric feature space as the native latent space for multi-view diffusion but struggled with generative fidelity and stability on in-the-wild scenes due to a limited training distribution; GeoVerse keeps that geometric-latent formulation and adds video generative priors plus a persistent 3D memory on top of it. The paper provides the full method derivation: Eq. (1) to Eq. (3) describe merging context and target views into the encoder input, Level-1 flow-matching denoising, and the Level-0 conditional cascade, and state that Level-1 is chosen as the boundary balancing geometric accuracy and visual fidelity.
A ControlNet-style adapter injects multilevel features from a frozen Wan2.2 VACE model into the geometric denoiser as additive residuals, so video-learned appearance priors supply the structural completion that the geometric latent space lacks in unobserved regions. Unlike applying video diffusion models directly or sequentially to 3D scenes, which accumulates cross-frame inconsistency, GeoVerse extracts video features once and reuses them across geometric denoising steps, without sampling a complete video or invoking the VAE decoder. Ablation shows removing the Wan2.2 prior causes the most severe visual degradation (on ScanNetV2, PSNR drops from 19.65 to 16.83 and LPIPS rises from 0.398 to 0.463); Appendix A states features come from blocks 0, 5, 10, and 15 of the low-noise component of Wan2.2-VACE-Fun-A14B at 5,120 dimensions, with zero-initialized projections keeping the pretrained denoiser initially unmodified.
A global colored point-cloud spatial memory continuously aggregates observed and synthesized content and performs depth-aware splatting toward target cameras, producing RGB-D guidance and a validity mask that anchors multi-round view expansion to a shared scene representation. Single-pass consistency is not enough for successive view expansion, which needs a persistent representation across multiple inference steps; GeoVerse follows ViewCrafter's point-cloud memory idea and encodes the projected memory as target-side conditioning in Eq. (1), where valid projections anchor known structure and masked regions are left to the generative prior. Ablation shows disabling spatial memory degrades both visual fidelity and cross-view alignment (PSNR 19.65 to 17.87, ATE 0.009 to 0.014); in three-round long-sequence ScanNetV2 evaluation, PSNR rises from GLD's 15.33 dB to 19.65 dB.
Piecewise reflow distillation of the Level-1 denoiser compresses denoising to 4 steps (three Level-1 steps plus one Level-0 cascade step), cutting inference latency substantially while preserving quality. The paper reports over 17 times speedup relative to GLD, with per-inference latency dropping from over 100 seconds to under 10 seconds; distillation applies only to Level-1, leaving the cascade and decoding heads unchanged. The step sweep in Table 2 shows that reducing Level-1 from three steps to one lowers runtime from 9.24 s to 6.07 s but worsens LPIPS from 0.339 to 0.367, whereas cutting the Level-0 cascade from 49 steps to one gives LPIPS 0.342 versus 0.339 at much lower latency, supporting the current budget allocation.
Perspective
This work targets novel view synthesis from sparse reference images, covering view interpolation and a degree of extrapolation, and is validated on two in-domain benchmarks (DL3DV, RealEstate10K) and two out-of-domain benchmarks (Mip-NeRF 360, ScanNetV2), with ScanNetV2 used for long-sequence evaluation. The training mixture spans 15 real and synthetic multi-view datasets totaling 87,886 scenes and 23,760,939 images, with dataset sources sampled with probabilities proportional to the square root of their RGB frame counts, and RealEstate10K and Waymo using Pi3-estimated poses and depth. Training runs in two stages on 32 NVIDIA A800 GPUs: a 300k-iteration adaptation phase and up to 50k iterations of reflow distillation. For research and engineering settings that want to bring video generative priors into a geometric latent space, or that need multi-round view expansion under latency constraints, this recipe offers reusable component boundaries: a frozen geometry encoder, a frozen video feature extractor, trainable adapters, and a persistent spatial memory.
The paper states two limitations: peak visual fidelity is constrained by the appearance capacity of the DA3 feature space, so prioritizing speed through compact feature conditioning may not fully match the fine texture richness of full-scale video diffusion models; and multi-round consistency depends on both the geometry backbone and synthesized RGB coherence, so in textureless or complex regions initial depth errors and generated appearance drift can accumulate through spatial memory updates, occasionally compromising cross-view alignment over long trajectories. The paper also notes that some baselines achieve lower reprojection or translation errors on individual datasets, indicating performance is not uniformly ahead on every metric. Appendix A notes that the available model card does not specify an absolute training-set size for the particular Wan2.2-VACE-Fun checkpoint used, and Appendix D notes that the inventory counts describe the available training pool rather than the number of distinct images consumed by a completed run. The relative advantages of the MoT variant versus additive residual injection still require comparison under matched training settings.
