Writing Geometry into the Generative State: GAE Unifies Perception and Generation with a Geometry-Native Latent Space
Synopsis
The work introduces the geometry-native autoencoder (GAE), which compresses the four feature levels of the frozen geometry foundation model DA3 into a single compact latent of 64 or 128 channels that decodes jointly to RGB, depth, cameras, and point maps, and in controlled comparisons holding the generator and training protocol fixed it lowers FVD on RealEstate10K and DL3DV and roughly halves camera-trajectory error on RealEstate10K relative to the strongest non-GAE latent.
Interpretation
It proposes GAE: a learned bottleneck placed between the frozen DA3 encoder and the frozen geometry head fuses and compresses the four-level hierarchy into one latent and reconstructs the whole hierarchy in one pass, so the same latent is read by the frozen geometry head as depth, camera rays, and point maps while a learned RGB head decodes appearance. Prior designs either add geometry to an appearance-led latent as an auxiliary output, control signal, or post-training objective, or generate directly over the high-dimensional geometry-foundation feature hierarchy (such as GLD's cascade or latent Riemannian flow matching jointly evolving four normalized levels on their product manifold); GAE instead learns one compact Euclidean reparameterization of the full hierarchy evolved by a single flow model. The method is specified with a two-stage training scheme (codec first, then a frozen codec with a conditional flow model) and formal objectives; in the controlled comparison GAE-128 matches or slightly exceeds raw L0 features on PSNR, LPIPS, rFID, and rFVD while using 24 times fewer channels.
It argues that how a latent is organized matters as much as what it reconstructs: a reconstruction-only codec is compact and well-conditioned yet weak in transport smoothness and semantic neighborhoods; token-wise alignment to C-RADIO improves transport smoothness and semantic neighborhood consistency but sharply degrades pairwise spatial structure, which matching pairwise DINOv2 similarities restores. It moves representation alignment from an intermediate denoiser representation (as in REPA) to the codec latent itself, and shows that position-wise alignment alone destroys relational structure and needs a complementary relational objective. Table 1 reports a three-step progression of representation diagnostics (transport ambiguity, condition number, effective rank, LNC@5, LDS, SRSS, xLNC), and Table 8 shows the improved latent quality transfers to downstream appearance, cross-view geometry, and camera recovery.
Under a fixed flow model, training budget, and camera-conditioning mechanism, the GAE latent improves both appearance quality and independently measured 3D coherence: GAE-64 gives the strongest overall results on both datasets, reducing FVD on RealEstate10K and DL3DV relative to the best non-GAE controlled latent and greatly reducing ATE on RealEstate10K. Appearance-led latents, semantic representation autoencoders, and raw geometry features each have distinct weaknesses (pixel VAEs are well-conditioned but weakly structured, semantic RAEV2 has strong semantic neighborhoods but no native geometry readout, raw DA3 features are geometry-native but high-dimensional and poorly conditioned); GAE gives the best overall balance of transport, conditioning, and structure among geometry-native representations. The controlled setup uses a shared held-out pool of 64 scenes, nine views per scene, one reference view, 50 Euler steps, and a fixed CFG; camera metrics are reconstructed by VGGT, which is independent of all evaluated latent backbones, with DA3-GIANT included only as a backbone-overlap cross-check.
The same conditional flow model supports text-to-image, camera-controlled video, and reference-conditioned novel-view synthesis through condition dropout, and the work showcases an 81-view long rollout and text-conditioned generation; reference views are kept outside the ODE state as clean conditioning tokens to avoid a set-encoding train-inference mismatch. Separating reference evidence from the noisy state lets conditional and unconditional sampling share one ODE parameterization without repeatedly clamping a reference slot, and metric Plücker-ray conditioning retains the translation scale discarded by per-scene pose normalization. Ablations show clean conditioning tokens lower FVD and reduce the reference-target point-cloud gap relative to placing the reference latent in the ODE state, metric pose substantially improves scale-sensitive camera recovery over per-scene normalized pose, and text-to-image co-training improves FVD and LPIPS without sacrificing MEt3R.
Perspective
The results target visual generation and 3D perception settings where geometric consistency matters, and apply to world-model-style generation tasks that require camera control and cross-view consistency; the controlled conclusions rest on shared held-out pools from RealEstate10K and DL3DV under a fixed flow-model protocol, while the larger-scale, higher-resolution, longer-sequence results are qualitative and not part of the controlled comparison.
Several numbers in the abstract and main text appear as placeholders in the evidence bundle (such as the FVD reduction and the condition-number range), so exact values should be checked against the original tables; in addition, geometry supervision uses pseudo-targets read from the frozen DA3 head on original features rather than dataset depth labels, and the long-rollout and text-conditioned results are currently qualitative capability checks, leaving open how the approach behaves in other data domains, over longer sequences, or with different geometry backbones.
