RDGSplat decodes render-dedicated geometry on a frozen 3D foundation model, lifting WM2.0 on RE10K from 20.918 to 24.266 dB while leaving its metric predictions unchanged
Synopsis
RDGSplat introduces a two-pass framework that decodes a render-dedicated geometry from a frozen 3D foundation model by duplicating pretrained decoders and optimizing them under photometric supervision alone, together with a target-pose conditioned adapter that reformulates the representation those decoders read, while the model's metric predictions remain unchanged. Across four benchmarks and three feed-forward backbones, it improves novel view synthesis, raising WM2.0 on RE10K from 20.918 to 24.266 dB with 205.5 M added trainable parameters, and the depth and pose the same model predicts are bit-exact.
Figure 2 : The two geometries of a single scene. Both are decoded from the same frozen backbone and the same context views. The metric geometry thickens the far wall into an opaque shell, and the views rendered from it lose the corridor behind it. The render-dedicated geometry resolves those surfaces into thin layers, through which the corridor is rendered.
arXivInterpretation
RDGSplat lets one frozen 3D foundation model serve both rendering and reconstruction: the render branch decodes a render-dedicated geometry from the same frozen weights, while the pretrained metric heads continue to output their original depth, pose and pointmaps. Prior adaptations of foundation models for rendering, such as NoPoSplat and AnySplat, update the backbone weights, thereby discarding the metric predictions the model was built for and requiring repetition for every new backbone; RDGSplat updates no pretrained weight, so the metric output is identical to the pretrained model's by construction. On RE10K, WM2.0 rises from 20.918 to 24.266 dB; on the KITTI depth benchmark, each predicted frame is identical to the pretrained model's with a largest absolute deviation of zero, and the pose AUC agrees to four decimal places.
Render-Dedicated Geometry Decoding (RDGD) duplicates the pretrained depth head and Gaussian decoder and optimizes the duplicates under photometric supervision alone, so primitive placement is constrained by rendering quality rather than a metric estimate. Earlier methods either make one set of weights satisfy both the metric and rendering criteria or finetune the backbone; RDGD gives each geometry its own decoders, letting the duplicates depart from the inherited metric surface and place primitives where they render best. The ablation shows RDGD alone reaches 23.485 dB on RE10K and 24.145 dB on ACID; the supplement decomposes this into 1.538 dB from the Gaussian decoder, 0.882 dB from the depth head and 0.147 dB from the residual pose head, while training the same three decoders from random weights reaches only 17.449 dB, 3.468 dB below the frozen backbone.
The Target-Pose Conditioned Adapter (TCA) injects a condition within the backbone's residual stream, making the representation the decoders read a function of the view to be rendered rather than of the context images alone. A conventional adapter modifies the single representation, so one forward pass cannot serve both behaviors; TCA reformulates the representation in the second evaluation conditioned on the target camera pose, and no target intensity enters the condition, so the primitives are still decoded from the context alone. The ablation shows TCA alone reaches 22.613 dB on RE10K and 23.615 dB on ACID; in the supplement, adding the eight adapter blocks alone reaches 21.520 dB, and conditioning them reaches 22.613 dB, indicating most of the contribution comes from the condition rather than the added capacity.
The design improves novel view synthesis consistently across four benchmarks and three feed-forward backbones, at a training cost far below full finetuning. RDGSplat attaches to existing backbones as added modules, requires no geometric annotation and does not depend on the supervision a backbone was trained under; it optimizes 205.5 M parameters against the 1.4 B a full finetune of the same backbone would require. All backbones improve on all metrics across RE10K, ACID, DL3DV and DTU with no metric regressing; WM2.0 gains 2.676 and 2.043 dB on DL3DV and DTU; of the 205.5 M parameters, only 39.0 M are instantiated from scratch and 166.5 M duplicate modules the backbone already carries.
Perspective
The method targets backbones that expose a token stream and a frozen depth head, and suits settings where metric depth and pose must be retained alongside novel view synthesis, such as 3D tasks that consume structure as well as images. Training uses RE10K, with zero-shot testing on ACID, DL3DV and DTU, and metric geometry evaluated on KITTI and RE10K. The render-dedicated geometry is scored on rendering alone and coexists with, rather than replaces, the metric geometry; camera translation keeps the pretrained estimate as a fixed anchor and learns only a penalized residual. At inference the backbone runs twice, raising latency by a factor of about 1.60 to 1.65 with a peak-memory increase of at most 0.04 GB.
As a measurement, the render-dedicated geometry is slightly less accurate than the pretrained head: the supplement reports median translation error rising from 1.589 to 1.609 degrees, with AUC falling by 0.1 at 10 degrees and unchanged at 20, while rotation is identical. Depth admits no such comparison, because the condition of the second evaluation requires a target frame and a preliminary render, which a single-sequence depth benchmark does not supply. The separate gains of the two modules exceed their joint gain, indicating they partly address the same limitation, and where their division of labor lies remains worth clarifying. Training also draws on one random 20% sample of the RE10K training split, so behavior at larger data scales and on further backbone types remains an open question.
