MoonGS reconstructs lunar terrain from image pairs with a frozen vision foundation model and semantic ranking loss, reaching 23.80 dB PSNR on LuSNAR
Synopsis
MoonGS is a feed-forward 3D Gaussian Splatting framework for lunar scenes: given only two images it predicts pixel-aligned Gaussian primitives in a single forward pass with no per-scene optimization, using robust depth features from a frozen vision foundation model (DUSt3R, or VGGT), semantic priors plus a semantic ranking loss to regularize weakly textured background depth, and an entropy-guided heuristic resampling strategy to augment sparse observations. On the LuSNAR benchmark and the authors' weak-texture MoonBlender dataset it outperforms feed-forward NeRF/3DGS baselines on PSNR, SSIM and LPIPS while keeping sub-second inference.
Ref.1
arXivInterpretation
Introduces MoonGS, a feed-forward 3D Gaussian Splatting framework tailored to lunar scenes that regresses pixel-aligned Gaussian parameters from an image pair and renders novel views without any per-scene optimization. Prior 3DGS methods require per-scene training from scratch and dense captures, while feed-forward methods such as pixelSplat and MVSplat are not designed for planetary scenes with limited overlap and weak texture; MoonGS explicitly formats sparse lunar data as image pairs with a dedicated architecture. Compared against Du et al., MuRF, pixelSplat and MVSplat on LuSNAR and MoonBlender using PSNR/SSIM/LPIPS, with baselines retrained on these datasets; MoonGS reaches 23.80 dB PSNR on LuSNAR versus the best baseline's 18.89 dB, at 0.200 s inference and 18.1 M trainable parameters.
Uses a frozen vision foundation model to supply cross-view depth features, replacing feature extraction that relies on epipolar constraints or image-to-image matching, so depth remains reliable where viewpoint overlap is insufficient. pixelSplat's DINOv2 is independent of depth estimation and MVSplat relies on UniMatch matching between two images, which is less robust under occlusion; MoonGS instead uses DUSt3R features pretrained on large-scale data that generate per-view pointmaps, keeping the backbone frozen and training only downstream modules. Freezing the backbone reduces trainable parameters and training cost, consistent with prior work using frozen visual backbones in low-data regimes; ablations show multi-level feature fusion with a 2D U-Net adds about 0.5 dB PSNR on LuSNAR and nearly 1 dB on MoonBlender.
Integrates semantics in two ways: concatenating one-hot semantics with visual features to refine Gaussian parameter estimation, and applying a semantic ranking loss to regularize depth in weakly textured background regions. The ranking loss enforces ordinal farther-nearer relations rather than metric depth, so it needs no absolute far distance and avoids the loss of genuine far-range structure and scale bias that zeroing background splat opacity or clamping depth to a fixed far plane would cause; the fusion branch can be gated off when semantics are unreliable. Semantic fusion yields roughly 0.8 dB PSNR improvement on LuSNAR and about 0.2 dB on MoonBlender; the ranking loss is enabled with weight 0.01 during a 500-step fine-tune after 30,000 training steps, and the authors report it transfers to pixelSplat and improves sky depth.
Proposes an entropy-guided heuristic resampling strategy and releases MoonBlender, a weak-texture lunar dataset. Resampling uses the entropy of splat footprint coverage as a view priority score to self-supervise selection of the most informative distant candidate viewpoints without ground-truth images; MoonBlender renders 1500 images at 1200x1200 along five traversable routes with varied Sun azimuth and elevation, with a larger share of weaker-textured sky pixels. On the LuSNAR sequence, resampling achieves the highest 24.09 dB PSNR, above fixed-interval resampling (22.46-23.97 dB range); the authors randomly sampled 96 CE4 TCAM images, of which 53 contained sky background with a mean background pixel ratio of 42.11%, indicating that sky-dominated frames are common in practice.
Perspective
The work targets feed-forward novel view synthesis from image pairs, suited to lunar rover settings with known relative camera poses and available semantics (expert annotations or segmentation models such as SAM); the authors note the framework can extend to three or more views because the foundation model supports multi-image inputs and the cost volume formulation generalizes to pairwise matching, and larger maps can be obtained by fusing outputs from multiple inferences. The number of Gaussian primitives per forward pass equals the number of input pixels, so per-pass memory depends on image resolution and per-primitive state rather than terrain size. For a reader, this positions the method as a representation layer for onboard terrain understanding and visual localization that can keep improving by swapping in stronger foundation models.
Semantics are used in several modules; the authors provide a fallback mechanism that preserves stability when semantics are unavailable, and plan to reduce this dependence using edge cues, regionness cues, uncertainty-aware pseudo-labels and surface normals. Inference is fast but the network is not yet lightweight enough for constrained devices, with distillation, pruning and quantization proposed as next steps. The backbone stays frozen during training, and the authors plan to partially unfreeze and fine-tune selected layers as larger, more diverse datasets become available. The pipeline assumes known relative camera poses, leaving pose-free variants to future work. In addition, MoonBlender is described by the authors as relatively simple in visual and environmental settings, so its role in evaluating reconstruction quality is limited, and the real-data validation is qualitative, covering three illumination conditions.
