VolS-GS solves volumetric subsurface scattering with a differentiable finite-volume solver on the Gaussian occupancy domain, improving relighting on three OLAT benchmarks
Related research and updatesSynopsis
VolS-GS turns the spatial support of a Gaussian scene into an occupancy field on a coarse grid and runs a differentiable finite-volume light-transport solve on it, with a small network predicting per-Gaussian scattering and absorption coefficients so light can propagate through the object's interior; across three OLAT benchmarks it consistently improves relighting quality on held-out lights and views over GS3, SSS-GS, RNG, and SSD-GS.
Figure 1: VolS-GS. proposes a high-quality relighting pipeline with three different modalities. Center: a rendered image can be decomposed into normal, diffuse, specular, subsurface scattering (SSS) and shadow. Left: VolS-GS faithfully synthesizes novel view renders. Right: relighting.
arXivInterpretation
A differentiable volumetric subsurface scattering defined on the Gaussian scene's occupancy, allowing light to propagate across primitives. Earlier relightable Gaussian methods such as SSS-GS and SSD-GS model the subsurface response as a function of one primitive's own attributes, while Radiosity-GS and Ref-Gaussian transport light only through the space outside the object; this work rasterizes the Gaussian cloud into a fractional occupancy field, fills the interior with a generalized winding number, and solves for the fluence on the occupied cells. The method is evaluated on three OLAT datasets and ablated one component at a time on seven scenes; on Bunny removing the volumetric SSS channel drops PSNR from 40.15 to 34.67, the largest effect of any arm on any scene.
The shadow term reads the same transport field, and a specular-leakage regularizer keeps the learned shadow and specular branches from trading against each other. The shadow decay is predicted from geometric visibility together with transmittance cues taken from the transport field, with the hint vector gradient-stopped; the regularizer penalizes specular energy only in pixels the shadow term predicts to be unlit. Ablations show that removing the transport cues drops Statue PSNR from 38.35 to 35.09, and removing the shadow term is the largest single change on six of seven scenes; the regularizer does not earn PSNR but collapses the specular footprint on Bunny at unchanged PSNR.
An adjoint formulation makes the transport solve differentiable, so backpropagation costs one extra solve and its memory does not grow with the iteration count. The operator is symmetric positive definite, so no transpose is needed and one solve yields gradients with respect to the source and to the diagonal of the operator; face conductances are held at stop-gradient so the medium is learned through the terms it enters directly. The appendix states propositions on well-posedness and stability, adjoint gradients for the full operator, and reciprocity in heterogeneous media, and the implementation reports that both the forward and the adjoint solve are truncated, so the gradient is the implicit-function estimator at an approximate solution.
Best average performance across three OLAT benchmarks, with the most consistent gains on the subsurface-dominant SSS-GS benchmark. Compared against GS3, SSS-GS, RNG, and SSD-GS retrained under matched resolution, splits, and training budgets, with 3DGS and GI-GS listed as additional reference points. Table 1 reports dataset means: on the SSS-GS synthetic scenes VolS-GS reaches PSNR 40.15 on Bunny and 43.64 on Candle, above every baseline, and it also leads on NRHints (32.37) and GS3 (31.59).
Perspective
The result targets OLAT capture with a known point light at a known position, used to re-render objects under novel lights and viewpoints; the authors note that because the solve is based on a diffusion approximation it may not suffice where radiance is still directional, and that it carries no path for light exchanged between surfaces on concave regions. The method takes the Gaussian cloud as its transport domain, and the grid carries no learnable parameters of its own, so it applies to objects represented by Gaussian primitives whose interior can be described by an occupancy field. The authors report training under 3.5 hours and under 8 GiB of GPU memory per scene on the SSS-GS synthetic scenes, and state that environment-map relighting can be obtained by linear superposition of sources without retraining.
The diffusion approximation is weakest at the boundary and for the first, still-directional scattering event, which the authors supplement with a separate single-scattering term integrated along the view chord, and they report that very thin geometry such as the Bunny ear tip still renders slightly dark, with grid refinement moving that region only by a few dB. Substituting a per-primitive dipole costs accuracy on six of seven scenes but wins slightly on Drums, indicating that the channel's benefit tracks how translucent the medium is. The specular-leakage regularizer does not earn PSNR; its value lies in keeping the decomposition identifiable, so it should be judged on the decomposition rather than on metrics. In addition, several equations, table values, and appendix numbers appear as placeholders in the provided text, so readers who need exact coefficients, iteration counts, and per-frame timings should consult the original equations and appendix tables.
