OX-NeRF attains the highest 3D reconstruction accuracy on ultra-sparse X-ray scenes with four to ten projections
Synopsis
The authors present OX-NeRF, which fuses a cross-scene convolutional encoder with a per-scene multi-resolution hash grid into an MLP and optimises them jointly by self-supervision, achieving higher 3D reconstruction accuracy than existing self-supervised radiance field methods on parallel-beam and cone-beam X-ray datasets in the ultra-sparse regime of four to ten projections.
Figure 1 : Overview of the OX-NeRF pipeline. Left: each scene is recorded by a few X-ray projections at known angles, and a batch of rays is drawn from them by the rule of Section 4.2 . Centre: the scene manager activates the hash grid of the current scene, so the hash encoder reads only that scene’s table, while the shared convolutional encoder reads a pixel-aligned descriptor from the source projections; the fusion module combines the two and a fully fused head returns an attenuation value. Right: attenuation is accumulated along the ray by Eq. 3 and compared with the measurement by mean squared error. Only the hash grids are specific to a scene.
arXivInterpretation
OX-NeRF stores transferable priors and scene-specific geometry separately: a shared ResNet-34 convolutional encoder learns features common to the whole set of scenes, each scene keeps its own multi-resolution hash grid for detail, and the two are concatenated, mapped by a fusion module and passed to an MLP that predicts attenuation. The earlier ONIX held its prior entirely in shared weights, limiting the capacity available to any single scene; OX-NeRF stores and updates shared priors and scene-specific representation separately, so per-scene detail no longer competes for shared capacity. The ablation on Ellipsoids at six training views shows the convolutional encoder paired with the hash grid gives the best configuration by a wide margin, and the hash grid outperforms NeRF's fixed sinusoidal positional encoding in every configuration but one.
Training uses residual-guided ray sampling: each target projection caches an error map of the absolute difference between rendered and measured pixels, rays are drawn preferentially where that error is large, a uniform term keeps part of the budget flowing to well-rendered regions, and the maps are refreshed on a fixed cyclic schedule. Unlike gradient-guided selection, which scores each pixel once by its Sobel image gradient, the residual rule follows the model's error as it moves rather than repeatedly sending rays to edges that are already reconstructed. In the ablation, residual-guided selection improves 3D SSIM in every configuration, with its largest effect on the unconstrained hash grid that has neither a prior nor error feedback.
Across four datasets (Ellipsoids, Shells, Lung CT, Voids), OX-NeRF renders the most faithful held-out projections in eleven of the twelve settings in Table 2, achieves the highest 3D SSIM in every setting and the highest 3D PSNR in all but one. Prior radiance field methods adapted to X-ray largely fit one scene at a time and had not been evaluated in the ultra-sparse domain; OX-NeRF is compared directly against existing self-supervised methods at every projection count from four to ten. All methods are scored with the same metric definitions and a single least-squares gain correction, ONIX and OX-NeRF both train on 50 scenes per dataset and receive the same four source projections, and the other methods use their published default settings.
The advantage narrows as views increase: on Lung CT, OX-NeRF still leads R2-Gaussian at 15 views, the two are on par at 20, and by 25 views R2-Gaussian has clearly overtaken. This delimits the method's operating range: the value of the angle-gap prior supplied by the convolutional encoder falls as the gaps between angles close, while R2-Gaussian spends its compute on rendering more rays, which becomes the better strategy as views are added. Table 4 reports 3D PSNR and 3D SSIM for both methods at 15, 20 and 25 views.
Perspective
The result targets the ultra-sparse regime of ten projections or fewer and applies to a set of related scenes recorded under a common acquisition setting, with one renderer serving both parallel-beam and cone-beam geometry. The method's value lies in a shared prior compensating for missing angular coverage, so it suits settings where angular gaps are large and a single scene's projections cannot determine the solution. The authors note that changing the input from a set of scenes to a temporally ordered sequence requires nothing else in the formulation to change, which gives a natural extension path to X-ray multi-projection imaging at synchrotrons and free-electron lasers, where a handful of projections are recorded per frame at kilohertz to megahertz frame rates; a sequence could also supply temporal regularisation, since geometry ambiguous in one frame could be resolved in another as objects move. The ablation shows both scene count and hash table size raise reconstruction quality and both were capped by memory, so whether gains continue on more capable hardware is an open question, and the shared components could in principle host domain-specific priors such as known anatomical structure.
Several points remain worth watching. Some hyperparameter values in Table 1 and Table 3 are not fully given in the body text, so reproduction requires the supplementary material. The ablation shows that 50 scenes and the default hash table size are not the strongest settings but the largest the memory trade-off allows, with Lung CT the limiting factor and training already peaking above 90 GB of VRAM, so the quality ceiling is tied to hardware scale. ONIX and CombiNeRF were modified before comparison, the former extended to cone-beam geometry and the latter given an X-ray transmission renderer, and the effect of those changes on their results is not quantified in the body text. The two tasks also diverge: at ten views most methods already render high and closely clustered 2D SSIM scores while the 3D reconstruction gap remains clear, indicating that rendering fidelity is not a substitute for volumetric correctness.
