Skip to main content
Back to timeline
arXivSource publication:

A zero-shot CT super-resolution framework that generates 2D projections with an X-ray diffusion prior and learns signed residuals via negative-density 3D Gaussians, outperforming CuNeRF in PSNR/SSIM on UHRCT and MELA

Synopsis

The work proposes a zero-shot 3D CT super-resolution framework: it first trains a diffusion model on abundant 2D X-ray data and uses DDNM/DDNM+ to upsample low-resolution CT projections into high-resolution projection priors, then applies a new Negative Alpha Blending Gaussian Splatting (NAB-GS) that models positive and negative Gaussian densities to learn the signed residual between diffusion-generated HR projections and upsampled LR projections for HR volume reconstruction; on the two public datasets UHRCT and MELA it achieves higher PSNR and SSIM than trilinear, cubic, NeRF, and CuNeRF zero-shot methods, is competitive with the supervised ArSSR, runs in about 15 minutes per volume, and two domain experts judged the 4× results to have clinical potential while 8× still needs improvement.

Source-provided article image: Zero-Shot CT Super-Resolution Using Diffusion-Based 2D Projection Priors and Signed 3D Gaussians
Fig. 1

Fig. 1. pre-trained diffusion model with 2D X-ray data is employed within the DDNM to generate HR 2D CT projection images from LR counterparts. (b) 3D CT reconstruction via NAB-GS: Using both positive and negative density Gaussians, we model a signed residual field between diffusion-generated HR projections and LR counterparts. For HR volume generation, the learned residual field is added onto the upsampled LR volume.

· Page 3

Interpretation

It proposes a two-stage framework that reformulates zero-shot 3D CT super-resolution as a 3D reconstruction task driven by diffusion-based upsampled 2D projection priors: stage one trains an unconditional diffusion model on large-scale 2D X-ray data such as ChestX-ray14 and CheXpert and generates HR CT projections within DDNM/DDNM+; stage two reconstructs the HR volume with NAB-GS. Unlike supervised 3D SR that needs paired HR-LR volumes and unlike zero-shot methods such as CuNeRF that rely only on information internal to a single LR volume, this framework brings a generative prior obtainable from abundant 2D X-ray data into 3D reconstruction, easing both the paired-data scarcity and the limited-LR-information constraints. The paper presents the two-stage pipeline and the DDNM/DDNM+ formulations (Eqs. 1 and 2) and reports results on UHRCT and MELA; the 2D diffusion model is trained on 112,120 ChestX-ray14 images and 80,845 CheXpert images, and the zero-shot learning uses only test sets.

It introduces NAB-GS, which replaces softplus with PReLU to allow negative densities and removes the non-negativity filter (αi≥ε) from the rendering formulation, so that a signed residual field between diffusion-generated HR projections and the upsampled LR baseline can be learned; the learned residual is added back to the upsampled LR volume and clipped with max(0,·). Standard 3DGS enforces non-negative density through softplus and alpha blending and therefore cannot express the positive and negative values inherent in the residual; NAB-GS relaxes this constraint and notes that because the formulation is a linear accumulation the gradient ∂C/∂αi=1 regardless of the sign of αi, keeping optimization stable, while pruning is adapted to discard Gaussians only when |z|<10⁻⁵. Ablations show the method outperforms R2-GS in PSNR (4×: +0.53, 8×: +0.71) and SSIM (4×: +0.0118, 8×: +0.0116); among softplus, ReLU, sine, and PReLU, PReLU performs best, while softplus and ReLU fail to capture boundaries in overestimated regions and sine introduces grainy artifacts.

It provides quantitative and qualitative validation on two public datasets plus an expert evaluation: UHRCT 4× reaches PSNR 25.42/SSIM 0.8957 and 8× reaches 21.96/0.8172; MELA 4× reaches 34.17/0.9525 and 8× reaches 30.81/0.9115, all above trilinear, cubic, NeRF, and CuNeRF, and competitive with the supervised ArSSR. Relative to CuNeRF, the only publicly available zero-shot 3D CT SR baseline, SSIM improves by roughly +0.04 to +0.06, and the framework takes about 15 minutes per volume (10 minutes for diffusion, 5 minutes for NAB-GS), faster than CuNeRF at about 1 hour and ArSSR at about 25 minutes; qualitatively, 4× restores fine structures such as bone boundaries more precisely. Quantitative results come from UHRCT (20/10 train/test split) and MELA (60/20/20 split) with PSNR and SSIM; the expert evaluation involves only two domain experts and compares only MELA at 4× and 8×, concluding that 4× shows potential for real-world clinical use while 8× requires further improvement and that the current visual gain does not substantially contribute to overall utility.

Perspective

The framework targets zero-shot 3D CT super-resolution: the input is a single LR volume, 100 projections uniformly spaced between 0° and 180° are generated with TIGRE, the projections are upsampled by the diffusion prior, and NAB-GS reconstructs the volume; it applies to 4× and 8× upsampling and does not require paired HR-LR training volumes. The authors note that experts considered the 4× results to have potential for real-world clinical use while 8× still requires further improvement, and that future work will enhance inter-slice continuity and conduct evaluations on real-world clinical data.

The expert evaluation involves only two domain experts and compares only MELA at 4× and 8×, so its clinical-potential judgment is a preliminary opinion; the authors also note that the current visual gain does not substantially contribute to overall utility and list inter-slice consistency and real-world clinical data evaluation as future work. In addition, this summary is based on the full text, and the specific visual details of Figures 2 and 3 cannot be fully verified from the prose, so the strength of the qualitative conclusions rests on the descriptions given in the text.

Sources