Skip to main content
Back to timeline
arXivSource publication:

Retrieval-augmented anatomical guidance lets text-to-CT generation beat text-only baselines on fidelity, clinical consistency, and spatial controllability at once

Synopsis

The work proposes a retrieval-augmented text-to-CT generation method: given a radiology report, a pretrained 3D vision-language encoder retrieves the semantically nearest case from a reference corpus, and that case's anatomical annotation is used as a structural proxy injected into a text-conditioned latent diffusion model through a ControlNet branch; on the CT-RATE dataset (27,514 training and 1,818 test volumes), retrieval-augmented generation improves image fidelity (FID 3D 0.004), clinical consistency (CT-Net AUC 0.787), and spatial controllability (Dice 0.772, HD95 3.072) over text-only baselines, while additionally enabling explicit spatial controllability that such approaches inherently lack.

Source-provided article image: Retrieval-Augmented Anatomical Guidance for Text-to-CT Generation
Fig. 1

Fig. 1: The proposed retrieval-augmented Text-to-CT generation framework.

· Page 3

Interpretation

Anatomical structure is reformulated as a retrievable latent proxy rather than a direct conditioning input required at inference. Prior structure-driven methods such as MAISI require ground-truth segmentation masks at inference, while text-only methods lack spatial constraints; this work uses the annotation of a retrieved related case as a coarse spatial scaffold, so no target-volume annotation is needed at inference. The method is implemented on CT-RATE, with retrieval performed exclusively over the training set; test-set reports are never included in the retrieval index and test-set masks are never provided as conditioning input, which the authors state ensures no information leakage at evaluation time.

The retrieved anatomical proxy is injected into a frozen text-conditioned latent diffusion model via a ControlNet branch, enforcing global anatomical consistency while preserving report-driven semantic variability. The diffusion backbone and vision-language encoder are kept frozen, and only the control branch and zero-initialized projection layers are optimized; zero-initialization makes the residuals approximately zero at the start of training, recovering the original pretrained generator. Control features are injected at all downsampling and bottleneck layers, the conditioning scale is fixed to 1.0 (selected via a sweep over [0.5, 2.0]), text conditioning uses classifier-free guidance with a scale of 5.0, sampling uses rectified flow with 30 steps, and the ControlNet branch is trained for 50 epochs.

Retrieval-augmented generation outperforms text-only baselines across three complementary evaluation axes and additionally gains explicit spatial controllability. Text-conditioned methods inherently lack spatial controllability; this work reaches FID 3D of 0.004 (versus 0.015 for Text-to-CT and 0.012 for MAISI), CT-Net AUC of 0.787 (versus 0.745 for Text-to-CT), Dice of 0.772, and HD95 of 3.072. Results use the official CT-RATE split, and statistical significance is assessed with a paired Wilcoxon signed-rank test across 5 independent runs with different random seeds (p < 0.05); CT-Net achieves an AUC of 0.824 on real CT volumes as a reference upper bound.

Retrieval quality systematically affects generation performance, with semantically aligned proxies yielding consistent gains. The ablation compares semantically nearest, semantically farthest, and random proxy selection, with all configurations sharing the same generative backbone and differing only in the proxy selection criterion at inference time. Semantically nearest retrieval performs strongest on clinical consistency and spatial controllability, while random and farthest retrieval degrade performance; the authors note that retrieval quality has a modest effect on low-level appearance statistics and a more prominent role in higher-level semantic and anatomical alignment.

Perspective

The result targets report-conditioned chest CT synthesis, in settings where the reference corpus already contains cases with anatomical annotations and retrieval can be performed within the training set; it offers a path to anatomical guidance without inference-time annotations for data augmentation, simulation, and privacy-aware learning, and could extend to other anatomical regions or modalities by swapping the reference corpus and annotation type. The authors state that future work will investigate pathology-specific evaluation and longitudinal scenarios, leveraging temporally related priors to model disease progression.

The retrieved proxy is explicitly defined as a coarse spatial scaffold rather than a precise template of the target anatomy, so spatial controllability measures adherence to the provided scaffold rather than correctness with respect to unknown ground truth; the authors also note that a model trivially copying the proxy would score high on spatial metrics yet degrade on clinical consistency, indicating the three axes must be read complementarily. The role of retrieval quality is modest for low-level appearance statistics and more prominent for higher-level semantic and anatomical alignment, and the boundary of this difference remains to be characterized further. In addition, validation is limited to chest CT and CT-RATE, with pathology-specific evaluation and longitudinal scenarios listed as future work, so it is not yet clear whether the conclusions hold for other anatomical regions or annotation schemes.

Sources