From Alignment to Synthesis: Contrastive Volumetric Grounding for Text-to-CT Generation
Synopsis
The work proposes a generation-oriented 3D-CLIP encoder trained with structured hard negatives constructed exclusively at the text level (Attribute-Aware Negatives, AAN, and Semantic-Aware Negatives, SAN) to strengthen contrastive learning under the small-batch constraints of volumetric encoders, and uses it to condition a fully end-to-end latent diffusion model operating directly in 3D latent space, achieving lower FID, higher pathology-classification AUC and precision, faster inference, and lower GPU memory than competing methods on CT-RATE across 18 pathological conditions.
Figure 1: Overview of the proposed Text-to-CT generation framework: (a) 3D-CLIP Training with Generation-Oriented Negatives : A contrastive learning setup aligns CT scans and radiology reports into a shared embedding space. To overcome the limitations of standard in-batch negatives under small batch sizes, training is augmented with two complementary miners: Attribute-Aware Negatives (AAN), which select reports sharing the same pathology but differing in clinical attributes such as laterality, severity, or affected lobe, and Semantic-Aware Negatives (SAN), which retrieve the most semantically confusable reports via nearest-neighbor search in a frozen text embedding space. (b) Diffusion UNet Training : A latent diffusion model is trained to denoise compressed CT representations, conditioned on textual embeddings via cross-attention. A pretrained VAE encoder compresses the volumes into latent vectors, which are noised and passed through a 3D U-Net. (c) Inference : A synthetic latent code is generated from noise using the textual prompt, then decoded into a high-resolution CT volume via the VAE decoder.
arXivInterpretation
A generation-oriented 3D-CLIP encoder is introduced that uses two text-level hard negative miners, AAN and SAN, to improve fine-grained semantic disambiguation under small batch sizes. Existing Text-to-CT methods condition generation on text encoders pretrained with language-only or 2D vision-language objectives, leaving the conditioning signal volumetrically blind; this work explicitly grounds the conditioning signal in 3D visual representations while keeping hard negatives text-only so no additional 3D memory is required. Trained on CT-RATE with the same contrastive objective and data as CT-CLIP, differing only by adding AAN and SAN, making it the most direct comparison; AAN selects reports differing in laterality, severity, or lobar localization under a label-overlap constraint, and SAN retrieves nearest neighbors in a frozen BioClinicalBERT text space while excluding true positives; ablations show each miner improves over the baseline and their combination performs best.
A fully end-to-end text-conditioned latent diffusion framework operating directly in 3D latent space is presented, avoiding the spatial artifacts and cross-slice inconsistencies of super-resolution pipelines. GenerateCT and MedSyn use cascaded pipelines with low-resolution generation followed by super-resolution, which introduce spatial artifacts and cross-slice inconsistencies; this framework compresses volumes with a pretrained 3D VAE, denoises in latent space, and decodes to full-resolution CT with the VAE decoder. The VAE is kept frozen and reconstruction on the test set reaches PSNR of 38.87 dB and SSIM of 0.963; the method attains the lowest 2.5D and 3D FID (3D FID 0.015 versus 0.053 for Report2CT, 0.169 for MedSyn, and 0.166 for GenerateCT), with the fastest inference at 24.26 s per volume and the lowest peak GPU memory at 19.43 GB among compared methods.
Through an ablation that fixes the generative backbone and varies only the conditioning encoder, an empirical link is established between volumetric vision-language grounding quality and downstream generative controllability. Prior work assumes that a richer text encoder is sufficient; this work isolates the conditioning component and shows that as grounding increases from BioClinicalBERT to 3D-CLIP, FID and classification metrics generally improve monotonically, with the gap being more pronounced on classification than on FID. With the same generative model, BioClinicalBERT reaches AUC 0.704 and precision 0.393, BioMedCLIP 0.666/0.380, Merlin 0.722/0.443, CT-CLIP 0.753/0.471, and the proposed 3D-CLIP 0.787/0.535; the exception of BioMedCLIP scoring slightly below BioClinicalBERT is attributed to visual bias from its heterogeneous 2D biomedical image pretraining.
It is argued that image fidelity metrics alone are insufficient to evaluate semantic controllability in text-conditioned medical image generation, and that task-oriented metrics such as pathology classification are more sensitive. Prior evaluation relies heavily on FID; this work shows a model can maintain low FID while systematically ignoring the conditioning report, and offers task-oriented metrics as a more clinically meaningful proxy. As the guidance scale varies from 0 to 15, FID 3D remains low and relatively stable, whereas precision varies and peaks at w = 5.0; additionally, CT-Net achieves AUC (macro) of 0.821 on real test data, providing a reference point for classification scores on generated volumes.
Perspective
The results target research and engineering settings that generate chest CT volumes conditioned on radiology reports, within the chest CT data distribution and official split represented by CT-RATE; their value lies in providing a reusable training and evaluation paradigm for the design choice that conditioning-signal quality governs semantic controllability, and in making the encoder itself usable for clinical tasks that do not involve generation, such as zero-shot pathology classification and volumetric retrieval. Because the method depends on paired text-volume data, extending it to other anatomical regions or protocols presupposes the availability of corresponding paired datasets.
Evaluation is conducted exclusively on CT-RATE, so generalization to other body districts and scanning protocols remains to be verified; hard negatives operate only at the text level, leaving open whether volumetric hard negatives would add further gains; the reader study is preliminary in scale (25 volumes per method, one board-certified radiologist), so a larger, multi-reader formal evaluation covering a broader range of pathological presentations remains an open direction; moreover, this is a full-text parse, and the per-class breakdown, per-volume statistics, and downstream classifier training results in the supplementary material are not included in the loaded text, which affects a complete judgment of the robustness of the conclusions.
