Skip to main content

Daily report

AI and science frontiers · 2025-06-01

Only content delivered through the publication boundary on this date is included.

arXiv

From Alignment to Synthesis: Contrastive Volumetric Grounding for Text-to-CT Generation

The work proposes a generation-oriented 3D-CLIP encoder trained with structured hard negatives constructed exclusively at the text level (Attribute-Aware Negatives, AAN, and Semantic-Aware Negatives, SAN) to strengthen contrastive learning under the small-batch constraints of volumetric encoders, and uses it to condition a fully end-to-end latent diffusion model operating directly in 3D latent space, achieving lower FID, higher pathology-classification AUC and precision, faster inference, and lower GPU memory than competing methods on CT-RATE across 18 pathological conditions.