Public articles linked to the same research event.
arXiv The work proposes a retrieval-augmented text-to-CT generation method: given a radiology report, a pretrained 3D vision-language encoder retrieves the semantically nearest case from a reference corpus, and that case's anatomical annotation is used as a structural proxy injected into a text-conditioned latent diffusion model through a ControlNet branch; on the CT-RATE dataset (27,514 training and 1,818 test volumes), retrieval-augmented generation improves image fidelity (FID 3D 0.004), clinical consistency (CT-Net AUC 0.787), and spatial controllability (Dice 0.772, HD95 3.072) over text-only baselines, while additionally enabling explicit spatial controllability that such approaches inherently lack.
The work proposes a retrieval-augmented text-to-CT generation method: given a radiology report, a pretrained 3D vision-language encoder retrieves the semantically nearest case from a reference corpus, and that case's anatomical annotation is used as a structural proxy injected into a text-conditioned latent diffusion model through a ControlNet branch; on the CT-RATE dataset (27,514 training and 1,818 test volumes), retrieval-augmented generation improves image fidelity (FID 3D 0.004), clinical consistency (CT-Net AUC 0.787), and spatial controllability (Dice 0.772, HD95 3.072) over text-only baselines, while additionally enabling explicit spatial controllability that such approaches inherently lack.
The work proposes a retrieval-augmented text-to-CT generation method: given a radiology report, a pretrained 3D vision-language encoder retrieves the semantically nearest case from a reference corpus, and that case's anatomical annotation is used as a structural proxy injected into a text-conditioned latent diffusion model through a ControlNet branch; on the CT-RATE dataset (27,514 training and 1,818 test volumes), retrieval-augmented generation improves image fidelity (FID 3D 0.004), clinical consistency (CT-Net AUC 0.787), and spatial controllability (Dice 0.772, HD95 3.072) over text-only baselines, while additionally enabling explicit spatial controllability that such approaches inherently lack.
The work proposes a retrieval-augmented text-to-CT generation method: given a radiology report, a pretrained 3D vision-language encoder retrieves the semantically nearest case from a reference corpus, and that case's anatomical annotation is used as a structural proxy injected into a text-conditioned latent diffusion model through a ControlNet branch; on the CT-RATE dataset (27,514 training and 1,818 test volumes), retrieval-augmented generation improves image fidelity (FID 3D 0.004), clinical consistency (CT-Net AUC 0.787), and spatial controllability (Dice 0.772, HD95 3.072) over text-only baselines, while additionally enabling explicit spatial controllability that such approaches inherently lack.