Skip to main content
Back to timeline
arXivSource publication:

TotalFM, an organ-separated 3D-CT foundation model, beats Merlin on 83% (25/30) of finding categories in zero-shot lesion classification while training at 32 batch/GPU

Synopsis

The study introduces TotalFM, an organ-separated 3D-CT radiology foundation model that uses TotalSegmentator and LLMs to automatically build roughly 340,000 organ-level volume-text pairs, then combines VideoMAE self-supervised pre-training with organ-wise contrastive learning; it reaches an average F1 of 0.708 in zero-shot organ-wise lesion classification (versus 0.515 for CT-CLIP and 0.650 for Merlin), a higher AUROC than Merlin in 83% (25/30) of finding categories, and report-generation performance comparable to Merlin while raising batch efficiency to 32 batch/GPU.

Source-provided article image: TotalFM: An Organ-Separated 3D-CT Foundation Model Leveraging Large-Scale Routine Clinical Radiology Data

Interpretation

It proposes an organ-separated contrastive learning framework that splits whole-body CT volumes into 192x192x32 organ patches and aligns them with corresponding report sentences in a shared embedding space, achieving organ-level volume-text representation learning. Prior 3D foundation models mostly contrast whole reports against whole volumes; this work refines the granularity to specific finding-sentence-to-organ-volume pairs and explicitly enforces organ-level semantic correspondence. In zero-shot organ-wise lesion classification, TotalFM reaches an average F1 of 0.708 versus 0.515 for CT-CLIP and 0.650 for Merlin, an average improvement of 37.5% over CT-CLIP and 6.5% over Merlin, with F1 above 0.6 for all evaluated organ labels.

The organ-separated, patch-based design substantially reduces computational cost, enabling large-batch contrastive learning on high-resolution 3D-CT. The paper's comparison shows CT-CLIP at 2 batch/GPU, RadFM at 1 batch/GPU, and Merlin at 18 batch/GPU, whereas TotalFM reaches 32 batch/GPU on H100 with an input resolution of 192x192x32 and about 180 GFLOPs. The batch-efficiency figures come from the authors' preliminary experiments in their own environment (H100, bf16, controlled organ patch counts) and are presented alongside related work in Table 1.

It builds a fully automated large-scale data pipeline integrating TotalSegmentator and LLMs to generate roughly 340,000 organ-level volume-text pairs from 140,000 CT series. The pipeline has five steps (report splitting, optimal series extraction, region/phase classification, finding-series matching, organ extraction) and adds about 110,000 rule-based negative pairs to learn normal anatomy. The training set contains 224,117 original-report-sentence pairs and 107,035 rule-based negative pairs; the region classifier achieves AUROC of 0.9974, 0.9654, 0.9708, 0.9897, and 0.9839 for head, neck, chest, abdomen, and pelvis.

In VLM fine-tuning it trains the Q-Former and LoRA fine-tuning of the LLM simultaneously and adds organ-specific embeddings to directly learn the connection to the LLM. Unlike the two-stage BLIP-2 scheme that first learns visual representations and then aligns to language space, this work trains the Q-Former and LoRA together to mitigate the modality gap, and generates reports by inputting only the specific organs identifiable from the target region rather than the whole CT volume. In the per-organ quantitative report-generation evaluation, TotalFM outperforms Merlin across all metrics for the gallbladder and pelvic organs and is generally comparable elsewhere; average BLEU is 0.164 (Merlin 0.151) and ROUGE-2 is 0.460 (Merlin 0.456).

Perspective

The framework targets organ-level 3D-CT analysis settings: teams that want to train high-resolution volumetric foundation models with limited GPU resources, and clinical research environments needing organ-level lesion screening, image-text retrieval, or a base for downstream CAD fine-tuning. The paper notes that running inference for each organ present in a CT scan enables automated diagnostic systems for the whole volume, for example an image-text retrieval system where text embeddings of finding sentences serve as a searchable index. It also presents the organ-wise classification task using real report sentences as a form of binary visual question answering that helps suppress performance fluctuations from prompt engineering and may serve as a candidate for standardizing 3D-CT foundation model evaluation.

Several open questions remain for a careful reader. First, findings that do not correspond to TotalSegmentator anatomical labels were excluded, with vascular (excluding the aorta) and lymphatic findings largely omitted from training data; future work needs open-vocabulary segmentation or visual grounding to localize regions corresponding to finding sentences. Second, the InfoNCE loss treats all non-paired elements in a mini-batch as negatives, which may introduce noise since some non-paired combinations might actually be positive findings. Third, the study selects a single representative CT series, whereas clinical diagnosis often requires comparing multiple series to observe temporal changes such as contrast enhancement dynamics. Fourth, the work focuses on individual findings, and integrating discrete findings into a holistic diagnosis remains future work. Fifth, in finding-wise classification, organs elongated along the z-axis or split into multiple labels, such as the aorta and lung fields, performed comparably to or slightly below Merlin, which the authors attribute to averaging similarities across multiple image embeddings under a sliding window diluting localized lesion signal. Sixth, the human evaluation of generated reports was performed by a single radiologist with over 10 years of experience as a qualitative example analysis. Seventh, the data are not publicly available due to patient privacy and ethical restrictions, and code and weights are planned for release.

Sources