Sparse autoencoders trained on 909,873 CT/MRI slices decompose medical foundation-model embeddings into language-describable concepts, recovering 87.8% of downstream performance with 10 features
Synopsis
Trained on 909,873 2D CT and MRI slices from the TotalSegmentator dataset, this study fits Matryoshka sparse autoencoders with BatchTopK sparsification on frozen embeddings from BiomedParse (biomedical) and DINOv3 (general-purpose) foundation models alongside a random-weight baseline, finding that sparse features reconstruct original embeddings with R2 up to 0.941, recover 87.8% of downstream performance with only 10 features (99.4% dimensionality reduction), preserve 97.7% of dense retrieval quality with five-feature fingerprints, correspond to monosemantic concepts expressible in language as verified by an independent LLM judge, and enable zero-shot language-driven image retrieval on a single clinical text query.
Interpretation
Sparse autoencoders can compress dense medical image embeddings into very few interpretable features while retaining semantic utility: BiomedParse's optimal configuration recovers 87.8% of dense ROC-AUC with 10 features and DINOv3 recovers 82.4%, with gains diminishing above N=10. Prior evidence for sparse autoencoders on imaging embeddings came from chest radiographs alone, a single modality and architecture with paired text supervision; this work extends that evidence to two modalities (CT and MRI), two architecturally distinct foundation models, and diverse anatomical regions. Systematic sweep over 96 configurations per model, with a test set of 14.1% of images withheld entirely from three institutions; dense baselines reach ROC-AUC 0.907 (BiomedParse) and 0.912 (DINOv3).
Reconstruction fidelity is not semantic utility: the random-weight BiomedParse baseline reaches only 0.606–0.651 AUC despite a comparable R2 range (0.587–0.915), while DINOv3's lower R2 (0.649–0.841) coexists with higher downstream AUC. This dissociation separates 'reconstructs well' from 'encodes meaningful representational structure,' showing a sparse code can faithfully reconstruct a random embedding space that encodes no semantic structure. Directly supported by the random-weight baseline control condition compared against both trained models under the same evaluation pipeline.
Sparse features correspond to monosemantic concepts expressible in language and emerge from self-supervised training without explicit anatomical labels: an independent VLM judge ranked DINOv3's true concept description first for 170/250 features (mean rank 1.60) versus 141/250 for BiomedParse (mean rank 1.91). Concept descriptions are generated automatically by a VLM from activating images and metadata and then discriminated by a separate VLM among five candidates (one true, four distractors), rather than relying on human annotation or paired text supervision. 250 most monosemantic features evaluated per model via LLM-as-judge ranking; monosemanticity scoring itself relies on metadata-derived organ labels and VLM-generated descriptions, making it proxy-based evidence.
Sparse feature concepts can bridge clinical language and abstract latent representations: for the query 'Axial CT of the abdomen and retroperitoneum in an elderly patient,' DINOv3 selected three anatomy- and modality-specific abdomen CT concepts and retrieved correct axial abdominal CT images, whereas BiomedParse lacked a modality-pure abdomen feature, selected mixed MRI/CT concepts, and retrieved thoracic images. This is a zero-shot, language-driven retrieval demonstration requiring no reference image or task-specific training, using automatically labeled sparse features for text-to-medical-image matching. A proof-of-concept demonstration on a single query; the authors explicitly note that aggregate evaluation across a broader query set remains for future work.
Perspective
The result is aimed at researchers and engineering teams who want to inspect the internals of frozen medical imaging foundation models without architectural modification, retraining, or task-specific labels, in the setting of 2D CT and MRI slices of normal anatomy across 10 institutions. It enables follow-on work on concept-level retrieval, feature-level auditing, and language interfaces, extending prior sparse-autoencoder evidence from chest radiographs to multi-modal volumetric imaging and two architecturally distinct foundation models.
Monosemanticity scoring relies on metadata-derived organ labels and VLM-generated concept descriptions rather than human annotation, providing scalable but proxy-based evidence; TotalSegmentator covers normal anatomy across two modalities and excludes pathological cases, and analysis operates at the 2D slice level rather than volumetrically; language-driven retrieval is demonstrated on a single query, and aggregate evaluation across a broader query set remains for future work; finer-grained constraints such as demographics remain an open direction. In addition, this is a fast parse of the text, so the specific image content of Figures 2, 3, and 4 and some table details cannot be fully reconstructed from the text; readers needing to verify configuration curves and retrieval examples should consult the original figures.
