Skip to main content
Back to timeline
arXivSource publication:

ΔRepresentation models phenotypes as representation increments over normal anatomy via counterfactual reasoning, improving lesion grounding and phenotype recognition across six medical VLMs

Synopsis

The work proposes ΔRepresentation, a visual phenotype representation learning framework for medical VLMs: BaseAnatomy applies geometry supervision to normal anatomical representations using ideal spatial distributions, while Phenotype uses counterfactual reasoning to estimate a normal anatomical reference for each pathological patch and learn phenotype-specific representation increments that cluster by phenotype. On ReXGroundingCT and LIDC-IDRI, it consistently improves lesion grounding and phenotype recognition across six pretrained medical VLMs.

Source-provided article image: $\Delta$Representation: Geometry Supervised Representation Learning of Phenotypes via Counterfactual Reasoning for Medical VLMs
Figure 1 ·

Figure 1: BaseAnatomy structures the anatomical visual representation space, while Δ \Delta Phenotype organizes lesion-specific increments relative to counterfactual normal anatomy.

arXiv

Interpretation

It defines a pathological phenotype as the visual change relative to its underlying normal anatomy and constructs that increment directly in the visual representation space rather than relying on semantic alignment. Prior fine-grained alignment methods (TRRG, PRIOR, MedKLIP, MCA-RG, RadAlign) use report sentences or pathology concepts as semantic supervision; text semantics describe the presence and subclass of a phenotype but do not explicitly represent the change relative to normal anatomy. This work instead defines the phenotype as the difference between a pathological patch and its normal counterpart in visual space. Supported by method argument plus experiments: across six backbones (RadFM, LLaVA-Med, Lingshu, Hulu-Med, MedGemma, MedGemma1.5) on ReXGroundingCT and LIDC-IDRI, average absolute VQA gains are 31.20% for grounding and 33.51% for phenotype, and average absolute RRG gains are 20.80% for grounding and 28.57% for phenotype.

BaseAnatomy uses annotation-defined ideal spatial distributions to provide geometric supervision that preserves both inter-structure geometry and intra-structure spatial continuity, establishing a stable normal anatomical reference space. Unlike alignment that relies on text-derived anatomical semantics, this module builds target distances directly from anatomical labels and normalized CT coordinates, optimizing with a smooth-L1 objective over matched pairwise cosine distances, plus rotation and scale consistency and balanced anatomical class sampling. High-dimensional diagnostics report Pearson correlations of 0.9742, 0.9751, and 0.9750 under rotation, scaling, and translation; in t-SNE, patches from different structures form more coherent clusters after training and vary gradually with position along the three CT axes; after downsampling, a single patch representation lies near the center of the original 16-patch distribution.

Phenotype constructs a counterfactual normal reference for each pathological patch and learns phenotype-specific increments that are continually refined as more lesion samples become available. Increments are estimated from anatomically matched mask-negative neighbors via normalized inverse-distance weighting (8–48 neighbors, grid radii 4–10); increments of the same phenotype are encouraged to be consistent while different phenotypes are separated, combined with localization supervision, anatomical preservation, and no-lesion suppression. Inferred counterfactual embeddings lie close to real normal representations with cosine distances of only 0.003 and 0.005 in two representative cases; normal-ring compactness is 0.9882 and mean/median abnormal-to-normal cosine distances are 0.0233/0.0210; phenotype increments reach a cosine silhouette of 0.3607 and nearest-prototype accuracy of 75.46%.

Phenotype representations improve continually as lesion-positive training data grows from 1% to 100%, indicating incremental learning refines rather than fixes the phenotype space. The work designs phenotype learning as a process that can incorporate new lesion samples while keeping the normal anatomical reference space frozen, progressively improving encoding and discrimination of pathological visual changes across anatomical contexts. Phenotype accuracy rises with data scale: ReXGroundingCT VQA from 25.80 to 43.11 and RRG from 15.55 to 52.65; LIDC-IDRI from 54.55 to 77.69 and RRG from 50.41 to 75.21; t-SNE shows the four phenotypes forming more compact and separable clusters as samples increase.

Perspective

The framework targets radiological interpretation where the unit of analysis is an axial chest CT slice; training requires anatomical annotations and lesion masks to build geometric supervision and counterfactual normal references, while inference receives only the unmarked CT slice, slice-position information, and a text prompt. It suits researchers and engineering teams who want to improve the vision encoder without replacing the language backbone, especially in VQA and report generation pipelines that need both lesion grounding and phenotype recognition. Because phenotype increments depend on the stable normal reference space provided by BaseAnatomy, benefits are most direct in anatomical regions with reliable annotation coverage.

Open questions a careful reader would still watch: the counterfactual normal reference is estimated from nearby mask-negative patches, and how its reliability changes near anatomical boundaries or for large lesions is not explored; phenotype increment separation is measured by a cosine silhouette of 0.3607 and nearest-prototype accuracy of 75.46%, which is moderate, so residual confusion among phenotypes deserves further observation; cosine silhouette and own-centroid cosine decrease modestly after spatial pooling, indicating representations are not strictly invariant to spatial aggregation, and the practical effect on downstream alignment needs more analysis; evaluation is limited to four pulmonary phenotypes in chest CT across two datasets, leaving cross-modality and cross-anatomy generalization open.

Sources