Skip to main content
Back to timeline
arXivSource publication:

K-MaT aligns prompt manifolds via optimal transport to transfer medical VLMs to low-end modalities without low-end training images

Synopsis

K-MaT is a prompt-learning framework that factorizes prompts, anchors them to clinical text descriptions, and aligns the low-end prompt manifold to the visually-grounded high-end space using Fused Gromov-Wasserstein optimal transport, transferring decision structures to low-end modalities without requiring low-end training images; across four cross-modal benchmarks including dermoscopy, mammography to ultrasound, and CT to chest X-ray it achieves state-of-the-art results, raising the average harmonic mean of accuracy to 44.1% from BiomedCoOp's 42.0% with macro-F1 of 36.2%, and on the challenging breast imaging task it mitigates the catastrophic forgetting seen in CoOp, which drops to 27.0% accuracy on the low-end.

AI-generated editorial illustration: K-MaT: Knowledge-Anchored Manifold Transport for Cross-Modal Prompt Learning in Medical Imaging

Interpretation

It proposes K-MaT, a prompt-learning framework that transfers decision structures to low-end modalities without requiring low-end training images, addressing the problem that large-scale biomedical vision-language models adapted on high-end imaging such as CT often fail to transfer to frontline low-end modalities such as radiography and collapse into modality-specific shortcuts. Unlike conventional adaptation that relies on low-end training images, K-MaT combines prompt factorization, clinical text anchoring, and Fused Gromov-Wasserstein optimal transport to align prompt manifolds, enabling zero-shot cross-modal deployment. The abstract reports evaluation on four cross-modal benchmarks, including dermoscopy, mammography to ultrasound, and CT to chest X-ray, with comparisons against BiomedCoOp and CoOp.

It achieves state-of-the-art results on four cross-modal benchmarks, with the average harmonic mean of accuracy reaching 44.1% versus BiomedCoOp's 42.0% and macro-F1 of 36.2%. It provides comparable quantitative gains in a cross-modal transfer setting rather than only single-modality adaptation performance. The abstract gives specific metric values and comparison baselines, though the visible text does not provide per-benchmark breakdowns or statistical testing details.

On the challenging breast imaging task it mitigates the catastrophic forgetting seen in standard methods such as CoOp, which drops to 27.0% accuracy on the low-end, while K-MaT preserves robust performance across modalities. It treats robustness as a core cross-modal transfer criterion rather than focusing only on average accuracy. The abstract uses CoOp's 27.0% low-end accuracy as a contrast to indicate that K-MaT avoids comparable degradation on that task.

Perspective

The work targets zero-shot deployment that transfers medical vision-language models adapted on high-end imaging to low-end modalities, suited to frontline imaging settings lacking low-end training images; the method relies on clinical text descriptions as anchors and aligns prompt manifolds via Fused Gromov-Wasserstein optimal transport. Evaluation covers four cross-modal benchmarks, including dermoscopy, mammography to ultrasound, and CT to chest X-ray.

The visible text is only the abstract plus page navigation, missing per-benchmark result tables, ablations, statistical significance, and details on low-end modality coverage, so how metrics such as 44.1% and 36.2% distribute across tasks, and the boundary conditions of the robust breast imaging result, remain open questions a reader would need the full text to confirm.

Sources