Skip to main content
Back to timeline
arXivSource publication:

SentZero pairs abstract-level sentence mapping with patch-level false-negative alignment to beat prior multi-task zero-shot methods on chest X-ray classification and grounding after MIMIC-CXR pretraining

Synopsis

SentZero is a sentence-centric vision-language pretraining framework for chest X-ray: it uses an LLM to extract phrases from radiology reports and map them to concise topic-presence sentences that expand positive-pair diversity, adds an auxiliary loss that attracts the highest-similarity image patches toward clinical sentences recurring across studies to mitigate false negatives, and applies sentence-conditioned residual modulation to visual embeddings; after pretraining on MIMIC-CXR it outperforms prior multi-task zero-shot methods on most metrics for zero-shot multi-label classification on Open-I, ChestXray14, PadChest, ChestXDet10 and CheXpert and for zero-shot grounding on ChestXDet10 and MS-CXR.

AI-generated editorial illustration: SentZero: An Enhanced Sentence-Centric Vision-Language Pretraining for Multi-Task Zero-Shot Chest X-Ray Analysis

Interpretation

SentZero proposes a sentence-centric vision-language pretraining framework that transfers directly to a variety of downstream tasks without any task-specific fine-tuning. Most prior CXR vision-language pretraining methods still require task-specific supervised fine-tuning after contrastive pretraining, whereas SentZero targets zero-shot multi-task transfer by design. The paper pretrains on MIMIC-CXR (377K chest X-ray images, 227K radiographic studies, 65,379 patients) and evaluates on Open-I, ChestXray14, PadChest, ChestXDet10 and CheXpert for classification and ChestXDet10 and MS-CXR for grounding; the authors state that evaluation images were excluded from both SentZero training and the pretraining data of the RAD-DINO checkpoint used.

Abstract-level sentence mapping converts detailed finding sentences into concise topic-presence statements (for example, 'There is mild opacity in the bilateral lung base' is simplified to 'There is opacity') and adds them as extra positive pairs, expanding positive-pair diversity. Earlier sentence-level methods used LLM-extracted clinical phrases but treated each extracted sentence as an independent caption of its own image, leaving the hierarchical semantic structure of radiology reports unexploited; SentZero generates supervision at two levels of granularity from the same phrases. Ablation shows that adding abstract-level mapping on top of multi-layer feature aggregation raises OpenI from 0.8741 to 0.8856, ChestXray14 from 0.8164 to 0.8323, ChestXDet10 classification from 0.8093 to 0.8395, and ChestXDet10 grounding from 0.6655 to 0.6811.

For false negatives caused by clinically equivalent sentences recurring across patients, SentZero leaves the contrastive pair assignments intact and instead uses an auxiliary loss to attract the top 20% most similar patches of each false-negative pair toward that sentence. Prior work such as CoNNs relabels or excludes cross-patient pairs using structured clinical concepts; the paper reports that directly masking or relabeling false-negative pairs underperforms no handling on most benchmarks, while selective patch-level attraction improves results while preserving the original contrastive objective and pair assignments. In Table 3, false-negative transition drops ChestXDet10 grounding to 0.5138 and MS-CXR to 0.7066, and masking also falls below the 'None' baseline on most datasets; the full model reaches best or near-best results at OpenI 0.8891, ChestXray14 0.8389, ChestXDet10 grounding 0.7315 and MS-CXR 0.9222.

Sentence-conditioned residual modulation uses FiLM to predict scale and shift from the sentence embedding and applies them to the attention-pooled visual feature, so the visual representation adapts to each input sentence. Earlier text conditioning was mostly applied to spatial attention; SentZero extends conditioning to the aggregated visual feature and keeps the original visual content through a residual connection. In the ablation, text conditioning alone scores higher than the full model on PadChest (0.8652) and CheXpert (0.9162), but the full model leads on the other five benchmarks by larger margins; Appendix Table 4 shows that replacing text features with duplicated aggregated patch features lowers performance on most datasets.

Perspective

The work targets research and engineering settings that use paired image-report data for chest X-ray vision-language pretraining: the pretraining corpus is MIMIC-CXR, evaluation covers multi-label classification on Open-I, ChestXray14, PadChest, ChestXDet10 and CheXpert and grounding on ChestXDet10 and MS-CXR, and the method transfers directly without task-specific fine-tuning. It suits teams that want natural-language prompts to drive classification and grounding and are willing to introduce LLMs for sentence structuring and mapping; abstract-level mapping relies on Qwen3-Next-80B and phrase extraction reuses RadZero's LLaMA3-70B-Instruct output, so reproducibility depends on those external models and prompt templates.

The reported results are zero-shot numbers on the listed public benchmarks and do not yet cover prospective clinical workflows or real deployment; abstract-level mapping and phrase extraction depend on external LLMs, and the main text does not systematically examine how output quality or prompt-template variation affects results; the appendix tables on the false-negative loss coefficient and patch sampling ratio show different settings trading off across individual datasets, so the best configuration may vary by task; and the main text presents results as tables and attention maps without per-class error analysis, so readers interested in failure modes for specific pathologies would need the appendix and the underlying data.

Sources