Skip to main content
Back to timeline
arXivSource publication:

Open-PMC-18M builds 18M medical image-text pairs via subfigure splitting and context summaries, lifting average retrieval by 27%

Synopsis

Starting from the BIOMEDICA corpus, the authors used DAB-DETR-based subfigure detection, Qwen2.5-VL-32B-Instruct subcaption extraction, and Qwen2.5-14B-Instruct summarization of inline context to curate and release OPEN-PMC-18M, a dataset of 18 million subfigure-text pairs, then trained vision-language encoders evaluated on 6 retrieval and 19 zero-shot classification tasks across radiology, microscopy, and visible light photography, reporting an average recall of 21.64, a 27% relative improvement over the previous best model OPEN-PMC, and an overall zero-shot average F1 of 39.77.

Source-provided article image: Open-PMC-18M: A High-Fidelity Large Scale Medical Dataset for Multimodal Representation Learning
Fig. S2

Supplementary Material, Fig. S2). Amplification of the der(1) junction fragment by suppression PCR retrieved the 5q31.2 breakpoint, which was confirmed by FISH (Fig. <xref ref-type="fig" rid="ddv00401">1</xref>E; E;

· Page 1

Interpretation

The authors curated and released OPEN-PMC-18M: 18 million subfigure-text pairs, where each text entry combines a subcaption extracted verbatim from the full caption with a summary of inline context, spanning microscopy (about 73%), radiology (about 18%), visible light photography (about 8%), and other clinical imaging (about 1%). Prior PMC-derived datasets (PMC-OA, PMC-15M, BIOMEDICA) pair whole compound figures with full captions and do not jointly perform subfigure extraction and contextual summarization; the authors state no prior work combined both in a unified biomedical vision-language dataset. Dataset scale and modality shares are reported in the main text; subcaption extraction was manually reviewed on a random sample of 1,000 instances with 94% reported correct, and 13.28% of subcaptions are exact copies of their full captions.

The authors trained a DAB-DETR-based subfigure detector on 500,000 programmatically generated compound figures (validated on a 20,000-image holdout), reaching 98.58% mAP and 99.96% F1 on the synthetic validation set and 36.88% mAP and 73.55% F1 on ImageCLEF 2016. Prior comparable work (Lin et al.) trained a DETR on 2,069 manually annotated compound figures from MedICaT; the authors instead use large-scale synthetic data to obtain ground-truth bounding boxes, which they describe as the first such corpus in the biomedical domain. Compared against a MedICaT-trained DETR and a zero-shot prompted Qwen2.5-VL-32B-Instruct on the same evaluation sets, with higher values on both metrics; the gap between the synthetic validation set and the real ImageCLEF 2016 benchmark is also reported.

Encoders trained on OPEN-PMC-18M reach an average recall of 21.64, a 27% relative improvement over the previous best model OPEN-PMC, and rank first in 13 of 19 zero-shot classification tasks and second in 1, with an overall average F1 of 39.77. The authors retrained PMC-6M and PMC-OA under a unified architecture (PubMedBERT text encoder plus ImageNet-pretrained ViT-B/16) and training protocol, making dataset composition rather than model differences the main variable. Retrieval is reported as Recall@200 (with Recall@10 and Recall@50 in the appendix), classification is reported under both zero-shot and linear probing, and perturbation robustness is tested with the Wilcoxon signed-rank test, significant at p<0.01 on Quilt and DeepEyeNet.

An ablation shows that replacing full captions with subcaptions alone yields no retrieval gain and even a slight drop, while adding contextual summaries on top produces a substantial increase. This provides a direct controlled comparison supporting the idea that subcaptions are locally precise but too sparse, and that contextual summaries restore broader semantic and clinical grounding, an aspect largely unaddressed before. The ablation is restricted to the radiology subset of OPEN-PMC-18M (3.2 million image-text pairs), with all variants sharing the same encoder architecture, optimization settings, and contrastive objective, compared on MIMIC-CXR.

Perspective

The dataset targets radiology, microscopy, and visible light photography images mined from PubMed Central open-access literature and is intended for research settings that train retrieval and classification encoders with contrastive learning; the authors state the models are not intended for clinical deployment and note that adapting them to clinical applications without rigorous validation and attention to clinical safety poses serious risks. They also point future work toward integrating the encoders with language model decoders for generative tasks such as medical report generation and visual question answering, and toward evaluating factual consistency there.

The authors note that additional analysis is required to fully assess generalization, and caution that data sourced from open-access repositories may reflect biases tied to specific institutions, imaging protocols, or publication norms, potentially limiting generalizability to underrepresented populations or distinct clinical settings. Factual consistency of generative downstream systems and interpretability remain open challenges. In addition, some figures in the text (such as the robustness plot and embedding visualizations) are presented as images, so specific values require consulting the appendix tables.

Sources