Skip to main content
Back to timeline
arXivSource publication:

MedQ-Engine lifts an 8B medical image quality model past GPT-4o by over 13% and cuts the gap to human experts to 4.34% using 10K annotations

Synopsis

The work proposes MedQ-Engine, a closed-loop data engine that iterates through evaluate-explore-evolve phases: it clusters model failures on a development set into failure prototypes, uses the visual component of those prototypes as retrieval anchors over a roughly one-million-image pool spanning five modalities, and combines progressive human-in-the-loop annotation with entropy-guided routing and quality-assured fine-tuning, so that with only 10K annotations an 8B-parameter model surpasses GPT-4o by over 13% on medical image quality assessment and narrows the gap to human experts to 4.34%, with more than 4x sample efficiency over random sampling.

Source-provided article image: MedQ-Engine: A Closed-Loop Data Engine for Evolving MLLMs in Medical Image Quality Assessment
Fig. 1

Fig. 1. Overview of MedQ-Engine. Our closed-loop data engine iteratively im- proves MLLMs for Med-IQA via three phases: Evaluating (failure clustering on a dev set), Exploring (prototype-based retrieval and human-in-the-loop annotation), and Evolving (quality-assured fine-tuning with re-evaluation).

· Page 2

Interpretation

It proposes a closed-loop data engine for medical image quality assessment that turns data-driven error analysis into systematic model improvement. Prior approaches either produce modality-agnostic scalar scores or are confined to specific modalities; this work applies an iterative evaluate-explore-evolve loop to Med-IQA so that data collection adapts as model weaknesses evolve. The paper reports experiments and ablations across five medical imaging modalities and describes the full three-phase loop with an overview figure.

It replaces predefined error categories with a data-driven failure discovery mechanism: the model is evaluated multiple times on an independent development set, samples whose error rate exceeds a threshold gamma become failure cases, and agglomerative clustering with the silhouette criterion yields failure prototypes. Failure patterns are characterized directly from model behavior rather than imposed categories, and only the visual component of each prototype is used as a retrieval anchor against the unlabeled image pool. The method section formalizes the failure threshold, the clustering criterion, and the retrieval similarity threshold tau_sim; the ablation reports a 9.46% drop when all engine components are removed (random baseline).

Progressive human-in-the-loop annotation with entropy-guided routing substantially lowers expert cost: at cold start GPT-4o pre-annotates 2K samples that experts review, and in later iterations routing by signals such as trajectory-level entropy reduces human review to 18% of samples. Compared with full human review, the 63% GPT-4o accept rate at cold start cuts per-sample review time from 5.1 minutes to 0.5 minutes, and the overall expert cost drops by more than 5x. The paper provides annotation statistics including a 63% accept rate, 29% edit rate, 8% reject rate, and average review times of 0.5 and 5.1 minutes.

Data scaling analysis shows failure-driven sampling helps most on hard cases: base models do well on artifact-free images but lose over 30% accuracy on mild and severe degradation, where MedQ-Engine yields the largest gains, and after training modality-specific questions consistently outperform general questions. This indicates failure-driven data collection does target the most challenging scenarios, and that domain-specific training data helps models build specialized knowledge of modality-characteristic degradations. Based on scaling curves from 2K to 40K samples across two architectures (InternVL3-8B and Qwen2.5-VL-7B) and on comparisons between failure-driven and random sampling.

Perspective

The result targets medical image quality assessment specifically, in settings that need descriptive quality assessment with clinical reasoning and where expert annotations are scarce; the authors argue the evaluate-explore-evolve paradigm generalizes to other specialized domains with scarce expert annotation and non-uniform model weaknesses. The method relies on a roughly one-million-image pool spanning five modalities plus an available development set and human review step, so its gains are easiest to reproduce for teams with comparable data pools and expert resources.

The loaded text is a fast parse in which some tables and formulas are incomplete: details of the three routing paths in entropy-guided routing, the full form of the trajectory-level entropy equation, and parts of the ablation tables are not fully rendered, so those settings should be confirmed against the original figures and tables. The specific values for the number of failure prototypes, the similarity threshold, and the adaptive sampling weights are also not given in the text, so reproduction requires checking the original. Results are based on the paper's own development set and evaluation benchmark, and how they hold on other institutions' data distributions remains an open question.

Sources