CARE's contrastive multi-agent adjudication lifts zero-shot melanoma-vs-atypical-nevus accuracy from 66.5% to 77.6%, but still trails Gemini-3-Pro on chest X-rays
Synopsis
In a zero-shot, training-free, tool-free setting, the authors benchmark multimodal LLM agents on two imaging-only proxy tasks (melanoma vs. atypical nevus and pulmonary edema vs. pneumonia) and propose CARE, a multi-agent framework in which two disease-specific agents generate opposing evidence and a third judge adjudicates it against the original image; CARE raises Gemini-3-Flash accuracy from 66.5% to 77.6% (Youden 0.552) on dermoscopy and from 60.2% to 64.6% on chest X-rays, yet remains below Gemini-3-Pro's 70.9% on the chest task and overall below clinical deployment requirements.
Interpretation
The authors propose CARE (Contrastive Agent REasoning), a training-free, tool-free multi-agent framework: two disease-specific agents each enumerate visual evidence for only their assigned hypothesis and are barred from making a final diagnosis, while a third judge agent receives the original image plus both evidence sets, cross-checks claims against the image, flags unsupported or contradictory claims, and weighs the remaining arguments to reach a diagnosis. Unlike prior remedies that rely on extra annotated data for fine-tuning, repeated sampling for consistency, tool-use collaboration, or training recipes against overconfidence, CARE structures disagreement explicitly through prompting alone and operates in a zero-shot setting. The method is fully specified in Section 2 with the workflow shown in Fig. 2; the authors note CARE is implemented via structured prompting only and does not compute explicit likelihoods or numerical functions.
On the dermoscopy task (melanoma vs. atypical nevus), CARE built on Gemini-3-Flash reaches 77.6% accuracy, F1 0.769, and Youden 0.552, an improvement of over 11 percentage points over the single-agent Gemini-3-Flash baseline (66.5% accuracy, Youden 0.328). The gain holds under compute-matched controls: Self-Check (3x) and Majority-Vote (3x) also invoke the API three times per case but show only limited improvement, indicating the benefit comes from structured contrastive reasoning rather than extra sampling or ensembling. The dataset comprises 509 studies curated from derm7pt (257 atypical nevi, 252 melanomas), used test-only; the difference versus Gemini-3-Flash is significant at p < 0.0001 by McNemar and permutation tests.
On the chest X-ray task (edema vs. pneumonia), CARE reaches 64.6% accuracy, F1 0.619, and Youden 0.287, outperforming its base model Gemini-3-Flash (60.2%) but falling short of Gemini-3-Pro's 70.9%. This shows the contrastive-adjudication gain points in the same direction across both modalities but differs in magnitude; the authors also report that CARE's difference from Gemini-3-Pro on the dermoscopy task is not statistically significant (McNemar and permutation p = 0.192). The dataset comprises 1,739 studies from MIMIC-CXR after XOR filtering and exclusion of low-confidence statements (878 edema, 861 pneumonia); CARE's improvement over Gemini-3-Flash is significant at p < 0.001 on all three metrics.
Ablations show that a judge agent with access only to the two specialists' textual arguments (Blind-CARE) reaches 73.9% accuracy on melanoma, better than both Self-Check variants but worse than CARE; in qualitative cases CARE flags contradictory findings, recalibrates cross-agent evidence weighting, and uses multi-view checking to reject the pneumonia agent's localized consolidation claim. This supports, at both experimental and case level, that direct access to visual evidence is essential for detecting fabricated or misinterpreted claims, and suggests errors in visually confounded cases may be correlated rather than random, limiting the corrective effect of simple majority aggregation. Ablations cover Self-Check (2x/3x), Majority-Vote (3x), and Blind-CARE, compared with CARE under matched compute; qualitative evidence comes from representative cases in Fig. 3.
Perspective
The work targets controlled zero-shot, imaging-only, forced mutually exclusive binary proxy tasks, and is meant for researchers and engineering teams who want to evaluate or design medical multi-agent reasoning without fine-tuning or external tools; the authors suggest future work could examine whether segmentation models or image retrieval improve performance, and call for further methodological advances and more rigorous evaluation.
Label quality carries uncertainty: part of the dermoscopy data was histologically verified while the rest relied on expert diagnosis, and chest X-ray labels were automatically extracted from radiology reports, reflecting report impressions rather than an independent reference standard such as CT or clinical adjudication; the mutually exclusive setting conflicts with patients potentially having both edema and pneumonia; agents were evaluated without external tools. In addition, some case text in Fig. 3 appears in truncated form in the source, so the qualitative evidence can only be read from the fragments shown.
