Skip to main content
Back to timeline
arXivSource publication:

ConRad fine-tunes a medical vision-language model with GRPO and a logarithmic scoring rule, cutting report-level ECE from 0.642 to 0.106 and improving zero-shot IU-Xray ECE from 0.310 to 0.034

Synopsis

The work introduces ConRad, a reinforcement learning framework that fine-tunes medical large vision-language models with the GRPO algorithm and a logarithmic scoring-rule reward so they emit verbalized confidence alongside radiology reports at report or sentence level, reducing report-level ECE to 0.106 and sentence-level ECE to 0.239 on MIMIC-CXR and improving zero-shot IU-Xray ECE from 0.310 to 0.034 while keeping report quality (GREEN) stable.

Source-provided article image: Calibrated Confidence Expression for Radiology Report Generation

Interpretation

ConRad extends scoring-rule-based confidence calibration from text-only large language models to multimodal long-form radiology report generation, letting the model state verbalized confidence together with the report text. Prior work on uncertainty in radiology reporting largely reproduced radiologists' uncertainty phrasing as report language, or estimated uncertainty via stochastic sampling, semantic consistency, or post-hoc auditing; ConRad instead trains the model's own internal confidence calibration and supports both report-level and sentence-level granularity. On MIMIC-CXR the authors fine-tune only the language-model component of medgemma-4b-it with LoRA, training report-level confidence on 3000 samples for one epoch and sentence-level on 1500 samples for one epoch, and compare against Verbalize Base, Verbalize Supervised, Sequence Probability, P(True), Self-Consistency, and Trained Probe.

Report-level calibration improves markedly: ConRad reaches ECE 0.106, below Trained Probe (0.121), Verbalize Supervised (0.206), and Verbalize Base (0.642), while achieving the highest Pearson correlation of 0.431. The paper reports an ECE reduction of over 80% relative to the base model, a calibration curve closer to the ideal diagonal, and a shift from the base model's extreme overconfidence to a more balanced distribution; GREEN stays at 0.252, close to Verbalize Base's 0.257, indicating calibration training did not alter report generation quality. Table 1 reports 95% confidence intervals, e.g., ConRad ECE 0.106 [0.093, 0.118] versus Verbalize Base 0.642 [0.630, 0.654]; GREEN serves as a control to verify stable report quality.

Sentence-level confidence also improves: ConRad reaches ECE 0.239 and AUROC 0.633, better than Verbalize Base (0.438 and 0.558) and Verbalize Supervised (0.446 and 0.552). The paper reports roughly a 40% sentence-level ECE reduction relative to the base model and argues that with binary correctness targets the RL formulation still learns fine-grained confidence, whereas supervised fine-tuning can only reproduce the binary targets; it also notes sentence-level calibration remains more challenging than report-level. Table 2 reports 95% confidence intervals, e.g., ConRad ECE 0.239 [0.225, 0.253] and AUROC 0.633 [0.620, 0.647]; GREEN remains stable after training (0.262).

Filtering sentences by confidence raises factual precision but sacrifices completeness, so the paper argues low-confidence sentences should be flagged for human review rather than automatically removed. Table 3 shows Precision GREEN rising monotonically with the threshold for ConRad, from 0.421 on all sentences to 0.642 at confidence equal to 10, while the number of remaining sentences drops from 6317 to 53 and standard GREEN falls from 0.262 to 0.097; filtering with Verbalize Base confidences yields only marginal precision gains (0.425 to 0.447). The conclusion rests on sentence-level Precision GREEN statistics after filtering by predicted confidence on MIMIC-CXR, with the paper reporting both the precision and completeness sides of the trade-off.

Perspective

The work targets clinical AI assistance for chest X-ray reporting in the MIMIC-CXR style, using medgemma-4b-it as the base model with only the language-model component fine-tuned, and expressing confidence as an integer from 0 to 10 normalized to [0,1]. It enables report-level triage and sentence-level targeted review: report-level scores can prioritize human review, and sentence-level scores can flag statements needing closer inspection. The paper also demonstrates strict zero-shot transfer on IU-Xray and a small-scale clinical evaluation on 50 reports with three raters, indicating that calibration gains transfer to expert judgment.

Sentence-level calibration remains more challenging than report-level, and the paper notes sentence-level results are overall weaker. Confidence training depends on GREEN, an automated report-quality metric, as the correctness signal, so calibration quality is tied to that metric's reliability. The clinical evaluation is small (50 reports, three raters), so its conclusions are best read as an initial signal rather than broad clinical validation. Filtering by confidence creates a precision-completeness trade-off, and the paper recommends flagging rather than removing low-confidence content, but how to set thresholds in real workflows and how review costs change after flagging remain open questions. The method is also currently tied to chest X-ray reporting and a specific base model, so transfer to other imaging modalities, languages, and larger deployments needs further study.

Sources