Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

ConRad fine-tunes a medical vision-language model with GRPO and a logarithmic scoring rule, cutting report-level ECE from 0.642 to 0.106 and improving zero-shot IU-Xray ECE from 0.310 to 0.034

The work introduces ConRad, a reinforcement learning framework that fine-tunes medical large vision-language models with the GRPO algorithm and a logarithmic scoring-rule reward so they emit verbalized confidence alongside radiology reports at report or sentence level, reducing report-level ECE to 0.106 and sentence-level ECE to 0.239 on MIMIC-CXR and improving zero-shot IU-Xray ECE from 0.310 to 0.034 while keeping report quality (GREEN) stable.