Public articles linked to the same research event.
arXiv The work introduces ConRad, a reinforcement learning framework that fine-tunes medical large vision-language models with the GRPO algorithm and a logarithmic scoring-rule reward so they emit verbalized confidence alongside radiology reports at report or sentence level, reducing report-level ECE to 0.106 and sentence-level ECE to 0.239 on MIMIC-CXR and improving zero-shot IU-Xray ECE from 0.310 to 0.034 while keeping report quality (GREEN) stable.
The work introduces ConRad, a reinforcement learning framework that fine-tunes medical large vision-language models with the GRPO algorithm and a logarithmic scoring-rule reward so they emit verbalized confidence alongside radiology reports at report or sentence level, reducing report-level ECE to 0.106 and sentence-level ECE to 0.239 on MIMIC-CXR and improving zero-shot IU-Xray ECE from 0.310 to 0.034 while keeping report quality (GREEN) stable.
The work introduces ConRad, a reinforcement learning framework that fine-tunes medical large vision-language models with the GRPO algorithm and a logarithmic scoring-rule reward so they emit verbalized confidence alongside radiology reports at report or sentence level, reducing report-level ECE to 0.106 and sentence-level ECE to 0.239 on MIMIC-CXR and improving zero-shot IU-Xray ECE from 0.310 to 0.034 while keeping report quality (GREEN) stable.
The work introduces ConRad, a reinforcement learning framework that fine-tunes medical large vision-language models with the GRPO algorithm and a logarithmic scoring-rule reward so they emit verbalized confidence alongside radiology reports at report or sentence level, reducing report-level ECE to 0.106 and sentence-level ECE to 0.239 on MIMIC-CXR and improving zero-shot IU-Xray ECE from 0.310 to 0.034 while keeping report quality (GREEN) stable.