Evaluating Large Language Models for Lay Summaries of Radiology Reports Using Tailored Prompting Strategies and Mixed-Method Assessment
Synopsis
Using 100 radiology reports from the BioNLP 2023 report summarization dataset, this study had five large language models generate patient-facing lay summaries under individually tailored prompting styles selected in pilot work (few-shot for GPT-4; generated knowledge for GPT-4o mini, Gemini 1.5 Pro and Gemini 1.5 Flash; zero-shot for Llama 3.1), then evaluated them through a mixed framework of two radiology fellows, two large reasoning models (Gemini 2.5 Pro and GPT-oss-120b), and readability metrics, finding that Gemini 1.5 Flash and Pro with generated knowledge ranked highest for actionable content, minimal-supervision usability and readability, that GPT-4 with few-shot achieved the highest human-rated accuracy (98%), and that expert and large-reasoning-model ratings aligned closely.
Interpretation
In the consolidated percentage-agreement ranking across human raters and large reasoning models, Gemini 1.5 Flash (generated knowledge) ranked first, Gemini 1.5 Pro (generated knowledge) second, GPT-4o mini (generated knowledge) and GPT-4 (few-shot) tied third, and Llama 3.1 (zero-shot) ranked last on every criterion. Prior work mostly compared models alone or prompts alone, largely with zero-shot prompting; this study ranks the model-plus-prompt combination as the real-world unit and tailors the prompt style to each model. Based on 100 reports and 500 summaries across five models, scored on a 5-point Likert scale by two human raters and two large reasoning models, with Friedman and post-hoc Nemenyi tests (e.g., Rater 1 P < 4.97 × 10⁻², Rater 2 P < 9.03 × 10⁻²¹).
On the 'No Supervision' criterion, Gemini 1.5 Pro (generated knowledge) received the highest agreement (71% among human raters; 61% combined human and large-reasoning-model), closely followed by Gemini 1.5 Flash (generated knowledge; 70% and 60%), while other model-prompt configurations fell below 30%. This criterion moves evaluation from text quality toward whether a summary could enter a clinical communication workflow directly, giving a quantified reference for workflow-level usability. Derived from blinded ratings by two radiology fellows with 7 and 8 years of experience, plus the same criteria applied by two large reasoning models across 500 summaries.
On readability, Gemini 1.5 Pro (generated knowledge) produced the most accessible summaries (Flesch Reading Ease 67.84, Flesch-Kincaid Grade Level 7.55), whereas GPT-4o mini (generated knowledge) was least accessible (FRE 53.69, FKGL 10.77); all models simplified substantially relative to the original reports, which required a Grade 12 or above reading level. The study places subjective ratings alongside objective readability metrics, showing that subjective ranking and readability ranking do not fully coincide, for example GPT-4 few-shot was rated most accurate by humans yet produced the lengthiest summaries. Five readability measures were used (FRE, FKGL, SMOG, ARI, Dale-Chall), with the Friedman test showing significant differences across models (P < 2.36 × 10⁻³³).
Human experts and large reasoning models agreed closely, with complete disagreement of only 0.96% for Gemini 2.5 Pro (LRM 1) and 3.4% for GPT-oss-120b (LRM 2) against human raters, and GPT-oss-120b did not preferentially favor summaries from its own GPT family. This is the first use of large reasoning models as evaluators in the radiology lay-summary task with a human-agreement comparison, offering preliminary evidence for scalable second-pass review. Based on confusion matrices and agreement statistics over 2500 Likert responses, with the two large reasoning models drawn from different vendors to reduce self-preference bias.
Perspective
The work concerns methods for generating and evaluating lay summaries of radiology reports, aimed at researchers and clinical teams who want to bring large language models into patient communication; its prompt templates, mixed evaluation framework and pilot-study design can be reused with other model versions and clinical text settings, and the results are explicitly framed as a snapshot-in-time evaluation of specific model versions rather than a performance benchmark.
Readers should still watch that inter-rater agreement among human experts was low to fair and one rater used the Neutral option more often; that the small sample may not cover rare pathologies or very long reports; that readability metrics capture syntactic and syllable complexity rather than true patient comprehension; that the criteria did not directly assess harmful information, so whether flagged 'inaccurate' content was dangerous or benign remains unclear; that large-reasoning-model evaluators may share model biases and need external validation on larger cohorts; and that a single generation run means phrasing and readability could vary across repeated generations.
