Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

PLOS digital health

Evaluating Large Language Models for Lay Summaries of Radiology Reports Using Tailored Prompting Strategies and Mixed-Method Assessment

Using 100 radiology reports from the BioNLP 2023 report summarization dataset, this study had five large language models generate patient-facing lay summaries under individually tailored prompting styles selected in pilot work (few-shot for GPT-4; generated knowledge for GPT-4o mini, Gemini 1.5 Pro and Gemini 1.5 Flash; zero-shot for Llama 3.1), then evaluated them through a mixed framework of two radiology fellows, two large reasoning models (Gemini 2.5 Pro and GPT-oss-120b), and readability metrics, finding that Gemini 1.5 Flash and Pro with generated knowledge ranked highest for actionable content, minimal-supervision usability and readability, that GPT-4 with few-shot achieved the highest human-rated accuracy (98%), and that expert and large-reasoning-model ratings aligned closely.