Public articles linked to the same research event.
PLOS digital health Using 100 radiology reports from the BioNLP 2023 report summarization dataset, this study had five large language models generate patient-facing lay summaries under individually tailored prompting styles selected in pilot work (few-shot for GPT-4; generated knowledge for GPT-4o mini, Gemini 1.5 Pro and Gemini 1.5 Flash; zero-shot for Llama 3.1), then evaluated them through a mixed framework of two radiology fellows, two large reasoning models (Gemini 2.5 Pro and GPT-oss-120b), and readability metrics, finding that Gemini 1.5 Flash and Pro with generated knowledge ranked highest for actionable content, minimal-supervision usability and readability, that GPT-4 with few-shot achieved the highest human-rated accuracy (98%), and that expert and large-reasoning-model ratings aligned closely.
Using 100 radiology reports from the BioNLP 2023 report summarization dataset, this study had five large language models generate patient-facing lay summaries under individually tailored prompting styles selected in pilot work (few-shot for GPT-4; generated knowledge for GPT-4o mini, Gemini 1.5 Pro and Gemini 1.5 Flash; zero-shot for Llama 3.1), then evaluated them through a mixed framework of two radiology fellows, two large reasoning models (Gemini 2.5 Pro and GPT-oss-120b), and readability metrics, finding that Gemini 1.5 Flash and Pro with generated knowledge ranked highest for actionable content, minimal-supervision usability and readability, that GPT-4 with few-shot achieved the highest human-rated accuracy (98%), and that expert and large-reasoning-model ratings aligned closely.
Using 100 radiology reports from the BioNLP 2023 report summarization dataset, this study had five large language models generate patient-facing lay summaries under individually tailored prompting styles selected in pilot work (few-shot for GPT-4; generated knowledge for GPT-4o mini, Gemini 1.5 Pro and Gemini 1.5 Flash; zero-shot for Llama 3.1), then evaluated them through a mixed framework of two radiology fellows, two large reasoning models (Gemini 2.5 Pro and GPT-oss-120b), and readability metrics, finding that Gemini 1.5 Flash and Pro with generated knowledge ranked highest for actionable content, minimal-supervision usability and readability, that GPT-4 with few-shot achieved the highest human-rated accuracy (98%), and that expert and large-reasoning-model ratings aligned closely.
Using 100 radiology reports from the BioNLP 2023 report summarization dataset, this study had five large language models generate patient-facing lay summaries under individually tailored prompting styles selected in pilot work (few-shot for GPT-4; generated knowledge for GPT-4o mini, Gemini 1.5 Pro and Gemini 1.5 Flash; zero-shot for Llama 3.1), then evaluated them through a mixed framework of two radiology fellows, two large reasoning models (Gemini 2.5 Pro and GPT-oss-120b), and readability metrics, finding that Gemini 1.5 Flash and Pro with generated knowledge ranked highest for actionable content, minimal-supervision usability and readability, that GPT-4 with few-shot achieved the highest human-rated accuracy (98%), and that expert and large-reasoning-model ratings aligned closely.