Large Language Models for Clinical Note Simplification: A Systematic Review and Experimental Evaluation of Medical Text Readability
Synopsis
Combining a systematic literature review with an experimental evaluation, this study tested ten freely available large language models on five synthetic German clinical notes using standardized prompts, finding that all models substantially increased text length and consistently reduced the density of technical terms and abbreviations, yet no model achieved consistent improvements across all readability indices, with Mistral, ChatGPT, and Copilot showing the highest efficiency in balancing linguistic simplification and text length, suggesting that conventional readability metrics should be extended with domain-specific measures.
Interpretation
All evaluated large language models substantially increased the overall length of simplified text while consistently reducing the density of technical terms and abbreviations. Prior work on clinical text simplification has largely focused on English or general text; this study targets German clinical notes and reports both length change and two domain-specific complexity indicators, terminology and abbreviation density. Based on an experimental evaluation of ten freely available large language models, five synthetic clinical notes, and standardized prompts, with consistent results on terminology and abbreviation density.
No model achieved consistent improvements across all readability indices, highlighting limitations of traditional readability metrics in the medical domain. Rather than relying on a single readability score, the study used five indices—Flesch Reading Ease, Wiener Sachtextformel, LIX, SMOG, and Coleman–Liau—and complemented them with an analysis of medical terminology and abbreviation density. A combination of five readability indices and domain-specific density measures across ten models and five synthetic notes.
Mistral, ChatGPT, and Copilot demonstrated the highest efficiency in balancing linguistic simplification and text length. The study provides a relative ranking among ten freely available models on the trade-off between simplification effect and text expansion. An experimental comparison based on standardized prompts and synthetic clinical notes, representing relative performance among models.
Large language models show strong potential to enhance the accessibility of clinical documentation for patients, but their effectiveness depends on model selection, prompt design, and evaluation methodology. By combining a systematic literature review with an experimental evaluation, the study explicitly identifies model selection, prompt design, and evaluation methodology as key factors influencing simplification outcomes. A research design combining systematic review and experimental evaluation, with conclusions based on ten models and five synthetic notes.
Perspective
This work applies to the automatic simplification of German clinical notes and is intended for researchers, clinical informatics professionals, and healthcare system designers seeking to improve patients' understanding of their medical records. Its conclusions are based on ten freely available large language models, five synthetic clinical notes, and standardized prompts, so they primarily apply to evaluating model simplification performance under controlled conditions rather than directly generalizing to real clinical workflows or all language settings. The study suggests that conventional readability metrics should be extended with domain-specific measures, a direction that can be validated in future work on real patient records and across different healthcare systems.
The study uses synthetic clinical notes rather than real patient records; real notes may differ in terminology density, abbreviation conventions, and structure, so how simplification performs in real settings remains to be observed. The study does not report patient or clinician understanding and satisfaction with simplified text, leaving open the question of whether improved readability metrics translate into actual comprehensibility. In addition, model versions, prompt wording, and the choice of evaluation metrics may all influence results, and transferability across languages and healthcare systems requires further validation.
