Public articles linked to the same research event.
Studies in health technology and informatics Combining a systematic literature review with an experimental evaluation, this study tested ten freely available large language models on five synthetic German clinical notes using standardized prompts, finding that all models substantially increased text length and consistently reduced the density of technical terms and abbreviations, yet no model achieved consistent improvements across all readability indices, with Mistral, ChatGPT, and Copilot showing the highest efficiency in balancing linguistic simplification and text length, suggesting that conventional readability metrics should be extended with domain-specific measures.
Combining a systematic literature review with an experimental evaluation, this study tested ten freely available large language models on five synthetic German clinical notes using standardized prompts, finding that all models substantially increased text length and consistently reduced the density of technical terms and abbreviations, yet no model achieved consistent improvements across all readability indices, with Mistral, ChatGPT, and Copilot showing the highest efficiency in balancing linguistic simplification and text length, suggesting that conventional readability metrics should be extended with domain-specific measures.
Combining a systematic literature review with an experimental evaluation, this study tested ten freely available large language models on five synthetic German clinical notes using standardized prompts, finding that all models substantially increased text length and consistently reduced the density of technical terms and abbreviations, yet no model achieved consistent improvements across all readability indices, with Mistral, ChatGPT, and Copilot showing the highest efficiency in balancing linguistic simplification and text length, suggesting that conventional readability metrics should be extended with domain-specific measures.
Combining a systematic literature review with an experimental evaluation, this study tested ten freely available large language models on five synthetic German clinical notes using standardized prompts, finding that all models substantially increased text length and consistently reduced the density of technical terms and abbreviations, yet no model achieved consistent improvements across all readability indices, with Mistral, ChatGPT, and Copilot showing the highest efficiency in balancing linguistic simplification and text length, suggesting that conventional readability metrics should be extended with domain-specific measures.