Public articles linked to the same research event.
Studies in health technology and informatics This study proposes a reference-free evaluation framework that combines embedding-based transcript-summary semantic similarity, stability analysis across repeated generations, and structured human evaluation to benchmark LLM post-processing models for clinical transcripts produced by a locally deployed Whisper system, applying it to 26 German physician-patient conversations in which four LLMs generated 416 summaries, and finding that GPT-OSS-120B and MedGemma-27B achieved the highest similarity scores across most embedding models, that cosine similarities across repeated runs typically exceeded 0.90, and that embedding-based similarity signals only partially agreed with human judgments of summary quality.
This study proposes a reference-free evaluation framework that combines embedding-based transcript-summary semantic similarity, stability analysis across repeated generations, and structured human evaluation to benchmark LLM post-processing models for clinical transcripts produced by a locally deployed Whisper system, applying it to 26 German physician-patient conversations in which four LLMs generated 416 summaries, and finding that GPT-OSS-120B and MedGemma-27B achieved the highest similarity scores across most embedding models, that cosine similarities across repeated runs typically exceeded 0.90, and that embedding-based similarity signals only partially agreed with human judgments of summary quality.
This study proposes a reference-free evaluation framework that combines embedding-based transcript-summary semantic similarity, stability analysis across repeated generations, and structured human evaluation to benchmark LLM post-processing models for clinical transcripts produced by a locally deployed Whisper system, applying it to 26 German physician-patient conversations in which four LLMs generated 416 summaries, and finding that GPT-OSS-120B and MedGemma-27B achieved the highest similarity scores across most embedding models, that cosine similarities across repeated runs typically exceeded 0.90, and that embedding-based similarity signals only partially agreed with human judgments of summary quality.
This study proposes a reference-free evaluation framework that combines embedding-based transcript-summary semantic similarity, stability analysis across repeated generations, and structured human evaluation to benchmark LLM post-processing models for clinical transcripts produced by a locally deployed Whisper system, applying it to 26 German physician-patient conversations in which four LLMs generated 416 summaries, and finding that GPT-OSS-120B and MedGemma-27B achieved the highest similarity scores across most embedding models, that cosine similarities across repeated runs typically exceeded 0.90, and that embedding-based similarity signals only partially agreed with human judgments of summary quality.