Skip to main content
Back to timeline
Studies in health technology and informaticsSource publication:

Embedding-Based Evaluation and Benchmarking Framework for Optimizing LLM Post-Processing of Medical Transcriptions from a Local Whisper System

Synopsis

This study proposes a reference-free evaluation framework that combines embedding-based transcript-summary semantic similarity, stability analysis across repeated generations, and structured human evaluation to benchmark LLM post-processing models for clinical transcripts produced by a locally deployed Whisper system, applying it to 26 German physician-patient conversations in which four LLMs generated 416 summaries, and finding that GPT-OSS-120B and MedGemma-27B achieved the highest similarity scores across most embedding models, that cosine similarities across repeated runs typically exceeded 0.90, and that embedding-based similarity signals only partially agreed with human judgments of summary quality.

AI-generated editorial illustration: Embedding-Based Evaluation and Benchmarking Framework for Optimizing LLM Post-Processing of Medical Transcriptions from a Local Whisper System.

Interpretation

It proposes a reference-free evaluation framework for benchmarking the quality of LLM post-processing of clinical transcripts. Whereas prior evaluation often relies on reference summaries, this framework replaces them with a combination of embedding-based transcript-summary semantic similarity, stability analysis across repeated generations, and structured human evaluation. The framework was applied to 26 German physician-patient conversations, with four LLMs each generating four summaries for 416 summaries in total, and a subset of summaries assessed by three human raters using six predefined quality criteria.

GPT-OSS-120B and MedGemma-27B achieved the highest transcript-summary similarity scores across most embedding models. This provides a relative comparison basis for selecting local LLM post-processing models in privacy-sensitive settings. Based on embedding similarity scores over 416 summaries, with relative model rankings remaining largely consistent even though absolute similarity values varied across embedding models.

Stability across repeated generations was high, but higher sampling temperatures reduced semantic similarity. By incorporating stability analysis, the framework reflects not only single-output quality but also generation consistency. Stability analysis showed cosine similarities across repeated runs typically exceeding 0.90, while higher sampling temperatures reduced semantic similarity.

Embedding-based similarity signals only partially agreed with human judgments of summary quality. This indicates that embedding metrics can serve as useful signals but cannot fully replace human quality assessment. A subset of summaries was assessed by three human raters using six predefined quality criteria and compared with embedding-based similarity signals, showing partial agreement.

Perspective

The framework targets locally deployed speech-to-text pipelines in privacy-sensitive settings and is suited to scenarios that require relative comparison and optimization of LLM post-processing models without gold-standard reference summaries, such as clinical summary generation after local Whisper transcription. It supports systematic model benchmarking and relative model comparison, serving researchers and developers who wish to select post-processing models without relying on reference summaries.

Embedding similarity and human quality judgments only partially agree, so the extent to which embedding metrics can serve alone as quality signals remains to be observed; absolute similarity values differ across embedding models, and although relative rankings are largely consistent, how to interpret absolute thresholds across embedding models remains an open question; human evaluation covered only a subset of summaries, so how far its conclusions extend to all 416 summaries deserves attention; higher sampling temperatures reduce semantic similarity, suggesting that the relationship between decoding parameters and evaluation results merits continued examination across broader settings.

Sources