Beyond word error rate: clinical risk as the necessary standard for ambient AI scribe evaluation: evidence from 77 global languages
Synopsis
This study constructed a multilingual clinical dictation corpus (five clinical dictation scripts spanning a complexity gradient, translated into 99 languages, rendered to synthetic speech under three acoustic conditions, and transcribed by a production ambient scribe), computed six frequency metrics, and had three independent large language model raters assess clinically meaningful error patterns using a Severity x Likelihood framework; across 59,819 genuine transcription-error occurrences, 58,329 (97.5%) were LOW risk and 251 (0.42%) CRITICAL or HIGH, none of the six frequency metrics showed a detectable association with serious clinical risk (absolute Spearman rho < 0.16), a Severity x Likelihood sum remained strongly correlated with WER (rho=0.
Interpretation
In a controlled multilingual corpus, aggregate frequency-based transcription metrics did not reliably track the sparse severe tail of clinically consequential errors. Prior ambient AI scribe evaluation relied mainly on frequency metrics such as WER; this study directly tested whether those metrics track consequential transcription errors and found no detectable association between six frequency metrics and serious clinical risk (absolute Spearman rho < 0.16). Based on 59,819 genuine transcription-error occurrences assessed by three independent large language model raters from external providers using a Severity x Likelihood framework informed by UK digital clinical-safety-risk-management principles.
A Severity x Likelihood sum remained strongly correlated with WER (rho=0.80), showing that the aggregate remained dominated by benign errors. This indicates that even a risk-weighted aggregate score still varies mainly with the large number of low-risk errors rather than the few critical or high-risk errors. Among 59,819 errors, 58,329 (97.5%) were LOW risk and only 251 (0.42%) were CRITICAL or HIGH, making the severe tail extremely sparse.
At complexity level 3, low-resource languages had worse WER than high-resource languages (beta=+0.078, 95% CI +0.045 to +0.111; p < 0.0001), without a detectable difference in CRITICAL/HIGH risk (OR 1.21, 95% CI 0.43 to 3.43; p=0.72). This suggests that language resource level affects transcription frequency quality but may not move clinical consequence risk in step, so the two need separate assessment. Based on stratified comparison at complexity level 3 in the multilingual corpus, reporting effect size, confidence interval, and p-value.
Consultation complexity was the principal predictor of serious risk (OR 3.06 per level, p < 0.0001). This shifts attention for drivers of serious risk from language or acoustic conditions toward the complexity of the consultation itself, offering a new focus for risk stratification. Analyzed across the complexity gradient in the controlled corpus, reporting odds ratio and p-value.
Perspective
This work applies to ambient AI scribe transcription evaluation in a controlled multilingual corpus, especially where transcription quality and clinical safety need to be distinguished; it is relevant to developers, evaluators, and clinical safety managers who wish to complement frequency metrics with context-aware assessment of error consequence.
The severe tail is extremely sparse (251 CRITICAL or HIGH, 0.42%), which may limit statistical power to detect associations with serious risk; results are based on synthetic speech and a single production system, so performance in real clinical settings remains to be observed; the degree of alignment between the three large language model raters' assessments and UK digital clinical-safety-risk-management principles, and inter-rater agreement, are questions a reader may continue to watch.
