Public articles linked to the same research event.
medRxiv This study constructed a multilingual clinical dictation corpus (five clinical dictation scripts spanning a complexity gradient, translated into 99 languages, rendered to synthetic speech under three acoustic conditions, and transcribed by a production ambient scribe), computed six frequency metrics, and had three independent large language model raters assess clinically meaningful error patterns using a Severity x Likelihood framework; across 59,819 genuine transcription-error occurrences, 58,329 (97.5%) were LOW risk and 251 (0.42%) CRITICAL or HIGH, none of the six frequency metrics showed a detectable association with serious clinical risk (absolute Spearman rho < 0.16), a Severity x Likelihood sum remained strongly correlated with WER (rho=0.
This study constructed a multilingual clinical dictation corpus (five clinical dictation scripts spanning a complexity gradient, translated into 99 languages, rendered to synthetic speech under three acoustic conditions, and transcribed by a production ambient scribe), computed six frequency metrics, and had three independent large language model raters assess clinically meaningful error patterns using a Severity x Likelihood framework; across 59,819 genuine transcription-error occurrences, 58,329 (97.5%) were LOW risk and 251 (0.42%) CRITICAL or HIGH, none of the six frequency metrics showed a detectable association with serious clinical risk (absolute Spearman rho < 0.16), a Severity x Likelihood sum remained strongly correlated with WER (rho=0.
This study constructed a multilingual clinical dictation corpus (five clinical dictation scripts spanning a complexity gradient, translated into 99 languages, rendered to synthetic speech under three acoustic conditions, and transcribed by a production ambient scribe), computed six frequency metrics, and had three independent large language model raters assess clinically meaningful error patterns using a Severity x Likelihood framework; across 59,819 genuine transcription-error occurrences, 58,329 (97.5%) were LOW risk and 251 (0.42%) CRITICAL or HIGH, none of the six frequency metrics showed a detectable association with serious clinical risk (absolute Spearman rho < 0.16), a Severity x Likelihood sum remained strongly correlated with WER (rho=0.
This study constructed a multilingual clinical dictation corpus (five clinical dictation scripts spanning a complexity gradient, translated into 99 languages, rendered to synthetic speech under three acoustic conditions, and transcribed by a production ambient scribe), computed six frequency metrics, and had three independent large language model raters assess clinically meaningful error patterns using a Severity x Likelihood framework; across 59,819 genuine transcription-error occurrences, 58,329 (97.5%) were LOW risk and 251 (0.42%) CRITICAL or HIGH, none of the six frequency metrics showed a detectable association with serious clinical risk (absolute Spearman rho < 0.16), a Severity x Likelihood sum remained strongly correlated with WER (rho=0.