Skip to main content
Back to timeline
Heart & LungSource publication:

Multimodal Large Language Models in Prehospital ECG Triage for Emergent Catheterization Laboratory Activation: A Retrospective Comparative Analysis

Synopsis

This study retrospectively analyzed 615 ECGs from 270 emergency medical service patient encounters with concern for acute myocardial infarction, using cardiology activation of the STEMI pathway as the reference standard, and compared three multimodal large language models with an ECG machine algorithm, finding that Gemini had the highest sensitivity (95.3%) but extremely poor specificity (9.4%), ChatGPT and Claude showed moderate sensitivity (68.1% and 67.2%) with limited specificity (42.3% and 46.5%), while the ECG machine algorithm was more balanced with sensitivity of 67.7% and specificity of 64.2%, suggesting that general-purpose large language models are not appropriate for ECG-based catheterization laboratory activation decisions in time-sensitive emergency workflows.

AI-generated editorial illustration: Large language models fail to reliably predict emergent catheterization laboratory activation from prehospital electrocardiograms.

Interpretation

The study evaluated multimodal large language models on prehospital ECGs using real-world cardiology activation of the STEMI pathway for emergent angiography as the reference standard. Prior assessments of automated ECG interpretation often compare against diagnostic labels, whereas this study anchors the comparison to actual clinical activation decisions, bringing the evaluation closer to real emergency workflow use. Retrospective analysis of 615 ECGs from 270 emergency medical service patient encounters, with cardiology STEMI pathway activation as the reference standard, representing a real-world decision benchmark.

The three multimodal large language models showed a marked imbalance between sensitivity and specificity: Gemini had sensitivity of 95.3% (95% CI 91.7-97.3) but specificity of only 9.4%, while ChatGPT and Claude had sensitivity of 68.1% and 67.2% with specificity of 42.3% and 46.5%, respectively. These results quantify differences in false-positive tendency across models on this specific prehospital ECG task, adding to the performance profile of general-purpose models in clinical interpretation tasks. Sensitivity, specificity, and a 95% confidence interval for Gemini sensitivity are reported, with a defined sample size and three models evaluated.

The ECG machine algorithm achieved a more balanced performance, with sensitivity of 67.7% (95% CI 61.4-73.4) and specificity of 64.2%, higher specificity than all tested large language models. It provides a machine algorithm comparison on the same set of prehospital ECGs, allowing the large language model performance to be understood against an existing clinical tool. Direct comparison with the three large language models on the same dataset and reference standard, with a confidence interval reported for sensitivity.

The study concludes that although some models achieved high sensitivity, poor specificity resulted in excessive false-positive activation recommendations, and that general-purpose large language models are not appropriate for ECG-based catheterization laboratory activation decisions. It translates model performance differences into a judgment about suitability for emergency workflows, indicating that high sensitivity alone is insufficient to support deployment in time-sensitive settings. The conclusion is based on the calculated sensitivity, specificity, positive predictive value, negative predictive value, and overall accuracy, representing evidence at the level of a retrospective comparative analysis.

Perspective

The study delineates the applicable scope of general-purpose multimodal large language models in prehospital ECG interpretation: the evaluation setting is emergency medical service patient encounters with concern for acute myocardial infarction, the reference standard is cardiology activation of the STEMI pathway, and the comparison involves three large language models and an ECG machine algorithm. These results can help emergency and cardiology teams clarify that, when considering large language models for prehospital triage or activation recommendations, their current role should be limited to assistance rather than replacement of existing interpretation tools; for researchers wishing to further explore model performance under specific thresholds or human-machine collaborative workflows, this study provides comparable baseline data.

Readers should still watch: this is a retrospective analysis, and its conclusions await prospective validation in real-time clinical settings; whether different model versions, prompting approaches, or image input conditions affect interpretation results is not elaborated in the original text; additionally, because the loaded content is at the abstract level and lacks figures and full methodological details, the specific values for positive predictive value, negative predictive value, and overall accuracy, as well as the statistical comparison methods for differences among models, remain open questions that require consulting the original article for complete information.

Sources