Skip to main content
Back to timeline
PLOS digital healthSource publication:

A Medically Grounded LLM Agent-Based Tool to Detect Patient Safety Events in Medical Records

Synopsis

The study presents SAFE-AI, a framework that restricts a large language model to zero-shot entity extraction from emergency medical services charts and then makes determinations through a deterministic rule-based computational graph built from a clinician-defined ontology of clinical guidelines, reporting 97.9% accuracy for detecting epinephrine overdose and 91.6% for detecting delays in epinephrine administration across 300 pediatric out-of-hospital cardiac arrest charts containing 18,402 lines of clinical information, outperforming the compared baseline models.

Source-provided article image: A medically grounded LLM agent-based tool to detect patient safety events in medical records.
PubMed

Interpretation

It introduces SAFE-AI (Structured and Automated Framework for Explainable AI), which converts a clinician-defined ontology of epinephrine errors based on established guidelines into a directed graph, uses the LLM only to extract initial graph nodes such as age, weight, dose, and timestamps, and performs the final determination through deterministic rule-based code. Unlike approaches that place guidelines, task definition, and case in a single prompt and ask the model to decide directly, this design separates information extraction from clinical decision-making, confining the source of hallucination to the extraction step. The method is described in full, including a four-step process, template-based prompting, JSON-structured outputs, and code generated in a sandbox and manually reviewed and validated by domain experts; the code repository is publicly available.

On 300 pediatric out-of-hospital cardiac arrest charts independently dually reviewed by experts, SAFE-AI achieved 97.9% accuracy in detecting epinephrine overdose and 91.6% accuracy in identifying delays in epinephrine administration, which the authors describe as similar to human experts and greatly outperforming baseline LLM models. Relative to the LLM-only and LLM-with-ontology baselines, SAFE-AI achieved a stronger overall balance between sensitivity and specificity and was more consistent in multi-step temporal reasoning and cumulative dose interpretation. The gold standard used two independent expert reviewers with disagreements resolved by complete consensus, and Gwet's AC1 was 0.85 for delays and 0.96 for incorrect doses; paired McNemar analyses showed b = 34, c = 0 for delay and b = 40, c = 0 for overdose versus the LLM-only baseline, with no discordant pairs (b = c = 0) versus the LLM + Ontology baseline.

Error analysis found that remaining SAFE-AI failures were dominated by ambiguity and inconsistency inherent to real-world clinical documentation, with the largest source of disagreement arising from inter-reviewer variability; among true model-related errors, most originated during information extraction rather than ontology execution. This analysis distinguishes model error from uncertainty in the underlying labels, suggesting that some apparent model errors reflect uncertainty in the clinical labels rather than algorithmic failure. Based on qualitative review of the 300 charts and a taxonomy of failure modes (Table 4), identifying formatting inconsistencies, incomplete timestamps, and conflicting documentation between structured and narrative components as the main extraction failure sources.

On efficiency, specialized clinician reviewers required more than 10 minutes on average to review a two-page chart, whereas SAFE-AI completed the same task in under one minute. This suggests symbolic clinical AI systems may help scale patient safety surveillance and quality improvement efforts while maintaining clinically auditable decision-making. Reported by the authors in the discussion as a time comparison, representing an initial observation within a single event type and dataset.

Perspective

The result applies to epinephrine administration as a specific safety event in pediatric out-of-hospital cardiac arrest, in emergency medical services charts that follow the NEMSIS standard and contain both structured fields and narrative text; the authors note the framework can be extended to other safety event domains but requires creating and validating new ontologies for each, and plan evaluation on larger multi-center datasets.

A careful reader would still watch how the framework generalizes beyond 300 charts and a single event type; whether the expert effort required for ontology construction can scale; how robust extraction is when structured fields conflict with narrative text; and how cases where reviewers themselves disagree should be counted as model errors. The threshold criteria and failure-mode taxonomy are presented as tables in the original and were not expanded in this parse, so reproducing them would require consulting the original tables.

Sources