Skip to main content
Back to timeline
JMIR Medical InformaticsSource publication:

Prediction Models for In-Hospital Delirium Using Routinely Collected Electronic Health Record Data: Systematic Review

Synopsis

This systematic review searched PubMed, MEDLINE, Embase, PsycINFO, and Web of Science from inception to November 11, 2025, and included 29 studies that developed, validated, or evaluated multivariable prediction models using routinely collected electronic health record or administrative data to predict acute mental status deterioration during adult hospital admissions, all operationalized as delirium; using CHARMS and TRIPOD/TRIPOD-AI for data extraction and PROBAST for risk of bias, it found that the evidence clustered into four overlapping prediction tasks (admission or early-stay risk stratification, perioperative or postoperative prediction, dynamic intensive care unit prediction, and external validation or workflow evaluation of existing tools), that most studies were retrospective co

Source-provided article image: Prediction Models for In-Hospital Delirium Using Routinely Collected Electronic Health Record Data: Systematic Review.
Figure 1. ·

PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) flow diagram of study selection [ 9 ].

PubMed

Interpretation

The review organizes the evidence on EHR-based in-hospital delirium prediction models into four overlapping prediction tasks: admission or early-stay risk stratification, perioperative or postoperative prediction, dynamic intensive care unit prediction, and external validation or workflow evaluation of existing tools. Rather than cataloguing isolated single-model reports, it uses the prediction task as the organizing thread, placing models from different clinical settings within one comparative frame. Based on 29 studies meeting the inclusion criteria, with data extraction informed by CHARMS and TRIPOD/TRIPOD-AI and risk of bias and applicability assessed using PROBAST.

Machine learning or hybrid approaches were common in this field (18/29, 62%), but more complex models did not consistently outperform statistical or rule-based approaches. This observation separates model complexity from actual predictive gain, indicating that the choice of method alone does not guarantee performance. Derived from a narrative synthesis of 29 studies, most of which were retrospective cohorts (20/29, 69%).

Reporting completeness was markedly uneven: internal discrimination was reported in 24 studies (83%; AUC range 0.77-0.97), external discrimination in 12, calibration in 15, and decision curve analysis in only 3, while prospective evaluation or workflow integration remained limited. By counting discrimination separately from calibration, clinical usefulness, and external and prospective validation, the review shows which links in the evidence chain are repeatedly reported and which remain sparse. Counts come directly from extraction of the 29 included studies, with performance, validation, calibration, and implementation features synthesized narratively.

Overall risk of bias was low in 8 studies, unclear in 10, and high in 11, mainly because of analysis-domain limitations; the authors conclude that routinely collected EHR data can support delirium risk prediction across hospital settings, but that no single algorithm is ready for routine adoption. By presenting risk-of-bias assessment alongside a clinical-readiness judgment, it offers a more complete basis for deciding whether deployment is warranted than performance metrics alone. Risk of bias and applicability were assessed with PROBAST, and the conclusion rests on the overall distribution across the 29 studies.

Perspective

This review is aimed at researchers and clinical informatics teams using routinely collected EHR or administrative data to predict delirium during adult hospital admissions, and its conclusions apply to risk-prediction designs in the settings covered, including general wards, mixed ward-intensive care units, intensive care units, and emergency departments; the authors recommend that future studies define the intended clinical use case before model development, evaluate calibration and clinical usefulness alongside discrimination, and test models across institutions, time periods, and workflows before deployment.

Readers should still watch for inconsistent outcome ascertainment across studies, heterogeneous prediction tasks, sparse reporting of calibration and decision-analytic measures, and limited external and prospective evaluation; in addition, the available content here is abstract-level text without the studies' figures and item-level data, so individual models' variables, sample sizes, and specific calibration results cannot be further verified and await consultation of the original article.

Sources