Development and External Validation of a Multimodal Artificial Intelligence Mortality Prediction Model of Critically Ill Patients Using Multicenter Data
Synopsis
Using 203,434 ICU admissions from 2001 to 2022 across more than 200 hospitals in the MIMIC-III, MIMIC-IV, eICU, and HiRID databases, the study developed and externally validated a multimodal deep-learning model that predicts subsequent inpatient mortality from time-invariant variables, time-variant variables, clinical notes, and chest x-ray images available within the first 24 h of ICU admission; with structured data alone the model reached an AUROC of 0.92 (95% CI, 0.90 to 0.93), an AUPRC of 0.53 (95% CI, 0.49 to 0.57), and a Brier score of 0.19 (95% CI, 0.18 to 0.20), external validation across eight eICU institutions yielded AUROCs of 0.84 to 0.92, and in the subgroup with both notes and imaging, adding text and images raised the AUROC modestly from 0.87 (95% CI, 0.85 to 0.89) to 0.
Fig. 1. Open multimedia modal Schematic of the multimodal deep-learning model for predicting mortality after 24 h of inpatient data. The architecture of our multimodal deep-learning model consists of four component models: (1) a multilayer perceptron for time-invariant variables; (2) a bidirectional long short-term memory model for time-variant data and time-series vital sign variables; (3) a BERT language model for clinical notes; and (4) a DenseNet convolutional neural network for chest x-ray imaging. Model outputs are then unified into a classification model that predicts the outcome. Generalizability, bias/fairness, and explainability were addressed. BERT, bidirectional encoder representations from transformers; CNN, convolutional neural network; EHR, electronic health record; LLM, large language model; LSTM, long short-term memory.
PubMedInterpretation
The work builds a multimodal deep-learning mortality model that integrates four data modalities: time-invariant variables through a multilayer perceptron, time-variant variables through a bidirectional long short-term memory network, clinical notes through a pretrained BERT language model, and chest x-rays through a DenseNet121 convolutional network, with component embeddings concatenated in a pooling layer before a classification head outputs the mortality probability. Prior deep-learning models of ICU outcomes typically used EHR-derived clinical data from a single modality; this work unifies structured time-series data, unstructured text, and imaging in one architecture and fixes a single prediction time point at 24 h after ICU admission. The architecture, layer sizes and activations, and training strategy (Xavier uniform initialization, weighted random sampling and weighted cross-entropy loss, Adam and AdamW optimizers, early stopping, and checkpointing) are described in the Methods, and the data harmonization code is publicly available in a GitHub repository.
With structured data only, the model achieved AUROCs of 0.92 (95% CI, 0.90 to 0.93) and 0.92 (95% CI, 0.91 to 0.93) in internal validation on MIMIC-III and MIMIC-IV, and was externally validated on eICU, HiRID, and a temporally separated MIMIC population from the 2020 to 2022 COVID-19 era, with an overall eICU AUROC of 0.87 (95% CI, 0.86 to 0.87) and institution-level AUROCs of 0.84 to 0.92 across eight sites. External validation was extended to a temporally separated MIMIC population, the single-center Swiss HiRID data set, and the multicenter United States eICU data set covering 208 hospitals, providing evidence of generalizability across regions and time periods. The sample comprised 203,434 ICU admissions with mortality rates of 5.2 to 7.9% across the four data sets; confidence intervals were computed with 1,000 bootstrap resamples and AUROCs were compared using DeLong's test.
In the 9,881 MIMIC-IV admissions with both clinical notes and chest x-rays, adding text and imaging raised the AUROC from 0.87 (95% CI, 0.85 to 0.89) to 0.89 (95% CI, 0.87 to 0.91), the AUPRC from 0.43 (95% CI, 0.36 to 0.51) to 0.48 (95% CI, 0.40 to 0.56), and improved the Brier score from 0.37 (95% CI, 0.36 to 0.39) to 0.17 (95% CI, 0.16 to 0.19). The analysis quantifies the incremental predictive value of unstructured text and imaging over structured time-series data, and shows that chest x-rays alone (AUROC 0.76; 95% CI, 0.721 to 0.794) or clinical notes alone (AUROC 0.77; 95% CI, 0.735 to 0.811) performed below structured time-variant data (AUROC 0.87; 95% CI, 0.847 to 0.897). The AUROC difference had a DeLong test P = 0.0167, and the comparison of all modalities against time-variant structured data alone had P = 0.017; the authors state that, because of the need for multiple testing, they did not perform statistical comparisons between models for F1, AUPRC, or Brier score.
The model outperformed existing clinical severity scores: on MIMIC-III the best-performing score, SAPS-II, reached an AUROC of 0.75 (95% CI, 0.73 to 0.77) and an AUPRC of 0.22 (95% CI, 0.19 to 0.25), whereas the deep-learning model, which also included these scores as inputs, reached an AUROC of 0.92 (95% CI, 0.91 to 0.93) and an AUPRC of 0.53 (95% CI, 0.49 to 0.57), with an AUROC difference of P < 0.05. The comparison places the multimodal deep-learning model against four widely validated scoring systems, SOFA, SAPS-II, OASIS, and APACHE-II, within the same data and evaluation framework. The comparison shows consistent trends on both MIMIC-III and MIMIC-IV, and is accompanied by a supplemental analysis of cases where the model and SAPS-II disagreed.
Perspective
The model is designed for a single prediction time point at 24 h after ICU admission, using only data known within that window, so its intended use is outcome benchmarking and early triggering of palliative care or goals-of-care discussions rather than continuous early warning throughout the ICU stay; external validation covered structured data modalities only because eICU and HiRID do not contain clinical notes or chest x-rays, so the incremental value of text and imaging was assessed only within a MIMIC-IV internal subgroup; the authors propose redesigning a more dynamic model for continuous early warning and call for prospective validation in real clinical settings.
Readers should note that the model predicts at a single time point, and the authors report that performance declined as the prediction window extended (1 day, within 7 days, and after 7 days); comorbidities were recorded as binary without severity stages; the study is based on retrospective data that may be confounded by missing data or unaccounted variables and requires prospective validation; fairness analyses showed higher false-positive rates for patients older than 75 yr and higher false-negative rates for the white population in MIMIC-IV, with the authors cautioning that age may act as a surrogate for unmeasured variables such as frailty; the withdrawal-of-care text audit identified 1.6% of notes corresponding to 1.5% of patients, and hospice information covered about 2% of patients in one data set, so those sensitivity analyses may be limited by event counts; in addition, the full performance metrics in the supplemental figures and tables were not loaded with the main text, so item-by-item verification of each modality's metrics requires consulting the supplemental material.
