Skip to main content
Back to timeline
Frontiers in PharmacologySource publication:

Clustering routine ICU data yields three overlapping HFpEF phenotypes but no phenotype-specific medication associations, with a TabPFN classifier reaching internal-validation AUCs of 0.951–0.969

Synopsis

In this multicohort retrospective study, K-prototypes clustering of first-24-hour ICU variables in 2,511 patients with HFpEF from MIMIC-IV produced three clinically interpretable but partially overlapping phenotypes—cardiorenal-metabolic, hypertensive-pulmonary, and low-blood-pressure/arrhythmia (K = 2 had a higher mean silhouette width than K = 3, 0.083 versus 0.060, while both showed high median resampling stability, ARI 0.940 versus 0.924)—with a graded 365-day mortality difference in the derivation cohort (38.3%, 31.5%, 23.

Source-provided article image: Machine-learning phenotyping and exploratory medication-outcome associations in heart failure with preserved ejection fraction: a multicohort retrospective study
FIGURE 1

FIGURE 1 Schematic illustration of the study design and analytical workflow. Abbreviations: ACEI, angiotensin-converting enzyme inhibitor; ARB, angiotensin receptor blocker; AUC, area under the receiver operating characteristic curve; FDR, false discovery rate; KNN, k-nearest neighbors; LightGBM, Light Gradient Boosting Machine; MIMIC, Medical Information Mart for Intensive Care; OW, overlap weighting; RFE, recursive feature elimination; SHAP, SHapley Additive exPlanations; sIPTW, stabilized inverse probability of treatment weighting; TabPFN, Tabular Prior-Data Fitted Network.

· Page 3

Interpretation

K-prototypes clustering of first-24-hour ICU variables in MIMIC-IV identified three clinically interpretable HFpEF phenotypes: a cardiorenal-metabolic phenotype (n = 733; median creatinine 2.2 mg/dL, diabetes 74%, renal disease 84%), a hypertensive and pulmonary phenotype (n = 922; chronic pulmonary disease 63%, hypertension 43%), and a low-blood-pressure/arrhythmia phenotype (n = 856; median systolic blood pressure 115 mmHg, arrhythmias 70%). Earlier HFpEF phenomapping work, such as the TOPCAT-based clustering study, focused mainly on phenotype identification and prognosis; this study places phenotype derivation, cross-cohort reproducibility, and phenotype-stratified medication associations within one framework and explicitly treats K = 2 as an equally plausible lower-resolution sensitivity solution. Derivation cohort n = 2,511; variables were Z-score standardized and rescaled to [-1, 1] before clustering; solutions from K = 2 to K = 10 were compared using within-cluster sum of squares and average silhouette width, and stability was assessed with 50 repeated 80% subsampling iterations using 15 random initializations each, summarized by the adjusted Rand index; the cluster-number decision rule was specified before the medication-outcome associations were examined, and medication exposure and survival outcomes were not used as clustering inputs.

Observed 365-day mortality was graded across the three clusters in the derivation cohort (38.3%, 31.5%, 23.0%; P < 0.001), but this prognostic separation was not reproduced as a statistically significant between-cluster difference in the internal validation or external application cohorts (log-rank P = 0.072 and P = 0.37). The result separates transportability of phenotype characteristics from transportability of prognostic stratification: cluster profiles could be aligned across cohorts, whereas cluster prevalence and prognostic separation depended on the population and setting, and K = 2 showed stronger internal-validation survival separation than K = 3, indicating that prognostic stratification was resolution-dependent. The three cohorts comprised 2,511, 1,524, and 340 patients, with 365-day deaths of 768 (30.6%), 472 (31.0%), and 127 (37.4%), respectively; survival comparisons used Kaplan-Meier curves and the log-rank test, and external cluster labels were assigned by the fixed classifier without refitting.

In the derivation and internal validation cohorts, no cluster-specific weighted association for diuretics, ACEI/ARB, or beta-blockers remained statistically significant after FDR correction, and all medication-by-cluster interaction tests were nonsignificant; external medication estimates were reported descriptively because within-group covariate balance was not achieved. The study states explicitly that failure to detect an interaction does not demonstrate identical medication associations across phenotypes and does not support selecting medication by cluster membership, while reporting the full trajectory of the ACEI/ARB estimate in Cluster 1, which had a raw P value below 0.05 (OW: HR = 0.46, 95% CI 0.26-0.82; P = 0.008) but an FDR-adjusted P of 0.098. A 24-hour landmark design was used; propensity scores included cluster membership, 14 prespecified covariates, and all cluster-by-covariate interaction terms; overlap weighting was primary with stabilized inverse probability of treatment weighting as sensitivity analysis, weights were truncated at the 1st and 99th percentiles, covariate balance was judged by absolute standardized mean differences below 0.10, the proportional-hazards assumption was assessed with scaled Schoenfeld residuals, and the Benjamini-Hochberg procedure was applied to the 12 medication-coefficient tests.

A fixed TabPFN classifier using 14 routinely available variables achieved one-versus-rest AUCs of 0.951-0.969 in internal validation with a 10-bin multiclass expected calibration error of 0.028, and was deployed as a publicly accessible web-based phenotype-assignment interface. The classifier was designed to reproduce clustering-derived phenotype labels rather than to predict mortality or treatment response; internal validation compared it against labels from independent de novo K-prototypes clustering in MIMIC-III, aligned one-to-one by prespecified Hungarian assignment on cluster-profile distances without using classifier predictions, survival outcomes, or medication exposures, and the model was not refitted in MIMIC-III. Recursive feature elimination compared candidate feature-set sizes of 8, 10, 12, 14, and 37 predictors; the 37-predictor model had the highest absolute cross-validated performance, while the 14-predictor model was the best reduced-input configuration and was retained to limit manual data-entry burden; single-feature ablation showed the largest macro-F1 reductions when hematocrit, respiratory rate, glucose, and renal disease were removed (Δmacro-F1 = 0.0320, 0.0253, 0.0240, and 0.0232), and SHAP values came from an auxiliary LightGBM surrogate.

Perspective

The framework targets ICU patients with HFpEF, LVEF of at least 50%, and prespecified heart-failure coding criteria, using routine vital signs, laboratory values, and ICD-coded comorbidities from the first 24 hours after ICU admission; it is intended for research-oriented phenotype description and standardized label assignment, with the classifier and web interface positioned as research tools rather than clinical decision support or treatment selection. Researchers wishing to reuse the workflow can adopt the 14-variable input configuration and the public interface, and should re-evaluate cluster structure and prediction confidence in new cohorts.

This is an incomplete reading: figures and supplementary material (for example Supplementary Table S4, S6, and S9 and Supplementary Figure S16, S17, and S20) were not loaded, so the full cluster-selection comparison, covariate-balance details, proportional-hazards diagnostics, and calibration curves cannot be checked here and remain open questions to confirm against the original. In addition, sensitivity of the cluster structure to variable selection, missing-data handling, and cohort composition; the absence of independent reference labels in the external cohort, which precluded calculating external classification accuracy; and the fact that SHAP values came from a LightGBM surrogate rather than directly explaining TabPFN are all directions a reader would continue to watch before adopting the framework, while whether phenotypes remain stable over time and what prespecified phenotype-stratified design would be needed to test treatment hypotheses await prospective multicenter study.

Sources