Machine learning for early post-ESWL risk stratification of pancreatitis in chronic pancreatitis
Synopsis
Using retrospective deidentified clinical data from 1370 inpatients with chronic pancreatitis who underwent extracorporeal shock wave lithotripsy (ESWL) at Changhai Hospital between May 31, 2016 and June 26, 2019, this study retained 55 of 109 variables and compared ten machine-learning and deep-learning algorithms (XGBoost, LightGBM, CatBoost, random forest, support vector machine, artificial neural network, multilayer perceptron, TabNet, Transformer, and Wide&Deep) across three modeling schemes (all preprocessed variables, 21 variables selected by univariate screening, and ADASYN-oversampled training data) using the F1-score as the primary selection metric; TabNet achieved the best held-out test performance in the first two schemes (F1 of 0.
Figure 1 Open multimedia modal Model performance and SHAP-based interpretation under different variable-processing strategies. (A) F1-Score results of the model based on the preprocessed variables. (B) F1-Score results of the model based on variables from univariate analysis. (C) F1-Score results of the model based on oversampled variables. (D) SHAP plot of the TabNet model based on preprocessed variables. (E) SHAP plot of the TabNet model based on variables from univariate analysis. (F) SHAP plot of the LGBM model based on oversampled variables. Panels A and D show the results obtained using the preprocessed variables, including the F1-score comparison of 10 machine-learning models on the test set (A) and the SHAP summary plot of the TabNet model (D). Panels B and E show the corresponding results based on variables selected by univariate analysis. Panels C and F show the corresponding results based on the oversampled training dataset, including the F1-score comparison (C) and the SHAP summary plot of the LGBM model (F). In the SHAP summary plots, each point represents one sample; the x-axis indicates the SHAP value (impact on model output), and the color reflects the feature value from low (blue) to high (red). ALP: Alkaline phosphatase; ALT: Alanine aminotransferase; AMS: Amylase; AST: Aspartate aminotransferase; DM: Diabetes mellitus; ESWL: Extracorporeal shock wave lithotripsy; Glu: Glucose; HCT: Hematocrit; PLT: Platelet; RBC: Red blood cell; RP: Recurrent pancreatitis; WBC: White blood cell.
PubMedInterpretation
The study developed and internally validated a machine-learning model based on routinely collected peri-procedural clinical data for early risk stratification of post-ESWL pancreatitis, addressing a gap in which prior AI models focused mainly on post-ERCP pancreatitis and few models had been established to recognize post-ESWL pancreatitis. The authors state that although AI models have been developed to identify post-ERCP pancreatitis, few such models have been established yet to recognize post-ESWL pancreatitis, so this work extends AI risk stratification to the specific ESWL setting. Based on a single-center retrospective cohort of 1370 patients, of whom 144 (10.51%) experienced post-ESWL pancreatitis, with a patient-level 8:2 split into training and held-out test sets and 10-fold cross-validation comparison of ten algorithms.
Across the three modeling schemes, TabNet achieved the best held-out test performance in both the all-55-preprocessed-variable and the 21-univariate-selected-variable schemes with an F1-score of 0.71, whereas in the ADASYN-oversampled scheme LightGBM reached a mean cross-validation F1 of 0.95 ± 0.02 but dropped to 0.66 on the independent test set, indicating that oversampling did not improve generalizability and supporting selection of the more parsimonious 21-variable TabNet model as final. By comparing variable screening and oversampling strategies, the study shows that univariate screening (P ≤ 0.1) improved cross-validated stability while maintaining test-set performance (the 21-variable model achieved a mean cross-validated F1 of 0.55 ± 0.09, precision of 0.80 ± 0.21, and recall of 0.43 ± 0.09), whereas the oversampled scheme showed a marked gap between cross-validation and test performance. Reports cross-validated means with standard deviations, held-out test F1-scores, and positive- and negative-class recall and precision, with supplementary figures and confusion-matrix-derived summaries.
Feature-importance analysis and SHAP consistently identified postoperative serum amylase as the most influential variable, ranking first in the 55-variable TabNet model (importance 0.15) and the 21-variable TabNet model (importance 0.147) and remaining a top contributor in the oversampled LightGBM model (importance 623); SHAP plots further revealed positive associations for postoperative serum amylase, white blood cell count, and recurrent pancreatitis, alongside negative associations for mixed pancreatic stones, preoperative indomethacin use, steatorrhea, and diabetes mellitus. The study emphasizes that before the full diagnostic criteria for post-ESWL pancreatitis are met, postoperative serum amylase functions as an early biochemical signal captured well by the model, supporting its role in early risk stratification rather than simple label leakage, and notes that negative SHAP directions for some diabetes- and glucose-related variables should be interpreted with caution as conditional contributions within a non-linear multivariable model. Based on SHAP summary plots and feature-importance rankings in supplementary tables across the three modeling schemes.
The model reliably excluded patients without the complication (recall 98.38%, precision 96.05%) but had more limited ability to identify positive cases (recall 62.96%, precision 80.95%), leading the authors to conclude that the current model is better suited for early triage and monitoring after ESWL than for pre-procedural planning. The study explicitly defines the model's intended setting and notes that because early postoperative biochemical variables contributed substantially to model performance, the model is positioned as an early post-ESWL triage and monitoring tool. Based on positive- and negative-class recall and precision in the held-out test set, together with the authors' statement that the relatively low event rate limited positive-case prediction and severity stratification.
Perspective
The intended scope is the early post-ESWL period in patients with chronic pancreatitis, with the model positioned for early triage and monitoring rather than pre-procedural planning because early postoperative biochemical variables contributed substantially to model performance. The study is based on retrospective data from 1370 inpatients at a single center (Changhai Hospital) between 2016 and 2019, using a patient-level 8:2 split and internal validation, and the authors explicitly state that future multicenter studies are expected to externally validate this framework and further examine whether models restricted to pre-procedural variables can support earlier clinical decision-making. For readers, this means the results are currently useful for understanding the potential value of routinely collected clinical variables and postoperative amylase for early risk stratification within single-center data, rather than for direct clinical deployment at other institutions.
Several open questions remain for a careful reader: the model was validated internally at a single center only, with no external independent cohort to test portability; the relatively low event rate (144/1370) limited positive-case prediction and severity stratification; the oversampled scheme showed a marked gap between cross-validation and held-out test performance, suggesting oversampling did not improve generalizability in this dataset; several clinically relevant variables, including some peri-procedural details and postoperative symptom measures, were unavailable or incomplete; and the negative SHAP directions for some diabetes- and glucose-related variables reflect conditional contributions within a non-linear multivariable model, which the authors also advise interpreting with caution. In addition, the loaded text is a fast parse, and the specific values in Supplementary Figures 2–7 and Supplementary Tables 1–8 are not included in the main text, so readers seeking a complete assessment of model performance and feature contributions should consult the supplementary materials.
