Validating machine learning models in an independent external cohort: an ensemble model predicted incident hydronephrosis in spina bifida patients with a concordance index of 0.77 and classified bladder dysfunction with 71% accuracy
Synopsis
This study adapted three previously developed machine learning models (a deep learning convolutional neural network using pressure-volume data, a deep learning imaging model using fluoroscopic imaging data, and an ensemble model averaging the risk data from both) to an independent external cohort of spina bifida patients who underwent videourodynamics at a single institution between 2016 and 2025, for prediction of incident hydronephrosis and classification of bladder dysfunction severity; among 70 patients in the hydronephrosis cohort, 13 (18%) developed incident hydronephrosis, the ensemble model achieved a concordance index of 0.77 (pressure-volume only: 0.70, imaging only: 0.72) and an overall AUROC of 0.75 (pressure-volume: 0.67, imaging: 0.
Interpretation
The ensemble model outperformed individual modalities for predicting incident hydronephrosis, with a concordance index of 0.77 versus 0.70 for pressure-volume data alone and 0.72 for imaging alone, and an overall AUROC of 0.75. The models were previously developed in an original cohort; this study is the first to evaluate adapted versions in an independent external cohort, providing external validation evidence. Based on a hydronephrosis cohort of 70 patients, of whom 13 (18%) developed incident hydronephrosis, using concordance index and AUROC as discrimination metrics and comparing against expert pediatric urologist reviewers.
For classification of bladder dysfunction severity in 95 videourodynamic studies, the ensemble model achieved 71% accuracy, higher than 60% for pressure-volume data alone and 66% for imaging alone, with no substantial disagreements with expert reviewers. The classification task was extended from the development cohort to an independent external dataset, showing that multimodal fusion yields higher classification accuracy than single modalities. Based on 95 videourodynamic studies, using expert pediatric urologist review as reference, reporting accuracy and degree of disagreement.
The previously developed machine learning frameworks could be adapted to an independent dataset and achieved discrimination and accuracy similar to those previously reported. Provides an externally transferable auxiliary analysis pathway for the interrater variability problem in videourodynamic interpretation. The study was conducted in spina bifida patients at a single institution between 2016 and 2025, using three models (pressure-volume convolutional neural network, fluoroscopic imaging deep learning model, and ensemble model), and states that data are available upon reasonable request.
Perspective
This study aimed to evaluate the adapted performance of previously developed machine learning models in an independent external cohort, applicable to spina bifida patients undergoing videourodynamic assessment, for predicting incident hydronephrosis and classifying bladder dysfunction severity. The results provide external validation evidence for clinicians and researchers, showing that the ensemble model fusing pressure-volume and fluoroscopic imaging data can achieve discrimination and accuracy similar to those previously reported. The findings support further evaluation of such models as auxiliary tools for videourodynamic interpretation at more institutions, but the current study is based on data from a single institution between 2016 and 2025, with data available upon reasonable request.
The currently loaded text is an incomplete abstract and declarations section, lacking the full paper's figures, model adaptation details, statistical analysis, and subgroup results, so performance differences across patient characteristics or institutions cannot be assessed. In addition, only 13 events occurred among 70 patients in the hydronephrosis cohort, and the small number of events may affect the stability of discrimination estimates; bladder dysfunction classification was based on 95 videourodynamic studies, and the specific definition and consistency measures for 'no substantial disagreements' with expert review are not elaborated in the abstract. These are questions readers may continue to watch when interpreting the external validation conclusions.
