Skip to main content
Back to timeline
Frontiers in EndocrinologySource publication:

Explainable XGBoost model predicts clinical pregnancy after frozen embryo transfer in 1,319 single-center cases, with AUC falling from 0.854 to 0.722 in temporal validation

Synopsis

Using retrospective single-center data from 1,319 patients undergoing their first frozen embryo transfer at Guangdong Provincial Hospital of Chinese Medicine, split by transfer date into a training cohort (n=1,013) and a temporal validation cohort (n=306), the authors used LASSO to select age, induced abortions, cycle type, good-quality embryos, transferred embryos, endometrial change, BMI, antral follicle count, embryo transfer-day endometrial thickness, and retrieval-transfer interval, then developed an XGBoost model and compared it with logistic regression; XGBoost achieved an AUC of 0.854 (95% CI 0.832-0.877) in training and 0.722 (95% CI 0.666-0.779) in temporal validation, with AUPRC values of 0.840 and 0.707, acceptable calibration in validation (Brier score 0.212, intercept 0.

Source-provided article image: Development and internal temporal validation of an explainable XGBoost model for predicting clinical pregnancy after frozen embryo transfer
Figure 1

(Figure 1). The final variables included in the model were age, induced abortions, cycle type, good-quality embryos, transferred embryos, endometrial change, BMI, AFC, embryo transfer-day ET,

· Page 4

Interpretation

In a single-center cohort of 1,319 frozen embryo transfer cycles, the XGBoost model discriminated clinical pregnancy with an AUC of 0.854 in training and 0.722 in temporal validation, AUPRC of 0.840 and 0.707, and validation calibration of Brier score 0.212, intercept 0.056, and slope 0.965. Prior regression models for frozen embryo transfer or IVF pregnancy outcomes have reported discrimination in the 0.60-0.70 range; this study reports discrimination, calibration, and decision curve metrics together within one cohort and adds a temporal validation result. Retrospective single-center study with 1,013 training and 306 temporal validation patients, 642 clinical pregnancies including 11 ectopic pregnancies; internal validation used 1,000 bootstrap iterations for confidence intervals, and the temporal validation cohort was excluded from feature selection and hyperparameter tuning.

XGBoost showed a numerically higher AUC than logistic regression in temporal validation (0.722 vs 0.690), but the difference was not statistically significant by DeLong's test (P = 0.090); at the 0.30 threshold, XGBoost had higher sensitivity in training (0.969 vs 0.918) while logistic regression was slightly higher in validation (0.941 vs 0.928). Rather than presenting the numerical lead as a stable advantage, the authors state that threshold-dependent classification performance may vary across cohorts and should not be read as consistent superiority of either model. Both models were compared on the same training and temporal validation cohorts with AUC, AUPRC, calibration, and DCA; sensitivity analyses showed a continuous-variable logistic model at C-index 0.690 and a restricted cubic spline model at 0.717, both numerically below XGBoost.

SHAP analysis ranked the most influential predictors as antral follicle count, good-quality embryos, age, endometrial change, cycle type, pre-progesterone endometrial thickness, retrieval-transfer interval, and BMI, while transferred embryos and induced abortions contributed minimally; higher AFC and more good-quality embryos corresponded to larger SHAP values, age showed a negative association, and BMI showed a nonlinear pattern with intermediate values associated with higher predicted probabilities. Beyond commonly used predictors, the model included "endometrial change" (transfer-day minus pre-progesterone endometrial thickness, grouped as decrease <= -1 mm, stable -1 to +1 mm, increase >= +1 mm) to capture the dynamic endometrial trajectory from preparation to transfer rather than a single static thickness. SHAP bar plots, beeswarm plots, and individual waterfall plots provide feature-level and individual-level explanations; the authors note that SHAP values reflect associations rather than causal effects and that the influence of AFC may be indirect.

Decision curve analysis showed both models exceeded the treat-all and treat-none reference lines in training and temporal validation; in the validation cohort, compared with treat-all, XGBoost would yield about 4 fewer true positives and 15 fewer false positives per 100 patients, while logistic regression would yield about 3 fewer true positives and 12 fewer false positives. The result translates model performance into comparable clinical trade-offs rather than statistical metrics alone, helping clinicians see the gains and costs of threshold choice. DCA was performed in both training and temporal validation cohorts, with sensitivity, specificity, PPV, and NPV at the 0.30 threshold reported in Supplementary Material Part 7.

Perspective

The model is intended for reproductive medicine clinicians and is designed to be applied on the day of frozen embryo transfer, when all required predictors including transfer-day endometrial thickness are routinely available, without reliance on post-transfer information and without requiring technical expertise beyond routine clinical assessment and data entry. The authors emphasize that it supports probabilistic risk communication rather than excluding or disadvantaging patients with low ovarian reserve who nonetheless have transferable embryos; if a predictor is unavailable or of poor quality, clinical judgment should prevail or the prediction should be deferred. The findings apply to patients undergoing their first frozen embryo transfer at this center between March 2023 and March 2025, and local recalibration would be needed before cross-center use.

The temporal validation AUC fell from 0.854 in training to 0.722, and the two cohorts differed in secondary infertility, history of cesarean section, number of transferred embryos, number of good-quality embryos, diminished ovarian reserve, and adverse reproductive history; the authors describe this as temporal case-mix drift and note that the small validation cohort prevented reliable quantification of how much this shift contributed to the drop in discrimination. The model did not include lifestyle factors such as smoking or alcohol consumption, nor genetic and molecular biomarkers; patients with two or more prior induced abortions were limited in number, so model performance in that higher-risk subgroup could not be reliably evaluated. SHAP values reflect associations rather than causal effects, algorithm complexity may limit clinical implementation, and the validation results suggest potential overfitting. This is a fast-parse version, so figures and supplementary material (including variable grouping criteria, model parameters, threshold metrics, and partial dependence and ICE plots) are not fully available, and some details can only be summarized from the main text.

Sources