Skip to main content
Back to timeline
Frontiers in MedicineSource publication:

A mechanism-informed ML framework predicts polymeric long-acting injectable release with XGBoost (test R² = 0.9774) and finds predictive importance nearly uncorrelated with intervention effects (r = 0.132 and −0.021)

Synopsis

The study builds a mechanism-informed machine learning framework that chains a pharmaceutical-knowledge-driven directed acyclic graph, causal structure discovery (PC, NOTEARS, DirectLiNGAM), XGBoost release prediction, explainable AI (SHAP), ATE/CATE intervention-effect estimation, and counterfactual formulation analysis; on a public dataset of 181 release profiles, 3,783 fractional release measurements, and 43 drug-polymer pairs, XGBoost reached test R² = 0.9774, RMSE = 0.0491, and MAE = 0.0339, and correlation and SHAP importance agreed strongly (r = 0.761) while both agreed poorly with intervention-effect estimates (r = 0.132 and −0.021), indicating that variables useful for prediction differ from those that can be manipulated.

Source-provided article image: Mechanism-informed machine learning for interpretable drug release prediction and virtual formulation intervention in polymeric long-acting injectables
FIGURE 1

FIGURE 1 Distribution of representative features before and after standardization.

· Page 4

Interpretation

The framework encodes pharmaceutical domain knowledge as a manually constructed directed acyclic graph used as a mechanistic prior, then merges it with data-driven causal discovery into a consensus graph for downstream effect estimation. Prior machine learning work on polymeric long-acting injectables focused mainly on predictive accuracy and feature importance, while causal discovery, treatment-effect estimation, and virtual formulation optimization were generally studied separately; this work places them in one pipeline. The knowledge graph has 29 edges, the correlation graph 122 edges, the mutual-information graph 205 edges, and the consensus graph retains 100 edges (only edges present in at least two of the three sources), with pairwise Jaccard similarity reported as an overlap measure.

XGBoost performed best for cumulative release prediction with test R² = 0.9774, RMSE = 0.0491, and MAE = 0.0339, and the smallest train-test generalization gap, so it was selected as the primary model for downstream interpretability and counterfactual analysis. In a comparison of five ensemble models (RandomForest, GradientBoosting, XGBoost, LightGBM, CatBoost) under the same data split, XGBoost was best on both validation and test sets, while CatBoost was weakest (test R² = 0.9337). A 70/15/15 train/validation/test split with fixed random seeds and 5-fold cross-validation; Table 2 reports R², MAE, RMSE, and MAPE for each model across train, validation, test, and full data.

Correlation and SHAP importance agreed strongly (r = 0.761), but both agreed poorly with intervention-oriented effect estimates (correlation vs. causal effect r = 0.132; SHAP vs. causal effect r = −0.021), showing that predictive importance and intervention relevance are different quantities. This directly addresses the paper's central question of whether the most predictive variables are also the most effective formulation intervention targets; the results show they do not coincide, so conventional explainable machine learning alone is limited for formulation decision-making. Three importance vectors (absolute correlation, TreeSHAP mean absolute SHAP value, and the Stage-4 causal effect) were normalized to [0,1] and compared pairwise with Pearson correlation, as reported in Table 4.

Counterfactual and intervention analyses yield comparable virtual formulation scenarios: the direction of the time effect on release was consistently positive across four estimators (DoWhy, DML, Causal Forest, Linear DML), while magnitude ranged from 0.1860 to 0.3790; a multi-feature scenario predicted release rising from 0.212 to 0.917 against a baseline of 0.442. The work translates causal effect estimation into an actionable ranking of virtual experiments and explicitly treats effect magnitude as estimator-dependent rather than a single definitive value, while CATE distributions show the same factor affects different formulations differently. Four estimators estimated the same (treatment, outcome, confounder) triple; sweeping time across empirical percentiles (P0 to P100) raised predicted release from −0.001 to 0.890, all within the observed range and thus interpolation; the authors stress these are model predictions, not experimental causal evidence.

Perspective

The framework targets cumulative release prediction and virtual formulation screening for polymeric long-acting injectables (mainly PLGA microparticles or cylindrical implants, plus PLA and PCL), and applies to interpolation within the data-supported feature range. For formulation developers and experiment designers, it can compare candidate formulation changes before laboratory validation, prioritize experiments, and identify which variables merit intervention; for methods readers, it offers a reproducible pipeline linking domain-knowledge graphs, causal discovery, and effect estimation (stage outputs are numbered for traceability).

The authors state clearly that counterfactual effects remain computational predictions and cannot substitute for experimental causal evidence, and that future work should validate them experimentally, assess the framework on independent formulation datasets, incorporate uncertainty quantification, and develop stronger graph-learning methods. Readers should also note that intervention effect magnitude varies across estimators from 0.1860 to 0.3790 and should be read as estimator-dependent rather than a single definitive value; the time sweep stays within the observed range and is interpolation rather than extrapolation; and this reading covered an incomplete scope, so the specific content of Figures 4 through 14 and some algorithm details could not be fully verified, and specific figure evidence should be checked against the original.

Sources