Skip to main content
Back to timeline
Frontiers in MedicineSource publication:

Leakage-aware ML predicts MOF drug loading and cell viability at in-domain R² of 0.557 and 0.774, but R² turns negative when whole publications are held out

Synopsis

Using data reconstructed from Wang et al.'s supplementary tables, this study built 161 loading-capacity observations with 110 descriptors and 444 cell-viability observations with 25 descriptors, optimized histogram-based gradient boosting regression (HGBR) and partial least squares regression (PLSR) with differential evolution under five-fold grouped cross-validation, and constrained exact duplicate records to the same partition to prevent information leakage; on the duplicate-safe 20% holdout HGBR was strongest for cell viability (R² = 0.774, RMSE = 11.655 percentage points, MAE = 8.153, AARD = 16.69%) while PLSR was best for loading capacity (R² = 0.557, RMSE = 0.345 g/g, MAE = 0.210 g/g), yet holding out entire source publications drove all R² values negative (viability −0.391 and −0.

Source-provided article image: Leakage-aware machine learning for data-driven performance prediction of metal–organic framework systems
Page 2

Interpretation

The study treats the two MOF drug-delivery endpoints as separate prediction problems rather than imposing a row-wise correspondence across publications. Prior work on the same endpoints often reported both within one high-accuracy pipeline, whereas this study explicitly separates the observation counts and feature spaces and selects a model family per endpoint. 444 viability observations with 25 predictors and 161 loading observations with 110 predictors; the two matrices differ in composition (viability includes metal indicators, linker/functional-group indicators, cell-category indicators, zeta potential, particle size, exposure time, and concentration; loading includes metal indicators, linker/functionality indicators, and 66 molecular fingerprint bits).

On the duplicate-safe 20% holdout the endpoint determined the better model: HGBR reached R² = 0.774, RMSE = 11.655 percentage points, MAE = 8.153, and AARD = 16.69% for cell viability, while PLSR reached R² = 0.557, RMSE = 0.345 g/g, and MAE = 0.210 g/g for loading capacity. Unlike workflows that rely only on a random row-level split, this study assigns exact duplicates a common group identifier, re-fits preprocessing parameters inside each training fold, and lets differential evolution select hyperparameters under grouped cross-validation. Viability HGBR grouped-CV R² = 0.723 versus holdout R² = 0.774, with training R² = 0.970 exceeding the cross-validated value; loading PLSR grouped-CV R² was only 0.071 while holdout R² was 0.557, indicating the single holdout happened to be more predictable than several cross-validation folds.

When entire source publications were held out, all models produced negative R², indicating the models capture interpolation within the material and experimental regimes already compiled rather than transfer across publication sources. This stress test answers a different question from the primary holdout: the former estimates interpolation among represented regimes, the latter probes transfer under shifts in synthesis, measurement, formulation, and reporting practices. Viability R² was −0.391 (HGBR) and −0.119 (PLSR); loading R² was −0.342 (HGBR) and −0.679 (PLSR); negative R² means prediction error exceeded the error from using the mean of the corresponding stress-test target values.

Permutation importance identified the variables each endpoint model relies on: concentration caused the largest RMSE increase for viability, followed by Zn, zeta potential, Fe, normal-cell status, and particle size; loading PLSR was most sensitive to 2,5-dioxidoterephthalate and Mg-related descriptors, followed by hydroxyl functionality, Cr, and selected molecular-fingerprint bits. The authors stress that these rankings reflect predictive dependence under the fitted model and may also encode study design, since particular MOFs, cell classes, or concentration ranges cluster within publications, rather than isolated biological mechanisms. Each predictor was randomly permuted thirty times and the mean increase in RMSE recorded; the authors note that rare binary indicators can acquire high leverage in a small dataset.

Perspective

The framework is intended for ranking candidates similar to those already represented within the compiled material and experimental regimes, and for researchers who want to prioritize MOF formulations from literature data rather than replace loading measurements or cytotoxicity testing; the authors' proposed next steps are to define standardized experimental protocols, record uncertainty and replicate measurements, include quantitative pore and stability descriptors, and reserve entire laboratories or newly generated batches for external validation, with graph-based drug and framework embeddings to follow once sample size supports them.

This is an incomplete reading: the content of Figures 1 through 5 and the full forms of Equations 1 through 7 are not included in the loaded text, so distribution shapes, residual structure, and optimization traces cannot be checked; in addition, the viability table contains no explicit drug-identity or molecular-fingerprint variable and the loading table lacks exposure and biological variables, so the endpoints cannot be combined into a valid two-output model, and binary structural indicators do not fully describe pore volume, accessible surface area, framework defects, drug ionization, solvent composition, or release kinetics, meaning cross-source predictive capability still awaits confirmation through standardized prospective measurements.

Sources