Skip to main content
Back to timeline

Keeping all 208 molecular descriptors and adding mixup virtual samples, random forest reaches R² 0.72 for corrosion inhibition efficiency, with SHAP pointing to Chi4n and other topological descriptors

Synopsis

Using SMILES strings of 317 organic corrosion inhibitor compounds, this study computed and retained all 208 two-dimensional molecular descriptors via RDKit, expanded the training set with mixup-interpolated virtual samples, compared Random Forest, Bagging and Gradient Boosting, and interpreted the models with SHAP, correlation analysis, partial dependence plots and a Williams plot; Random Forest performed best on the original data (test MAE 4.568, RMSE 6.122, R² 0.718), R² for all three models rose from roughly 0.66–0.70 to above 0.90 and then saturated as virtual samples grew from 100 to 10,000, and SHAP identified topological descriptors Chi4n, Chi3v, Chi1n, Chi4v plus MolMR as the most influential features.

AI-generated editorial illustration: Explainable machine learning prediction of corrosion inhibition efficiency from molecular descriptors using virtual sample generation

Interpretation

On the original 317 organic compounds, Random Forest achieved test-set MAE 4.568, RMSE 6.122 and R² 0.718, outperforming Bagging (4.617, 6.559, 0.677) and Gradient Boosting (5.187, 7.487, 0.579). Compared with the prior work cited by the authors on the same 317-sample dataset, this study lowers MAE from 4.80 to 4.57 and RMSE from 6.40 to 6.12 while holding R² at 0.72; Gradient Boosting reached training R² 0.998 but only 0.579 on test data, indicating boosting is more prone to overfitting on small data. Based on 10-fold cross-validation test metrics after Grid Search tuning, with reported optimal hyperparameters (e.g., RF n_estimators=400, max_depth=None; GB n_estimators=600, learning_rate=0.01).

Mixup feature-space interpolation generated virtual samples (x_new=λx_i+(1−λ)x_j, y_new=λy_i+(1−λ)y_j) applied only to the training set while validation kept original data to avoid leakage; as samples grew from 100 to 10,000, R² rose consistently across all three models, exceeding 0.90 at several thousand samples and then plateauing. Unlike prior work emphasizing feature selection, this study performs no feature reduction and retains all 208 descriptors, instead expanding data volume; IE distribution and PCA visualizations show augmented samples lie within the original feature space without forming new clusters, indicating gains come from increased density of valid data rather than unrealistic patterns. Evidence comprises R² curves across 100–10,000 samples, before/after IE distribution comparison, and two-dimensional PCA projection, consistent across RF, Bagging and GB.

SHAP analysis shows molecular connectivity indices Chi4n, Chi3v, Chi1n, Chi4v and Chi3n, together with MolMR (molecular refractivity, linked to electron polarizability and molecular volume), are the most important features for predicting inhibition efficiency, with LabuteASA and Ipc also contributing notably. The result links model predictions to inhibition mechanisms: topological complexity supplies more adsorption centers (heteroatoms, lone pairs, π-electron systems) and polarizability strengthens electron donor-acceptor interactions, forming a more stable protective layer; correlation analysis, partial dependence plots and SHAP corroborate one another. Based on the SHAP summary plot, the Chi4n dependence plot (including interaction with Chi1v), a single-sample waterfall plot (base value about 85.7 rising to about 95.4), scatter regressions of the top five descriptors against experimental IE, and partial dependence curves.

Williams plot applicability-domain analysis shows the vast majority of molecules fall within the standardized residual ±3 and leverage warning value of about 0.589, so the model gives reliable predictions for most molecules in the dataset. Leverage was computed on a PCA-reduced descriptor space (55 principal components retained) to improve numerical stability in high dimensions, defining the descriptor range over which the model can be extrapolated beyond reporting predictive performance. Based on the leverage-residual distribution in the Williams plot; a few high-leverage molecules are treated as structural outliers and a few large-residual molecules as response outliers, but their number is small relative to the whole dataset.

Perspective

The framework targets organic corrosion inhibitor datasets represented as SMILES with experimental IE labels, suited to settings where data are scarce and virtual samples expand the training set before ensemble regression and SHAP-style interpretation; for researchers screening candidate molecules early or judging whether a molecule falls inside the model's applicability domain, the paper offers a reusable workflow and criteria (leverage threshold about 0.589, residuals ±3). The authors note future work could add more physicochemical descriptors, quantum-chemical parameters or graph-based representations, and integrate the model into a high-speed virtual screening pipeline.

Virtual samples are interpolated within the training set, so their information derives from the original 317 compounds; the rise of R² above 0.90 therefore reflects fitting and convergence on the augmented distribution rather than new independent experimental evidence, and since validation always uses original data, extrapolation to genuinely new molecules still needs external data. A few structural and response outliers remain in the applicability-domain analysis, and prediction reliability for those molecules is not fully settled. In addition, this is a full-text read in which figures are described in prose, so specific numerical details (such as absolute SHAP values per descriptor or PDP curve inflection points) require consulting the original figures.

Sources