When Do Simple Models Win? Machine Learning Architectures for UV Absorption Prediction
Synopsis
Using UV absorption wavelength (λmax) prediction as a testbed across 18,415 solute-solvent pairs (after the Greenman/Song duplicate-handling protocol) with stratified 5-fold cross-validation and two external data sets totaling over 40,000 molecules, this work compares five models spanning four architecture families (Random Forest, XGBoost, a directed message-passing GNN, a BiGRU, and a pretrained Transformer) and finds that Random Forest is statistically indistinguishable from the best deep learning model for screening within known chemical space (RMSE 31.50 ± 1.47 vs 33.15 ± 3.27 nm, p = 0.16) while training in 15 min on a CPU, whereas on 16 novel UV-absorbing compounds the GNN (MAE 26.2 nm), Transformer (26.3 nm), and BiGRU (28.6 nm) all outperform RF (38.
Interpretation
Model superiority depends on the task: for interpolation-based screening within known chemical space, Random Forest with Morgan fingerprints is statistically indistinguishable from the best deep learning model at a fraction of the compute cost. Prior architecture comparisons for molecular property prediction have rarely paired statistical equivalence with a training-cost contrast in one testbed; here the RMSE comparison (31.50 ± 1.47 vs 33.15 ± 3.27 nm, p = 0.16) together with 15 min on a CPU, lightweight grid-search tuning, and no GPU makes the conditions under which a simple model suffices concrete. Based on 18,415 solute-solvent pairs, stratified 5-fold cross-validation, and two external data sets totaling over 40,000 molecules, reported with mean ± standard deviation and a p value.
Deep learning is superior for exploring novel scaffolds: experimental validation on 16 novel UV-absorbing compounds shows the GNN, Transformer, and BiGRU all achieve lower MAE than Random Forest. This places the claim that learned representations extrapolate where fixed fingerprints cannot onto specific compound-level validation rather than only benchmark splits. Validation covers 16 novel UV-absorbing compounds, with GNN MAE 26.2 nm, Transformer 26.3 nm, BiGRU 28.6 nm, and RF 38.5 nm; the sample is small and the evidence is directional.
Encoding solvent identity improves all models, without domain-specific descriptor engineering. A simple SMILES/fingerprint concatenation injects solvent information, reducing RMSE by 21-32% on chromophores with multisolvent training coverage (60% of test records) and 6-8% across the full benchmark (p < 0.02), indicating the gain comes from data organization rather than elaborate feature design. Reports stratified reduction ranges and a significance level, and states the share of test records corresponding to multisolvent training coverage.
Interpretability analysis shows RF feature importance and BiGRU gradient saliency independently highlight the same chromophore motifs, and architectures with local inductive bias systematically outperform global-attention Transformers on this locally determined property. It links architectural differences to the locality of the property itself, offering dual evidence within one task for the idea that local-bias architectures suit locally determined properties. Based on agreement between two independent attribution methods, this is supporting mechanistic evidence rather than a causal experiment.
Perspective
The results are scoped to UV absorption wavelength (λmax) prediction and apply to two settings: screening within known chemical space and exploration requiring extrapolation to novel scaffolds; the quantified solvent-encoding gains correspond to chromophores with multisolvent training coverage (60% of test records) and to the full benchmark. The proposed idea that local-bias architectures suit locally determined properties remains a hypothesis, with validation across property types explicitly left as future work.
The novel-scaffold conclusion rests on experimental validation over 16 novel UV-absorbing compounds, so its extrapolation range warrants continued observation; solvent-encoding reductions differ substantially by training coverage (21-32% vs 6-8%), so realized gains depend on the share of multisolvent coverage in the data; the interpretability conclusion comes from agreement between two attribution methods, and the local-bias hypothesis still needs testing across more property types. This reading is at summary scope and does not include figures or full methodological detail, so judgments about specific implementations and hyperparameters should be checked against the original text.
