Skip to main content
Back to timeline
Journal of Chemical Information and ModelingSource publication:

In-Context Learning Meets Small Molecule Property Prediction: Benchmarking Novel Machine Learning Approaches

Synopsis

This study comprehensively benchmarked tabular foundation models (TFMs) based on in-context learning for small organic molecule property prediction, comparing several TFMs with multiple machine learning methods across 11 data sets (regression, random and structure-aware splits, up to 10,000 molecules each), and found that TFMs consistently outperform XGBoost, CatBoost, multilayer perceptrons, and other descriptor-based methods even with careful hyperparameter selection for the baselines, achieve accuracy on par with or better than graph-based methods including those pretrained on chemical data, with Uni-Mol2 slightly outperforming TFMs in some experiments, while retrieval (selecting the 500 closest neighbors by Tanimoto similarity for each test molecule) yields further improvement on relat

AI-generated editorial illustration: In-Context Learning Meets Small Molecule Property Prediction: Benchmarking Novel Machine Learning Approaches.

Interpretation

TFMs consistently outperform XGBoost, CatBoost, multilayer perceptrons, and other descriptor-based methods across 11 small-molecule property prediction data sets, even with careful hyperparameter selection for the baselines. Prior evaluations of tabular foundation models were mainly in other tabular domains; this study systematically brings them to small-molecule property prediction and compares them directly with classical descriptor-based methods. Covers 11 data sets, regression tasks, random and structure-aware splits, up to 10,000 molecules each, with careful hyperparameter selection for baselines, making the comparison conditions relatively stringent.

TFMs demonstrate accuracy on par with or better than graph-based methods, including those pretrained on chemical data, while Uni-Mol2, a pretrained deep neural network operating on 3D atomic coordinates, slightly outperforms TFMs in some experiments. Places TFMs that use no chemical pretraining and only a minimalistic set of 2D molecular descriptors in the same benchmark as graph-based methods that use chemical pretraining and 3D coordinates. The comparison deliberately disfavors TFMs because they do not use any chemical pretraining and rely on a minimalistic set of 2D molecular descriptors without feature engineering, making their performance more notable.

Retrieval further improves results: for each test molecule, the 500 closest neighbors by Tanimoto similarity are selected from the training set, and TFM inference is performed on this local subset, yielding improvement on relatively large data sets and random splits. Introduces retrieval-based local inference into the TFM workflow for small-molecule property prediction as a complementary means of improving accuracy. The improvement appears on relatively large data sets and random splits, indicating that its effect depends on data scale and split type.

Perspective

This benchmark targets small organic molecule property prediction, covering regression tasks, random and structure-aware splits, and up to 10,000 molecules per data set, comparing against descriptor-based methods such as XGBoost, CatBoost, and multilayer perceptrons as well as graph-based methods including those pretrained on chemical data. In this setting, TFMs use no chemical pretraining and rely on a minimalistic set of 2D molecular descriptors without feature engineering; retrieval improvement appears on relatively large data sets and random splits. The results therefore apply to this molecule class, data scale, and split setting, providing a starting point for subsequent evaluation on larger scales or other molecular systems.

A careful reader may still watch: how TFMs perform on larger data sets beyond 10,000 molecules; how retrieval behaves under other similarity measures or neighbor counts; whether differences between TFMs and graph-based methods remain stable under structure-aware splits; and under which specific conditions Uni-Mol2 is slightly better. Because the current reading scope is summary-level and does not include figures or per-data-set results, these details cannot be confirmed from the available text and remain open questions for consulting the original.

Sources