RetroChimera: Improving Small-Molecule Retrosynthesis Prediction with a Learned Two-Model Ensemble
Synopsis
This work presents RetroChimera, a retrosynthesis prediction framework that combines a Transformer-based de-novo model, R-SMILES 2, with a graph-neural-network model grounded in reaction templates, NeuralLoc, through a learned, rank-dependent ensembling strategy, so that their complementary strengths are leveraged to perform strongly across both common and rare reaction classes and, in blind tests, to produce disconnections of complex molecules that PhD-level chemists preferred over those from its constituent sub-models, from more established approaches, and even from the test set itself.
Interpretation
It proposes a retrosynthesis framework built around two complementary sub-models: R-SMILES 2 predicts precursor molecules directly from the target molecule, while NeuralLoc encodes both the target molecule and reaction templates as graphs and predicts which templates to apply and where. Relative to a single modeling paradigm, this design accommodates both unconstrained flexible generation and predictions grounded in reaction patterns extracted from training data, and it characterizes their division of labor: R-SMILES 2 performs particularly well on reactions involving large changes over the course of the reaction, while NeuralLoc excels in reactions of low precedence and those involving more localized changes. The basis is the architectural description and the stated specialization across reaction types, which is a method-level design argument.
It uses a learned ensembling strategy to combine the ranked predictions of both sub-models: each model assigns a learned, rank-dependent vote to each predicted reactant set, and votes are added when both models propose the same reaction. Rather than simple selection or fixed-weight fusion, the strategy learns how much to trust each model at different ranks, thereby approximately matching the better-performing sub-model across reaction classes. The basis is the description of the ensembling mechanism and the reported behavior across reaction classes.
In blind tests, expert chemists preferred disconnections of complex molecules suggested by RetroChimera over those obtained from its constituent sub-models, from more established approaches, and even from the test set itself. It extends the evaluation criterion from comparison with recorded reactions to alignment with chemists' judgment, responding to the stated challenges of recalling rare but strategically important reactions, robustness beyond the training distribution, and aligning with chemists' expectations. The basis is a blind-test preference result involving PhD-level chemists; the text does not report sample sizes or statistics.
The implementation and weights are open-sourced on GitHub under an MIT license and accessible via Microsoft Foundry, and the work reports validation studies including zero-shot transfer and fine-tuning on proprietary datasets. It moves the method from a paper description to a resource the chemistry community can experiment with directly, and it covers transfer beyond the training distribution. The basis is the text's statements about open-sourcing, access channels, and the scope of the validation studies.
Perspective
The work targets retrosynthesis route planning starting from small-molecule targets, applicable to molecular science settings such as drug discovery and the design of smart materials; its value lies in helping chemists identify promising synthesis routes more efficiently and assess more, and more complex, candidate molecules at large scale, and the text positions it as a component that, paired with laboratory automation, moves toward closed-loop, self-improving systems. The open-sourced implementation and weights and the platform access channel let outside researchers test it directly on their own targets.
The scale and statistical basis of the expert preference in blind tests, the coverage of rare reaction classes, and the specific settings of zero-shot transfer and fine-tuning on proprietary datasets are not elaborated in this text; those details require the cited Nature publication and accompanying article. In addition, the expectation of acceleration toward closed-loop, self-improving systems is an outlook that awaits practical validation once combined with laboratory automation.
