Skip to main content
Back to timeline
Nature NewsSource publication:

Drug companies' private protein structures markedly improve AI folding models

Synopsis

A consortium of pharmaceutical companies called the AI Structural Biology (AISB) Network fine-tuned OpenFold3, previously trained only on PDB data, on 20,167 protein–ligand structures from five companies, and on a held-out test set of 1,056 protein–ligand structures the model reached high accuracy on more than half of them, versus about one-third for the public OpenFold3 and around 40% for the open-source model Boltz-2, while also outperforming models trained only on each company's own data, indicating that pooling data is more valuable than keeping it siloed.

AI-generated editorial illustration: Drug firms’ secret data supercharge AI protein models

Interpretation

Fine-tuning OpenFold3 with pharmaceutical companies' unpublished protein–ligand structures noticeably improves prediction accuracy for protein–small-molecule binding structures. Previously AlphaFold-family models and OpenFold3 relied mainly on the public Protein Data Bank (PDB), which contains relatively few experimental structures of proteins interacting with drug-like molecules; this work brings companies' internal structures into training to address that gap directly. Evaluated on 1,056 protein–ligand structures held out from training, the AISB model reached high accuracy on more than half, versus about one-third for the public OpenFold3 and around 40% for Boltz-2; however, the study is described in a blog post, has not been peer-reviewed, and the model is not publicly available.

Pooling data across multiple companies works better than training separately on each company's own data. This provides a direct empirical contrast between data silos and data sharing, rather than only a theoretical argument for sharing. The article quotes John Karanicolas, head of computational drug discovery at AbbVie, saying the AISB model also outperformed co-folding tools trained only on each company's individual data; specific comparison numbers are not given in the text.

The result is used to argue for building comparable public datasets to advance protein-folding and co-folding AI. It turns the outcome of one inter-company collaboration into a case for investment in public data infrastructure, pointing to the already-running OpenBind project. The article notes OpenBind is supported by up to £8 million (about US$10.8 million) in UK government funding and released hundreds of new protein structures last month, with thousands more in the works; this is a statement about related activity, not direct evidence of the model's performance.

Perspective

The result concerns prediction of protein–small-molecule ligand binding structures and applies to settings such as pharma consortia or similar alliances that hold large internal structural datasets; its value lies in showing that supplementing public PDB data with proprietary structures can improve co-folding model performance, and in providing an argument for public data projects such as OpenBind. For researchers without such data sources, near-term direct benefit depends on whether public datasets can reach comparable scale and diversity.

Readers will still watch for: once the formal paper appears, how robust the performance gain is across target classes and ligand chemical space; whether there is distribution overlap between the held-out test set and the training data; how much of the gain from pooling multiple companies' data comes from data volume versus data diversity; and whether the comparison can be reproduced externally given that the model is not public. The article does not provide these details, so the finding is best read as a promising early signal rather than established consensus.

Sources