Skip to main content
Back to timeline
NatureSource publication:

Drug Firms' Private Data Supercharge AI Protein Models

Synopsis

An AI Structural Biology (AISB) network of pharmaceutical companies fine-tuned the open-source model OpenFold3 on 20,167 proprietary protein-ligand structures from five companies; on a held-out test set of 1,056 protein-ligand structures, the model achieved high-accuracy predictions for more than half, compared with one-third for the public version of OpenFold3 and around 40% for the competing open-source model Boltz-2, and it also outperformed models trained only on individual companies' siloed data, indicating that pooling private data markedly improves protein-ligand interaction prediction.

Source-provided article image: Drug firms' secret data supercharge AI protein models.

Artificial-intelligence-based models of protein structure could be improved by incorporating data from pharmaceutical companies. Credit: Miyako Nakamura/Getty

PubMed

Interpretation

Fine-tuning OpenFold3 on 20,167 proprietary protein-ligand structures from five pharmaceutical companies enabled the model to predict more than half of 1,056 held-out test structures to a high level of accuracy. Previously, AlphaFold-like models relied mainly on public data from the Protein Data Bank (PDB), which contains relatively few experimentally determined structures of proteins bound to drug-like molecules—maybe just 10,000—and private data had not been used to train such models. The result comes from a non-peer-reviewed study described in a blog post by the AISB network, and the model is not publicly available; the test set comprises 1,056 held-out protein-ligand structures, compared against the public version of OpenFold3 and Boltz-2.

The public version of OpenFold3 achieved the same high-accuracy performance on just one-third of the structures, and Boltz-2 achieved around 40%, both lower than the AISB model. This provides a direct performance comparison under the same test conditions between a public-data model and a private-data-enhanced model. The comparison data come from the same blog post, with a test set of 1,056 structures, but the study has not been peer-reviewed.

The AISB model also outperformed co-folding tools trained only on each company's individual data, highlighting the benefits of pooling information across companies. Previously, individual pharmaceutical companies' data were siloed; this result is the first to show gains from cross-company data pooling for protein-folding models. This conclusion was stated by John Karanicolas, head of computational drug discovery at AbbVie, in the report, based on internal comparisons within the AISB network; specific numerical values are not given in the text.

Perspective

The result applies to protein-ligand interaction prediction in drug discovery, especially when public data are insufficient to cover the target chemical space. It shows that pooling proprietary structural data from multiple pharmaceutical companies can improve model performance, providing a rationale for public dataset projects like OpenBind. For teams able to access or generate large proprietary protein-ligand structure datasets, this approach may be directly applicable.

The study has not been peer-reviewed and the model is not publicly available, so its performance claims cannot be independently verified. The blog post does not provide detailed training procedures, definitions of evaluation metrics, or statistical significance tests. Additionally, the total size of the proprietary data vaults is unknown, and the quality and diversity of data from different companies are not disclosed. Readers may watch for the subsequent peer-reviewed paper and whether public dataset projects like OpenBind can reproduce similar performance gains.

Sources