Skip to main content
Back to timeline
NVIDIA ResearchSource publication:

NVIDIA, Google DeepMind and EMBL-EBI openly released predicted protein-complex structures for more than 2,800 viruses, about 30% of them interactions never previously documented

Synopsis

NVIDIA joined a coalition including Google DeepMind and EMBL-EBI to infer and openly release predicted 3D structures of protein complexes for more than 2,800 viruses using AlphaFold2 with the NVIDIA BioNeMo inference runtime, deposit them in the AlphaFold Database, and open-source the BioNeMo Structure Prediction Pipeline used to generate them, with about 30% of the protein interactions being entirely new to science.

AI-generated editorial illustration: How Open Science Can Help Researchers Prepare for the Next Pandemic

Interpretation

The coalition openly released predicted 3D structures of protein complexes for more than 2,800 viruses, available to any scientist through the AlphaFold Database. Structural knowledge had been absent for thousands of viruses; this dataset begins to fill that gap and extends coverage to lesser-studied viruses. Structures were inferred with AlphaFold2 and optimized with the NVIDIA BioNeMo inference runtime to scale to thousands of viral proteomes; the text states predictions are labeled by confidence and that high-confidence predictions can be verified experimentally.

About 30% of the protein interactions added to the database are completely new to science, showing interaction shapes never documented in the Protein Data Bank. It offers the biological community previously undocumented interaction shapes, described as an engine for hypothesis generation. The text grounds this proportion in a comparison against the Protein Data Bank and does not detail further statistical methodology.

NVIDIA openly released the BioNeMo Structure Prediction Pipeline, the GPU-accelerated workflow used to generate the dataset, so researchers can go from protein sequence to predicted 3D structure for their own targets. It makes the generating workflow itself public, not only the resulting data. The text identifies the pipeline as the workflow used to generate the dataset and points to an access route.

The dataset is positioned as an openly accessible foundational resource that lowers barriers especially for scientists in low-resource settings confronting outbreaks. Compared with structure resources serving only a few institutions, open access plus coverage of lesser-studied viruses widens the potential user base. This comes from statements by people at EMBL-EBI and Google DeepMind and is a project-positioning statement rather than an independent evaluation.

Perspective

The dataset is meant for researchers who need structural information on viral protein complexes, including vaccine and drug design, viral diagnostics and fundamental biology, as well as researchers in low-resource settings confronting outbreaks firsthand. Its purpose is to stockpile structural knowledge in advance so scientists do not start from zero when an unknown pathogen appears; predictions are labeled by confidence, high-confidence results can be further verified experimentally, and researchers can use the open BioNeMo Structure Prediction Pipeline to generate structures for their own targets.

The text gives no verification scale, accuracy figures or systematic comparison against experimental structures, and does not specify how the 30% figure was computed; these are questions a reader would still watch when judging how much downstream conclusion the dataset can directly support. It also notes that high-confidence predictions can be verified experimentally but does not say whether such verification has been carried out.

Sources