Skip to main content
Back to timeline
Nature NewsSource publication:

AlphaFold database adds more than 8,000 viral protein dimers and launches a pandemic preparedness portal

Synopsis

On 24 September, researchers added more than 8,000 predicted viral protein dimers from 23 virus families to the AlphaFold Protein Structure Database as part of a new pandemic preparedness portal; the structures were generated with AlphaFold2 from an analysis of 41,774 proteins from about 2,800 viruses, and 2,749 homodimers and 5,279 heterodimers were judged accurate enough to be included.

AI-generated editorial illustration: AlphaFold 'goes viral': database adds protein complexes of common viruses

Interpretation

Viral protein complexes entered the AlphaFold database at scale for the first time: more than 8,000 viral protein dimers covering 23 virus families with human-infecting members, launched as a pandemic preparedness portal. The database had already added about 1.7 million interacting protein pairs from 20 widely studied organisms including humans, mice and tuberculosis-causing bacteria, but viruses had remained a blind spot; this extends complex predictions to viruses. The report gives explicit figures: more than 8,000 dimers, 23 virus families, and a 24 September launch date; the database is maintained by EMBL's European Bioinformatics Institute and has more than three million users.

The workflow starts with sequence definition: researchers at the Swiss Institute of Bioinformatics identified accurate sequences for thousands of viral proteins cut from polyproteins, and researchers at various organizations worldwide then analysed the sequences of 41,774 proteins from around 2,800 viruses, including those causing mpox, measles and hepatitis B. The report notes that because some viral RNA is first translated into a polyprotein and then cut up, it is not always clear from the genetic sequence where a viral protein begins and ends, which can lead to incomplete structure predictions; resolving cut boundaries before predicting addresses that difficulty. The report quotes Joe Grove, a molecular virologist at the University of Glasgow, on the polyprotein-cutting problem, and gives the concrete scale of 41,774 proteins from around 2,800 viruses.

Predictions far outnumber inclusions: AlphaFold2 was used to predict structures for 40,746 homodimers and nearly 1.7 million heterodimers, but only 2,749 homodimers and 5,279 heterodimers were deemed accurate enough for the database, while all predictions were made publicly available. This separates making predictions from meeting inclusion criteria, showing that database entries are an accuracy-filtered subset while the full prediction set is released separately. The report gives paired prediction and inclusion counts and states that predictions below the standard were also made public.

The report states the current boundaries: the predictions lack the sugar molecules that adorn many viral proteins and help them evade immune detection, and many viral proteins work in complexes larger than dimers, such as the three-identical-protein spike of SARS-CoV-2 and other coronaviruses and HIV's envelope-entry protein; dimer predictions for such complexes were often not accurate enough to be included. This scopes the new data to the dimer level and flags that glycosylation and higher-order assembly are not yet covered. Stated directly by Grove, who was part of the effort, with the spike protein and HIV envelope protein as examples of trimers.

Perspective

This work is aimed at researchers who need structural clues about viral proteins, especially teams studying viruses such as mpox, measles and hepatitis B, and those working on antiviral targets and vaccine-related analyses; the applicable setting is dimer-level structure prediction queries, with predictions freely available. The report also notes that predictions below the accuracy standard, while not in the database, remain accessible, so users can choose between database entries and the full prediction set.

The report does not specify the exact accuracy criteria or the distribution of entries across virus families; the absence of glycosylation and of higher-order complexes such as trimers means inferences about full viral protein behaviour from these predictions still warrant caution. In addition, no papers were attached to this evidence bundle, so all information comes from the news report and the original methodological details and data tables could not be checked.

Sources