Skip to main content
Back to timeline
bioRxivSource publication:

The ModelSEED Biochemistry Database, 2026 update: grading multi-source thermodynamics

Synopsis

This update expands the ModelSEED Biochemistry Database to roughly 46,000 compounds, 56,000 reactions and 37,000 metabolic structures, widens thermodynamic handling from two sources to four (group contribution, eQuilibrator 3.0, dGPredictor and experimental values) each kept with its own uncertainty, assigns gold, silver or bronze evidence grades to about 33,000 reactions, and releases reaction directions predicted by an ensemble of large language models alongside a new conflict-resolution pipeline that documents structural choices across sources.

AI-generated editorial illustration: The ModelSEED Biochemistry Database, 2026 update: grading multi-source thermodynamics

Interpretation

The database grew and its source composition changed: compounds rose 34% and reactions 55%, with reaction growth now split across MetaCyc, KEGG and the newly integrated Rhea, which contributes 8,339 reactions no other primary source supplies. Compared with the 2020 release dominated by two sources, this version adds Rhea reactions and makes each source's unique share and completeness explicit (90% of MetaCyc, 76% of KEGG and 66% of Rhea reactions can be assigned complete structures). Based on counts of reactions per primary database and the share of complete reactions (Figure 1), with the note that Rhea identifies compounds through ChEBI so its completeness depends on ChEBI structures.

The thermodynamic layer now publishes multiple sources side by side, each retaining its own uncertainty and direction inference, exposing where sources agree and where they diverge. The 2020 release combined two sources and leaned toward a single value, whereas this update carries four source records and shows that reported uncertainties differ by more than an order of magnitude (medians of 0.63, 10.41 and 17.01 kcal/mol). Calibrated against experimentally anchored reactions: eQuilibrator understates its error (only 28.7% within one reported standard deviation), group contribution overstates it (73.6% within one), and dGPredictor is best scaled (root mean square 1.15, 87.9% within one).

A gold, silver and bronze evidence grading scheme combines each source's calibrated self-confidence with cross-source agreement into an interpretable reliability label. Rather than taking reported uncertainties at face value, the scheme first calibrates uncertainty into a probability by isotonic regression and then uses a weighted comparison to judge whether sources corroborate or contradict one another. 33,099 reactions were graded (3,434 gold, 18,388 silver, 11,277 bronze); when re-graded on predictors alone, the gold subset of 806 anchored reactions falls within 2 kcal/mol of measurement 96.9% of the time versus 91.3% for silver.

An LLM ensemble (three predictors, one auditor, one adjudicator) predicts reaction direction from reaction name and stoichiometry alone as a complementary view beyond thermodynamic rules. The predictions cover about 46,000 reactions and abstain on roughly 2%, far more coverage than the thermodynamic rules commit to, and they take no part in evidence grading. The authors report that the ensemble's returned confidence does not distinguish correct calls relative to thermodynamics and is not appropriate as a downstream filter; its outputs and prompts are stored in the repository for review.

Perspective

The database targets general biochemistry across kingdoms of life, with all values reported at pH 7.0, ionic strength 0.25 M, pMg 3.0 and 298.15 K; transport reactions are scored from stoichiometry and energy alone, without membrane potential or pH gradient estimates. The grades are meant for settings that need to judge the reliability of reaction direction, such as genome-scale reconstruction, flux analysis and pathway design, with gold and silver recommended for metabolic models and bronze reserved for review.

The three predictors report uncertainties on scales that are not directly comparable, and grading depends on calibration against experimental anchors, so the meaning of a grade for compounds and reactions outside the anchored set warrants care. The LLM ensemble's confidence has not been shown to distinguish correct calls, leaving its relationship to thermodynamic rules an open question. Structure conflict reports still list about 1,000 compounds with differing structures and 56 compounds with differing elemental composition, whose effect in specific models is worth watching. In addition, the loaded text is a preprint in which figures appear as textual descriptions, so checking exact value distributions would require the original figures.

Sources