Skip to main content
Back to timeline
arXivSource publication:

GRACE predicts collision cross sections with geometric residual and early-fusion adduct conditioning, reaching 1.67%, 2.11%, and 2.36% mean percentage difference on random, scaffold, and adduct-sensitive splits

Related research and updates

Synopsis

GRACE is a 3D collision cross section (CCS) predictor that adapts a pretrained molecular geometry encoder through geometric residual adduct conditioning via early fusion; on over 9,000 experimental molecule-adduct CCS records it achieves the best mean percentage difference among the evaluated learned models on random, scaffold, and adduct-sensitive splits (1.67%, 2.11%, and 2.36% respectively), shows consistently lower error across four independent external test sets, and attains the lowest mean percent difference against four previously reported physics-based workflows on a held-out set.

Source-provided article image: Predicting Collision Cross Sections with GRACE: Geometric Residual Adduct Conditioning via Early-fusion
Figure 1 ·

Figure 1: CCSBench statistics ( n = 9,209 n=9{,}209 CCS measurements). (A) Distribution of CCS values by adduct type: [M+H] + , [M − - H] - , and [M+Na] + . Adducts differ systematically in CCS range and distribution shape. (B) Number of molecules measured in one, two, or three adduct forms (3,911 / 1,743 / 604), showing the adduct coverage of the dataset. (C) Maximum Tanimoto similarity (ECFP4) to the training set for molecules in the random, scaffold, and adduct-sensitive test sets. The scaffold and adduct-sensitive splits have lower similarity to training, confirming structural diversity between partitions. (D) CCS as a function of exact molecular mass; molecular mass alone explains 92.7% of variance ( R 2 = 0.927 R^{2}=0.927 ), motivating the linear Ridge baseline used in residual learning.

arXiv

Interpretation

GRACE combines two inductive biases for 3D CCS prediction: a residual objective relative to an adduct-aware physical descriptor baseline, and adduct conditioning inside the encoder via a learned adduct token and low-rank attention adapters. Most existing predictors either ignore explicit 3D structure or treat adduct identity as a late categorical feature, which limits their ability to capture adduct-dependent geometric effects; GRACE moves adduct information into the encoder and layers residual learning on top. The abstract reports evaluation on over 9,000 experimental molecule-adduct CCS records with random, scaffold, and adduct-sensitive splits designed to separate interpolation, scaffold generalization, and adduct-driven generalization.

GRACE achieves the best mean percentage difference among the evaluated learned models on all three splits: 1.67% on the random split, 2.11% on the scaffold split, and 2.36% on the adduct-sensitive split. This extends error evaluation beyond a single random split to scaffold and adduct-sensitive splits, so scaffold generalization and adduct-driven generalization are measured separately. The numbers come from the three split results reported in the abstract and are comparisons among the evaluated learned models.

Diagnostic analyses suggest that residual learning stabilizes training by removing the dominant mass-CCS trend, while early fusion improves adduct-sensitive prediction relative to late fusion. This offers a mechanism-level account of what each inductive bias contributes, beyond the aggregate error figures. The abstract phrases this as what diagnostic analyses suggest, so it is analytical evidence rather than a causal conclusion from a controlled experiment.

Across four independent external test sets GRACE shows consistently lower error than the other evaluated models, and on a held-out set it attains the lowest mean percent difference compared with four previously reported physics-based workflows. The external test sets and the physics-workflow comparison extend evaluation from internal splits to cross-dataset and cross-paradigm settings. The abstract reports consistently lower error and the lowest mean percent difference without giving per-test-set values.

Perspective

The work targets fast prediction of molecule-adduct CCS in ion mobility mass spectrometry, suited to settings where CCS serves as a descriptor for molecular annotation, for example using a learned model to replace or complement physics-based workflows under random, scaffold, and adduct-sensitive generalization conditions. Its design intent is to predict more accurately when explicit 3D structure and adduct identity both matter, while keeping error low across datasets and against other method paradigms.

The abstract describes the roles of residual learning and early fusion as what diagnostic analyses suggest, so the strength of these mechanism explanations depends on the control setup in the full text, which readers may want to examine. The four independent external test sets are reported only as consistently lower error without per-set values, leaving cross-dataset error distributions an open question. The comparison with four physics-based workflows rests on a held-out set, and its comparability conditions (such as descriptor and workflow settings) need confirmation in the full text. In addition, the abstract does not state model scale, training cost, or inference speed figures, so the practical magnitude behind the word fast awaits the full text.

Sources