Skip to main content
Back to timeline
arXivSource publication:

DiffGCMS infers molecular structures directly from GC-EI-MS spectra via discrete graph diffusion, with LLM repair and reranking of candidates

Related research and updates

Synopsis

The work presents DiffGCMS, a spectrum-conditioned discrete graph diffusion model for de novo structure elucidation from GC-EI-MS, together with a two-stage framework in which DiffGCMS first generates candidate molecular structures from input spectra and a large language model then uses mass spectral information to validate, repair, and rerank the candidates while providing interpretable analysis of fragment-ion peaks; on a test set of 13,696 spectra from NIST 20 the generative model reached Acc@1 of 6.01% and Acc@10 of 15.76%, and on the subset of molecules with no more than 10 heavy atoms LLM-assisted molecular graph repair and reranking raised Acc@1 from 21.28% to 21.95%, Acc@10 from 46.91% to 47.99%, and candidate validity from 91.04% to 100%.

Source-provided article image: Explainable Molecular Structure Inference from GC--MS with Diffusion Models and LLM Reranking
Figure 1 ·

Figure 1: Overall two-stage workflow for molecular structure inference from GC–EI–MS. The query spectrum is encoded as the conditioning vector y y , and the molecular formula determines the heavy-atom set X X . DiffGCMS progressively denoises the bond-adjacency state from A T A^{T} through A t A^{t} to the predicted clean molecular graph A 0 A^{0} . Ten sampling runs are performed for each spectrum to generate 10 candidate SMILES. In the second stage, DeepSeek uses bidirectional SMILES–MoleCode conversion to inspect, retain or revise, and rerank the candidates, returning ranked SMILES together with mass-spectral explanations.

arXiv

Interpretation

DiffGCMS is a spectrum-conditioned discrete graph diffusion model that generates candidate molecular structures directly from GC-EI-MS spectra without relying on reference spectral library matching. Conventional methods rely heavily on reference spectral library matching, which limits identification of compounds absent from those libraries and inference of complete molecular structures directly from fragmentation information; this work formulates structure elucidation as spectrum-conditioned discrete graph diffusion generation. On a test set of 13,696 spectra from NIST 20, the generative model achieved Acc@1 of 6.01% and Acc@10 of 15.76%.

A two-stage framework integrates DiffGCMS with second-stage LLM reasoning, in which the LLM uses mass spectral information to validate, repair, and rerank candidates and provides interpretable analysis of fragment-ion peaks. Spectrum-aware postprocessing is added beyond the generative model so that candidate molecular graphs are repaired and reranked, with traceable evidence supporting the final ranking. On the test subset containing molecules with no more than 10 heavy atoms, LLM-assisted molecular graph repair and reranking increased Acc@1 from 21.28% to 21.95%, Acc@10 from 46.91% to 47.99%, and candidate validity from 91.04% to 100%.

The framework can generate plausible molecular structures for compounds absent from reference spectral libraries and provide traceable evidence supporting its decisions. Relative to conventional library-matching pipelines, the framework targets compounds outside reference libraries and presents ranking rationale as auditable, traceable explanations. The abstract reports structure generation for compounds absent from reference spectral libraries and traceable explanations, with candidate validity changing from 91.04% to 100% as a quantitative indication of the postprocessing effect.

Perspective

The results address structure elucidation for volatile and semivolatile compounds by GC-EI-MS, in settings where users want candidate structures and their ranking rationale without relying on reference spectral library matching. For practitioners, this means candidate molecular structures can be obtained for compounds absent from reference libraries, together with interpretable analysis of fragment-ion peaks and traceable evidence; the two-stage design also suggests the generative model and the spectrum-aware postprocessing can be evaluated and replaced separately. The quantitative results reported in the abstract center on the NIST 20 test set and the subset with no more than 10 heavy atoms, so the direct scope of application follows that data and molecular-size setting.

At the abstract level it remains unclear how the diffusion model represents and denoises graphs, what form of mass spectral information the LLM uses for repair and reranking, and how candidate validity reaching 100% is determined; the marked difference between metrics on the full test set and the subset with no more than 10 heavy atoms is worth examining in the full text. The abstract also does not describe validation across spectral libraries or instrument conditions, which are directions a careful reader may watch.

Sources