Skip to main content
Back to timeline
PRX IntelligenceSource publication:

SMARTERS uses an Attention U-Net to predict planar molecular structures directly from simulated TERS hyperspectral images, reaching a mean test Dice of 0.842

Synopsis

The work introduces SMARTERS, an Attention U-Net encoder-decoder trained on simulated TERS hyperspectral images of 1,840 planar small molecules, which maps vibrational spectral images directly to 2D atomic position maps (mean test Dice similarity coefficient 0.842) and to chemically resolved maps per element (H, C, N, O; mean test 0.810), showing that automated molecular structure identification from TERS images is feasible on simulated data without conventional manual comparison and case-by-case quantum-chemistry calculations.

AI-generated editorial illustration: Automated Structure Discovery for Tip-Enhanced Raman Spectroscopy

Interpretation

The model regresses 2D atomic position maps directly from simulated TERS hyperspectral cubes, with a mean test DSC of 0.842 (train 0.862, validation 0.833) and consistent behavior across the three splits. Previously, comparison of TERS images with simulations relied mainly on case-by-case visual inspection; the authors state that automated molecular structure discovery from TERS hyperspectral images had not been demonstrated before this work, which formulates the task as image-to-image translation and trains it end to end. Trained on simulated data for 1,840 planar molecules split 80/10/10, for 50 epochs with the Adam optimizer and Bayesian hyperparameter search via Optuna; in the sparse atom-map setting DSC is strongly sensitive to sub-angstrom displacement (Appendix C shows DSC drops rapidly when a predicted map is translated), so a high score reflects localization accuracy rather than merely a correct atom count.

The same architecture outputs four element channels (H, C, N, O), jointly predicting atomic position and elemental identity, with mean test DSC 0.810, broken down as H 0.856, C 0.827, N 0.759 and O 0.799. Compared with single-channel structure prediction that only marks the presence of an atom, this task folds chemical identity into the same prediction, which the authors describe as going beyond simpler classification-style tasks. Per-element DSC is averaged over the four channels to offset unequal elemental abundance; however, H and C are far more abundant than N and O in the dataset, nitrogen attains the lowest median DSC and the widest spread, and the gap between best and worst element on the test set is roughly 0.10 in DSC, a trend that tracks the relative abundance of each element in the dataset.

Molecular planarity dominates performance: training separate models on datasets with progressively relaxed RMSD thresholds of 0.05, 0.1 and 0.5 angstrom lowers median test DSC from 0.842 to 0.803 to 0.771 while widening the performance distribution (standard deviation rising from 0.112 to 0.150). The authors name this effect planarity-induced variability (PIV) and note that although a larger sample would normally reduce variance, the spread instead grows here, indicating that out-of-plane distortion weakens or obscures near-field contributions and makes the structure-to-signal mapping less deterministic. Three datasets were independently trained, validated and tested, and the trend is consistent across training, validation and test splits; on this basis the authors recommend that RMSD thresholds for TERS machine-learning dataset curation balance structural diversity against the physical limits of near-field sensitivity.

On experimental TERS data for FePc/Ag(110), the model does not produce a usable structure: on raw experimental data it predicts several atomic positions that do not resemble the phthalocyanine structure, and applying Gaussian denoising greatly reduces the number of incongruent positions but leaves only a few placed at random locations. The authors attribute this to the gap between simulation and experiment, including the relatively sharp 5 angstrom FWHM tip assumed in simulation, the absence of substrate effects, and the unknown influence of tip changes during acquisition, and they state plainly that the current model is not yet capable of routinely handling experimental data. This is an ablation in Appendix D using previously published experimental data (FePc adsorbed on Ag(110), 300-2000 cm-1) against simulated data spanning 0-3200 cm-1; the authors also note a large mismatch between simulated and experimental TERS spectral sums, and that Fe is absent from the training dataset, which limits prediction at the molecule center.

Perspective

The result targets isolated planar molecules lying flat on weakly interacting substrates such as NaCl, with the tip plasmon modeled as a 3D Gaussian of 5 angstrom FWHM, a simulation field of view of 18x18 square angstroms rasterized at 256x256 pixels, and frequency-axis binning over 0-4000 cm-1 (100 channels finally used, i.e. 40 cm-1 per bin). Within this setting the model can rapidly produce 2D atomic position maps and element-resolved maps from simulated TERS hyperspectral images, offering a methodological starting point for automated scanning probe platforms; the authors also frame the work as a proof of concept and suggest moving to 3D molecular graphs with graph neural networks for non-planar systems, curating datasets with greater elemental and functional-group diversity to improve N and O identification, and using augmentations closer to experiment to narrow the simulation-to-experiment gap.

The authors list three open directions in the conclusions: non-planar systems (subsurface atoms do not contribute directly to the TERS signal but influence vibrational modes through their bonds), the reliability of rarer elements such as N and O limited by the chemical imbalance of the dataset, and bridging the simulation-to-experiment gap. Appendix D shows that prediction fails on experimental data, so behavior on real measurements remains an open question; in addition, vibrational selection effects mean modes oriented parallel to the surface contribute weakly, so some structural regions may be absent from the signal, which introduces ambiguity in defining the ground-truth image itself. Coordinate-level metrics depend on a Laplacian-of-Gaussian binarization threshold, so the authors keep the threshold-agnostic DSC as the primary measure and report threshold-dependent metrics only as corroboration.

Sources