A CNN trained only on synthetic multipole-Padé spectra infers plasmonic poles in one forward pass and transfers to first-principles and experimental dielectric spectra
Synopsis
The work builds a one-dimensional convolutional neural network trained entirely on synthetic dielectric spectra generated by randomly sampling the analytical multipole-Padé expression, so that it directly inverts a spectrum into pole parameters (energy, broadening, spectral weight) without material-specific data; on synthetic validation spectra, RPA first-principles spectra (Al, Ca, Ti) and experimental spectra (Si, ZnO, Ta, Mo), a single forward pass reconstructs the dominant spectral structure, and an imaginary-part-only input still recovers the pole representation.
Figure 1: Schematic representation of the CNN architecture used to infer the MPA representation of the dielectric spectra. The input consists of the frequency grid and the real and imaginary components of the dielectric response, { ω i , ℜ 𝔢 𝔜 ( ω 𝔦 ) , ℑ 𝔪 𝔜 ( ω 𝔦 ) } 𝔦 = 1 𝔫 ω \{\omega_{i},\mathgothic{Re}\,Y(\omega_{i}),\mathgothic{Im}\,Y(\omega_{i})\}_{i=1}^{n_{\omega}} , sampled on n ω = 1000 n_{\omega}=1000 frequency points. The convolutional part of the network comprises four one-dimensional convolutional layers with kernel sizes k = 101 k=101 , k = 21 k=21 , k = 7 k=7 , and k = 3 k=3 , respectively, with GELU activations and max-pooling operations after the first three convolutional layers. An adaptive average-pooling operation reduces the resulting feature representation to 128 × 50 128\times 50 , which is flattened and passed through a fully connected layer with 128 neurons and a GELU activation. The final linear layer outputs the 3 n p 3n_{p} MPA parameters { ℜ 𝔢 Ω 𝔭 , ℑ 𝔪 Ω 𝔭 , 𝔯 𝔭 } 𝔭 = 1 𝔫 𝔭 \{\mathgothic{Re}\,\Omega_{p},\mathgothic{Im}\,\Omega_{p},r_{p}\}_{p=1}^{n_{p}} . The spectra shown at the left and right illustrate the input response and its corresponding reconstruction from the predicted MPA parameters, respectively; the upper and lower panels show their real and imaginary components.
arXivInterpretation
A one-dimensional CNN is trained to learn the inverse mapping from a dielectric response to its multipole-Padé (MPA) pole representation, outputting complex poles and real positive weights, with physical constraints embedded in the output parameterization (pole imaginary parts kept closer to the real-frequency axis, weights positive). Previously the MPA representation was constructed from limited frequency samples or by nonlinear fitting of the full-frequency response; here the inverse mapping itself is the learning target, giving pole parameters in a single forward pass and replacing a fitting procedure sensitive to initialization and fitting choices. On synthetic validation spectra the reconstructed responses closely match the targets; for the dataset with several overlapping poles the predicted parameters agree closely with the targets with small absolute errors; training and validation losses fall by orders of magnitude over up to 300 epochs and approach saturation while remaining close.
The training set is generated entirely from the analytical MPA expression with randomized parameters, with no material-specific spectral database, so the learned mapping transfers to first-principles and experimental spectra outside the training distribution. Unlike predicting a dielectric function from material information or training on material-specific data, this uses a material-unbiased synthetic ensemble to learn the general relation between a spectral response and its pole structure. The synthetic ensemble spans up to a number of poles on a homogeneous frequency grid over a 100 eV interval, with a Rayleigh-like criterion excluding pole pairs whose energy separation is too small; training and validation are split 80% versus 20%. Transfer tests cover RPA spectra of Al at two momentum transfers, Ca and Ti, and experimental spectra of Si, ZnO, Ta and Mo.
Inversion difficulty grows with the number of overlapping excitations: mean, median and 95% absolute errors of the pole parameters increase monotonically with the number of poles, while the reconstructed-response error need not grow comparably because weak-weight poles contribute little to the total spectrum. Separating parameter error from response-reconstruction error shows that dominant spectral structure can still be recovered accurately even when small-weight poles have larger position or broadening errors. Table 1 reports mean, median and 95% absolute errors for all four models across pole numbers; the four models show similar trends, errors increase monotonically with pole number, and differences between noisy and clean training are generally smaller than the increase tied to spectral complexity.
An imaginary-part-only input is sufficient to recover the MPA structure, offering a route for experiment: once pole parameters are determined from the measurable imaginary part, the corresponding real part is accessible without explicitly performing a Kramers-Kronig transformation over an extended frequency interval. The full-complex input is generally slightly more accurate, indicating the real part carries complementary information, but the imaginary-only models perform relatively closely, showing most of the information needed to infer MPA parameters is contained in the imaginary component. Comparison of the four models (full-complex versus imaginary-only, each with and without noise) shows the full-complex models generally achieve lower parameter errors, while the imaginary-only models remain relatively close; with noise, the full-complex model shows a somewhat weaker increase in mean error, particularly for pole positions.
Perspective
The framework targets settings that need a compact pole representation extracted from a dielectric response, such as high-throughput screening and automated spectral analysis, for users interpreting many-body perturbation-theory calculations and electron-energy-loss, optical or absorption spectra. It applies to spectra dominated by a finite set of discrete poles: the synthetic training set spans controlled excitation energies, broadenings, spectral weights and multipole complexity, with a Rayleigh-like criterion imposed for multiple poles. The authors note the synthetic ensemble can be extended to include additional spectral structures, closer and more strongly overlapping poles, larger numbers of poles, and experimentally relevant forms of noise and background variation, which would require larger training sets and, where necessary, greater network capacity; increasing the number of poles requires only enlarging the final output layer.
The experimental part is framed by the authors as a proof of feasibility rather than a representation optimized for all experimental spectral features: the ZnO spectral onset, two weaker shoulder structures around the main Ta peak, secondary structures in Mo and ZnO, an incomplete low-energy tail, and experimental noise differing from the training perturbations all lie beyond the current synthetic ensemble. Pole-parameter errors grow with the number of poles, and weak-weight poles carry larger position and broadening uncertainty, so they should be assessed together with response-reconstruction error. The sampling ranges, the Rayleigh-like criterion threshold and the noise-residue bound are given as equations in the text; this parse did not retain the specific numerical values inside those equations, so their exact values remain an open question to confirm against the original.
