Replacing full-Hessian supervision with random projected Hessian-vector products matches second-order accuracy while speeding each epoch by more than 24x
Synopsis
The authors introduce Projected Hessian Learning (PHL), which supervises second-order curvature in machine-learning interatomic potentials (MLIPs) through stochastic Hutchinson-trace-based Hessian-vector product (HVP) projections instead of explicit Hessian construction; on a ωB97XD/6-31G(d) dataset of reactants, products, transition states, IRC and normal-mode-sampled geometries they compare four schemes (E-F, E-F-HVP with one-column or Hutchinson probes, and E-F-H), finding that with probe vectors resampled each minibatch both HVP methods are statistically indistinguishable from full-Hessian training in energy, force and Hessian accuracy while giving more than 24x speedup per epoch, and that under the more realistic fixed-vector regime with one HVP per molecule Hutchinson projections con
Interpretation
PHL casts second-order supervision as a stochastic loss that depends only on Hessian-vector products: treating the l2 Hessian-error loss as a matrix trace and applying the Hutchinson estimator tr A ≈ v^T A v yields the unbiased approximation L_H ≈ |H̃v − Hv|²/(3N)², so the (3N)×(3N) Hessian never has to be explicitly constructed or stored. Prior Hessian-supervised training either built the full Hessian explicitly (quadratic memory and compute) or sampled a single random Hessian column per step (one-hot probing); PHL uses random probe vectors to aggregate multiple curvature directions in expectation, reducing second-derivative supervision cost to near force-level complexity. The paper derives the unbiasedness of the Hutchinson estimator and, in the Supplementary Information, gives analytic mean-squared-error expressions for both estimators: under the locality assumption that Hessian errors decay with interatomic distance, Hutchinson MSE scales as O(N) while one-hot MSE scales as O(N²).
With probe vectors resampled each minibatch, the one-column and PHL HVP methods are statistically indistinguishable in energy, force and Hessian accuracy and approach full-Hessian training; relative to the E-F baseline, NMS energy RMSE drops by about 29%, force RMSE by about 48% and Hessian RMSE by about 77%. This indicates that most of the accuracy benefit of explicit Hessian supervision can be recovered through stochastic curvature sampling without paying the full-Hessian cost; paired t-tests found no significant difference between the two estimators on any dataset or property (all p > 0.05). Results span the Test, IRC and NMS datasets, with each method trained as an ensemble of five independent random seeds; reported values are ensemble means with standard deviations, assessed by paired two-sample t-tests and Bland–Altman analysis.
In the data-limited regime with only one fixed HVP per molecule, Hutchinson-based PHL further reduces NMS energy RMSE by 6.2%, force RMSE by 5.6% and Hessian RMSE by 11.2% relative to one-column probing, and is significantly better on Hessian accuracy across all three datasets (p < 0.01). Fixed probes better reflect practical quantum-chemistry settings where only limited second-derivative information is available; here random Hutchinson vectors sample curvature directions more uniformly, whereas one-column probing is confined to a single coordinate direction whose directional bias becomes limiting in data-sparse and far-from-equilibrium regimes. Differences are supported by paired t-tests over five independently trained models: NMS energy p = 0.006, IRC and NMS force p = 0.013 and 0.006, and Hessian p < 0.01 on all three datasets; the Supplementary Bland–Altman analysis shows confidence intervals excluding zero under fixed probes.
On computational efficiency, full-Hessian training averages about 326.5 s per epoch versus about 13.6 s for one-column and 13.3 s for PHL, roughly a 24x speedup over E-F-H and only about three times the cost of standard E-F training (about 4 s per epoch); at the quantum-chemistry level a single HVP can be obtained with two force-like operations (forward-over-reverse or reverse-over-reverse automatic differentiation, or finite differences of forces along a fixed direction v), costing on the order of two force evaluations rather than a full Hessian. This moves second-derivative supervision from quadratic scaling with system size to near force-level cost, making curvature information practical both for data generation and for MLIP training. Training times were recorded per epoch on NVIDIA A6000 GPUs and averaged over multiple epochs; DFT cost scaling was compared with Gaussian16 for systems up to 100 atoms across energies, forces, full Hessians and HVPs.
Perspective
The work targets MLIP development where second-order response properties are needed, particularly reactive potential energy surfaces and data-limited second-order supervision pipelines. It lets curvature information enter the training loss without explicit Hessian construction, suiting settings where quantum chemistry can supply only limited HVPs and large-scale training pipelines that want second-order regularization at near force-level cost. The authors note that when the probe vector is fixed, reference HVPs can be obtained by finite differences of forces along v, extending curvature supervision to bulk materials properties requiring large supercells, such as surfaces and defects. The method is implemented in ANI and HIP-NN, with code and the OpenREACT-CHON-EFH dataset publicly available for reproduction and extension on similar architectures.
The experimental systems are small molecules, with a median of about 14 atoms per system, while the advantage of Hutchinson over one-column probing derives mainly from analytic scaling predictions in the large-N limit, so the benefit for condensed-phase or supercell systems of hundreds to thousands of atoms still needs empirical confirmation. The effects of hyperparameters such as the number of random probes, probes per structure K, and the loss weight λ_H on accuracy and cost are not systematically scanned in the main text. In addition, the comparison focuses on energy, force and Hessian targets; the end-to-end impact on downstream response properties such as phonons and elastic constants is not directly evaluated here.
