Skip to main content
Back to timeline
arXivSource publication:

Generalized Laplace Active Subspaces: Scalable Neural-Network Uncertainty Quantification with a Handful of Curvature Directions

Synopsis

The work proposes a low-rank Laplace approximation for neural-network uncertainty quantification within a generalized Bayesian framework, restricting the posterior to an active subspace spanned by dominant generalized Gauss-Newton curvature directions of the empirical mean loss, computing eigenpairs matrix-free via Lanczos and calibrating the prior variance by empirical Bayes; across regression and molecular-generation classification tasks, standard Bayesian scaling contracts posterior variance with data size and forces retention of many weak-curvature directions, whereas mean-loss scaling yields calibrated, coherent predictive intervals with only a few directions.

Source-provided article image: Scalable AI Uncertainty Quantification via Generalized Laplace Active Subspaces
Figure 1 ·

Figure 1: The matrix-vector product G ​ v ∈ ℝ D Gv\in\mathbb{R}^{D} with D = 2701 D=2701 , computed using the matrix-free method outlined in Section 4.1.5 , and by explicitly forming matrix G G via ( 28 ). The absolute error | G ​ v − ( G ​ v ) e ​ x ​ p ​ l ​ i ​ c ​ i ​ t | |Gv-(Gv)_{explicit}| is plotted on the right vertical axis.

arXiv

Interpretation

Introduces generalized Laplace active subspaces (GLAS): starting from a generalized Bayesian posterior defined through the empirical mean loss, it builds a local Gaussian approximation around pretrained weights, with posterior variances available in closed form in the retained subspace, prior variance calibrated by an empirical Bayes criterion, and active-subspace dimension selected by empirical coverage on held-out calibration data. Unlike variational Bayesian neural-network methods such as Bayes by Backprop that optimize a variational distribution during training, this is a post-hoc approximation requiring no additional training; unlike classical active subspaces built from an uncentered covariance of model gradients, the subspace here comes from dominant generalized Gauss-Newton curvature directions of the empirical mean loss, tying it directly to the local data-fit geometry. The method is demonstrated on three cases: a scalar regression with known noise whose weight-space dimension is small enough to explicitly form the full matrix (used to verify matrix-free vector products and eigenpairs, with absolute error at the torch.float32 level), the UCI Concrete Compressive Strength regression (with a large number of connection weights, assessing scalability), and a 3-layer LSTM molecular language model from REINVENT (256 neurons per layer, about 1.6M weights, trained on ChEMBL SMILES).

A central finding is that the scaling of the generalized posterior is not merely a technical detail: standard Bayesian scaling (corresponding to the summed negative log-likelihood) contracts the posterior variance in leading active directions as training data grows, scaling with the data size in regression and its proportion in classification, forcing the low-rank framework to retain many weak-curvature directions to reach nominal coverage. Prior low-rank Laplace subspace work (e.g., Faller and Martin) constructs subspaces in the linearized Laplace setting and still requires dimensions in the tens to hundreds; this work identifies data-size-driven posterior contraction as the source of that dimension inflation under standard scaling and avoids it with mean-loss scaling, achieving calibrated coverage with only a few curvature directions. In the known-noise regression, generalized scaling needed only 2 active directions to reach 90% coverage with coherent intervals, while standard scaling required more directions and produced slightly wider, less coherent intervals; with more training data, generalized scaling remained 2D while standard scaling required 43 curvature directions. In the concrete-strength case, generalized scaling met the target at a small active dimension, whereas standard scaling never reached it within the allowed range, with 90% intervals covering only about 60% of calibration data even at 50 curvature directions.

When posterior samples are propagated through the nonlinear network, the extra weak-curvature directions required by standard scaling introduce larger parameter perturbations that can move samples outside the local region where the Laplace approximation is reliable, degrading the predictive mean relative to the pretrained model and widening or destabilizing predictive intervals; linearized Laplace removes this nonlinear sampling artifact, but the excessive posterior contraction under standard scaling persists. The analysis links low-rank posterior dimension choice to predictive coherence under nonlinear propagation, showing that in the full nonlinear-network setting the size and coherence of parameter perturbations are central, and that generalized mean-loss scaling reduces both nonlinear sampling artifacts and the number of curvature-vector products needed. In the known-noise regression with standard scaling and more data, mean prediction and interval quality degraded markedly; linearized Laplace improved the predictive mean by construction (centered at the expansion-point prediction) and gave substantially narrower intervals, but at the maximum allowed dimension of 50 covered only about 75% and 40% of calibration data. In the concrete case, standard-scaling intervals were judged overconfident, while generalized-scaling means were generally centered on validation data with 90% intervals encompassing most data without being overly wide.

The active subspace is highly localized in the high-dimensional weight space: activity scores derived from the active-subspace eigenmodes show that in the concrete case the 1000 most-sensitive connection weights (just 0.1% of the total) rapidly saturate and account for most curvature sensitivity, with the remaining 99.9% of weights contributing about 5% of total sensitivity; in the molecular case the first 7.6% of activity scores already account for 95% of total sensitivity. Activity scores, originally first-order sensitivity indices for physical models, are repurposed here to examine whether curvature contributions are broadly distributed across the network or concentrated in a small subset of connection weights, showing the active subspace is localized in weight space as well as low-dimensional. Results are presented as cumulative activity scores on log-log plots (Figure 8 for concrete, Figure 12 for molecules); the molecular curve saturates less quickly but still shows the first 7.6% of activity scores accounting for 95% of total sensitivity.

Perspective

The method targets pretrained neural networks trained with an empirical mean loss and applies to post-hoc uncertainty quantification for regression and classification; in regression, the active-subspace dimension and prior variance are chosen jointly via empirical coverage on held-out calibration data, while in classification calibration entails ensuring the correct class is contained with approximately a given probability. Its design goal is to concentrate uncertainty in a few data-informed curvature directions, yielding calibrated and sharp predictive intervals while retaining the full nonlinear network. The authors identify extending this generalized-Bayes active-subspace Laplace approximation to fine-tuned language models (as in the Laplace-LoRA setting) as future work, which would shift the role of the Laplace approximation from probing generative sensitivity toward correcting overconfidence in fine-tuned autoregressive models.

Several open questions remain for a careful reader: standard scaling and generalized mean-loss scaling define different posterior objects, and the authors explicitly note that standard scaling is not intrinsically defective, since posterior contraction with increasing data is an expected feature of the standard Bayesian update, so the scaling choice should be judged against the specific low-rank UQ goal. In the molecular-generation case, generalized scaling loses validity as the active dimension grows, sequence lengths develop a long right tail, and the probability of SlogP above 5 rises significantly; the authors interpret this as perturbations pushing the autoregressive dynamics outside the chemically meaningful region, but the case is positioned as a sensitivity study of free-running molecular generation, and the pretrained REINVENT model is already reasonably calibrated at the token level, so the Laplace posterior is not expected to improve its predictive distribution. In addition, the molecular case computes eigenmodes from subsampled SMILES strings, and the authors observe an obvious outlier at one subsampling size, most likely from a Lanczos run that failed to converge, suggesting a replica-eigenvector inner-product matrix as a flag; the effect of this subsampling approximation on the final conclusions merits attention. This is a full-text reading, but some equations and figures are referenced by number without their numerical content being expanded in the text, so readers wishing to reproduce results should consult the authors' public PyTorch code and data.

Sources