Skip to main content
Back to timeline
arXivSource publication:

SoftServe builds positive-definite curvature estimates from a variational objective and reaches lower losses than Adam, Muon, and SOAP on ill-conditioned deep learning tasks

Related research and updates

Synopsis

The authors introduce SoftServe, a family of quasi-Newton methods that derives positive-definite curvature estimates from the variational objective of Berglund et al. (2025), remaining positive definite even under negative curvature, and offers diagonal and Kronecker-factored variants that preserve positive definiteness by construction; it relies on the stable coupled Newton-Schulz iteration to replace costly matrix decompositions with GPU-friendly matrix multiplications, and on severely ill-conditioned problems such as recurrent networks, deep autoencoders, physics-informed neural networks, and a 136M-parameter physics-informed diffusion model it often achieves lower losses than established baselines including Adam, Muon, and SOAP.

Source-provided article image: SoftServe: A Scalable Quasi-Newton Method for Deep Learning
Figure 1 ·

Figure 1: Training objectives against gradient evaluations: prediction MSE for RNN Adding (left), and reconstruction binary cross-entropy with ℓ 2 \ell_{2} regularization for full-batch MNIST (middle) and minibatch MNIST (right). Lines show three-seed means; shading spans the seed minimum and maximum. K-BFGS(L) with failed seeds are shown as individual traces.

arXiv

Interpretation

SoftServe derives positive-definite curvature estimates from the variational objective of Berglund et al. (2025), remaining positive definite even in the presence of negative curvature, without line searches or ad hoc curvature corrections. Prior quasi-Newton methods in deep learning were limited by non-convexity and enormous parameter sizes and often relied on line searches or ad hoc curvature corrections; this work grounds positive-definite curvature estimation in a variational objective, sidestepping both obstacles. The abstract states this constructively, indicating that positive definiteness comes from the variational objective rather than post-hoc correction; no theorem numbers or numerical verification details are given.

The authors develop diagonal and Kronecker-factored variants that preserve positive definiteness by construction and scale to massive neural networks. It turns the positive-definite quasi-Newton update into scalable parameterizations usable for very large networks, not only small convex problems. The abstract states that the two variants preserve positive definiteness 'by construction' and 'scale to massive neural networks'; this is a design-level statement without comparative parameter counts or memory figures.

SoftServe uses the stable coupled Newton-Schulz iteration for the required matrix operations, replacing costly matrix decompositions with GPU-friendly matrix multiplications. It replaces the matrix-decomposition bottleneck of quasi-Newton methods with GPU-suitable multiplications, lowering the cost of putting the method on hardware. The abstract names the iteration and the substitution explicitly, a method-level statement; no iteration convergence rates or measured comparisons against decomposition methods are provided.

On severely ill-conditioned problems, including recurrent networks, deep autoencoders, physics-informed neural networks, and a 136M-parameter physics-informed diffusion model, SoftServe often achieves lower losses than Adam, Muon, and SOAP. It extends the demonstrated effectiveness of quasi-Newton methods from convex optimization to ill-conditioned deep learning tasks and compares directly against widely used first-order and structured baselines. The abstract gives task types, one 136M-parameter model scale, and baseline names, describing results as 'often achieving lower losses'; no per-task numerical tables or statistical significance are given.

Perspective

The work targets severely ill-conditioned unconstrained optimization in deep learning, applying to training tasks such as recurrent networks, deep autoencoders, physics-informed neural networks, and large physics-informed diffusion models; the diagonal and Kronecker-factored variants preserve positive definiteness by construction and, together with the coupled Newton-Schulz iteration, replace matrix decompositions with matrix multiplications, making them suitable for scaling to very large networks on GPUs. For practitioners who want quasi-Newton methods without line searches or ad hoc curvature corrections, this family offers candidates directly comparable against Adam, Muon, and SOAP.

The abstract does not give specific loss values, training steps, or convergence curves for each task, nor the scale of tasks other than the 136M-parameter model, so the magnitude and stability of 'often achieving lower losses' still need confirmation in the main text. The positive-definite curvature estimates come from the variational objective of Berglund et al. (2025), and the specific derivation and applicability conditions within this family of methods are part of what further reading would need to cover. The actual memory and per-step computational costs of the diagonal and Kronecker-factored variants, as well as the convergence behavior of the coupled Newton-Schulz iteration, are not elaborated in the abstract. In addition, this summary is based on the abstract only; the main text, figures, and appendices were not read, so the judgments above are limited to what the abstract states.

Sources