Skip to main content
Back to timeline
SIAM Journal on Scientific ComputingSource publication:

Learning nonlinear finite element solution operators with MLPs and energy minimization yields small learning errors on parametric Poisson, Gaussian random field, and nonlinear elasticity cases, and speeds up Newton's method when the network prediction is used as an initial guess

Synopsis

The authors develop and evaluate a data-free, physics-informed method for learning solution operators: after a finite element discretization, a multilayer perceptron (MLP) maps problem data parameters (boundary conditions, coefficients, right-hand sides, etc.) to the finite element degrees of freedom, with a loss function given by the expected energy functional, plus a parallelizable training algorithm that can use random batches of mesh elements; on a parametric Poisson problem, a Gaussian random field coefficient problem, and a nonlinear neo-Hookean elastic beam, learning errors are small (e.g., for the parametric problem with the full mesh and a 4·512 network, mean relative energy error 2.0e-4, L2 0.0020, H1 0.

AI-generated editorial illustration: Learning Nonlinear Finite Element Solution Operators Using Multilayer Perceptrons and Energy Minimization

Interpretation

The target of operator learning is set to the standard finite element discrete solution rather than the exact solution, so that finite element theory can be used directly in error analysis. Unlike existing field-to-field maps (e.g., convolutional networks for linear problems) or mesh-informed networks based on the strong residual, this work uses an MLP to learn a tuple-to-field map into a finite element space, with both analysis and examples including nonlinearities. The paper presents a general framework (energy functional, discrete operator Ah, network Ah,θ), lemmas relating the energy difference to the H1 seminorm for linear and nonlinear energies, and an approximation error estimate combining the finite element error with the network training error.

An energy-functional-based loss and a parallelizable training algorithm with random batches of mesh elements are provided. The loss is the expected energy, avoiding training data generation; the energy can be assembled locally on each element and parallelized on GPUs, and for large problems only a small fraction of randomly chosen elements is used per iteration. The paper derives the element-wise decomposition of the loss and compares full-mesh versus batched training (e.g., 3277 elements, about 10%) in training time and error, showing that on a finer mesh batching reduces training time while the learning error increases.

Learning errors and inference speed are quantified in three examples, and two ways of combining with finite element software are demonstrated. The parametric Poisson problem reports relative energy, L2, and H1 errors for different network widths and batching settings; the Gaussian random field problem compares neural network and FEniCSx finite element times for computing quantities of interest; the nonlinear elasticity problem uses the network output as a Newton initial guess. Error statistics are based on means and standard deviations over K=10^4 samples (first two examples) or K=10^3 samples (elasticity), and timing comparisons are given on an A100 GPU and an Apple M1 CPU; in the elasticity case the network initial guess gives total times about 70% (basic bending) and 42% (extreme bending) of those with the zero function.

Using the network prediction as a Newton initial guess preserves final accuracy while speeding up the solve, with greater potential speed-up for larger deformations. Rather than treating the network prediction as the final solution, the learned finite element solution is used as an iteration initial guess, retaining finite element accuracy while exploiting millisecond-level network inference. In both the basic and extreme bending cases, the network initial guess and the zero-function initial guess converge to the same energy values (e.g., -0.014397 and 0.067984) but with shorter Newton solver times; the paper also estimates how many problems must be solved to offset training time (e.g., about 81,034 for basic bending and 20,921 for extreme bending with GPU training).

Perspective

The framework targets parametric PDEs that admit an energy functional, and its output is the standard finite element discrete solution rather than the exact solution, so it fits settings that want to combine with finite element theory or software, for example uncertainty quantification requiring many realizations or nonlinear problems needing a good iterative initial guess. For problems without an energy functional, the paper notes that a weak residual could be used instead, but the analysis and examples here center on energy. The examples use P1 triangular elements and MLPs, while the authors state that the framework allows other element types and higher-order polynomials.

The error estimate relies on a well-trainedness assumption (energy difference bounded by ε), which in practice is supported indirectly by the small errors in the examples but is not automatic for all problems. Batched element training can reduce training time on large meshes, but with the same number of iterations the learning error increases, so the trade-off between batch size and iteration count remains an open question. In the elasticity example the network can only provide an intermediate state rather than the final curled state, because during training the external forces are always applied to the unbent beam; the authors propose introducing time dependence and letting the network take the current beam position as input as a future direction. In addition, the training and inference time comparisons depend on the specific hardware (A100 GPU, Apple M1 CPU) and problem size, so they need to be re-evaluated when transferred to other settings.

Sources