Skip to main content
Back to timeline
arXivSource publication:

Kernel and neural-network solvers for HJB equations on unbounded domains: convergence and error bounds from empirical residuals

Synopsis

For two non-monotone, mesh-free approximations of Hamilton–Jacobi–Bellman equations on unbounded spatial domains—Wendland kernel collocation under a native space norm constraint and physics-informed neural networks with smooth activations under a Sobolev-norm constraint—the work derives error bounds from computable empirical squared residuals: norm constraints, an interpolation inequality and a controlled-diffusion truncation estimate convert the empirical loss into the pointwise residual control needed for viscosity-solution error estimates, yielding uniform convergence on compact sets plus a posteriori and high-probability quantitative bounds governed by the fill distance for kernels and by network and sample sizes for neural networks.

Source-provided article image: Convergence of kernel and neural-network methods for Hamilton--Jacobi--Bellman equations on unbounded domains
Figure 5 ·

Figure 5.1: Kernel realization, d = 1 d=1 , R = 3 R=3 , c = 1 c=1 , λ = 6 \lambda=6 . (left) Errors on the fixed compact core Q core = [ 0 , 1 ] × [ − 1 , 1 ] Q^{\mathrm{core}}=[0,1]\times[-1,1] and on the full truncation cylinder [ 0 , 1 ] × [ − 3 , 3 ] [0,1]\times[-3,3] , versus the spatial spacing h x h_{x} , with a reference slope h x 3 h_{x}^{3} . (right) The a posteriori residual bound γ n \gamma_{n} of ( 3.6 ), versus h x h_{x} , against an h x 3 / 2 h_{x}^{3/2} reference slope.

arXiv

Interpretation

A unified empirical-residual-to-error mechanism: norm constraints give locally uniform regularity of candidates, a Gagliardo–Nirenberg-type interpolation turns population residuals into pointwise residuals, and a controlled-diffusion truncation estimate converts a residual bound on a truncation cylinder into convergence on compact subsets. Viscosity-solution error estimates previously required pointwise residual control, whereas kernel collocation and PINN/DGM only minimize an empirical squared PDE residual and terminal mismatch; the paper recovers that pointwise control from computable empirical losses under norm constraints, quantifying discretization error for kernels and sampling error for networks. The mechanism is stated and proved as lemmas and theorems in Section 2 and is shared by both realizations; its inputs are only the unique bounded continuous viscosity solution supplied by (A1) and (A2), with no uniform nondegeneracy required.

Kernel realization: with an ansatz of translates of a Wendland function under a native space norm constraint, Theorem 3.3 gives uniform convergence on compact sets and Theorem 3.4 gives an a posteriori quantitative error bound for the unique bounded continuous viscosity solution, governed by the empirical loss and fill distance. Existing rigorous non-monotone theories were largely confined to bounded domains, Cordès or ABP operator classes and mesh-based discretizations; this work extends residual-to-error certificates to kernel collocation on unbounded domains with a posteriori bounds evaluable from computed empirical losses. Convergence relies on additional classical and native space regularity assumptions; the a posteriori bound (3.10) applies to every candidate satisfying the native space constraint, with a right-hand side involving only computed empirical losses and prescribed parameters.

Neural-network realization: with a feedforward network using a bounded smooth activation such as tanh under a Sobolev-norm constraint, Theorem 4.9 gives uniform convergence on compact sets in probability and Theorem 4.10 gives a quantitative high-probability error bound, with rate governed by the approximation power of the network class and by a Rademacher bound from a covering argument in the parameters. Prior PINN/DGM convergence and error estimates mostly concerned linear, semilinear or quasilinear equations; this work treats direct minimization of an empirical residual for the finite-horizon second-order HJB equation and decomposes the error into approximation, generalization, concentration and optimization terms. Bound (4.13) holds with the stated probability for every admissible parameter choice; explicit rates follow once an approximation rate is available, the tanh case citing De Ryck et al., and that rate exhibits the curse of dimensionality.

Numerical experiments: on one- and two-dimensional manufactured problems the kernel realization reduces core RMS error to about 1e-3 under spatial refinement, with bandwidth optimal near 1.25, and under theorem-scaled refinement of grid and truncation radius the core RMS error falls from about 7.9e-2 to about 5.7e-5; network error decreases with width. Experiments report both core and full-cylinder errors, showing that for large bandwidth errors concentrate near the artificial boundary while a larger bandwidth largely removes this layer; against a monotone explicit finite-difference baseline the kernel method is more accurate on the core while finite differences are far cheaper. Experiments use one manufactured equation, fixed time grids and short training budgets; network training does not enforce the theoretical Sobolev constraint and uses a product of layer norms only as a computable proxy, and the authors state the experiments illustrate rather than validate the theory.

Perspective

The analysis targets finite-horizon second-order HJB equations with a controlled-diffusion representation on unbounded spatial domains, and applies to kernel collocation and PINN-type methods that minimize an empirical residual under native space or Sobolev norm constraints. For the kernel method the a posteriori bound can be used to assess any candidate satisfying the constraint; for the network method the error bound holds with stated probability for admissible parameters and yields explicit rates once an approximation rate is available. The experiments address one- and two-dimensional manufactured problems to display the effects of grid refinement, bandwidth choice, truncation radius and network width, and allow comparison with a monotone finite-difference baseline on core and full-cylinder errors.

The convergence results rely on additional classical-solution and native space or Sobolev regularity assumptions that are separate from viscosity wellposedness; the network realization also needs generalization and tail conditions, and its explicit rate exhibits the curse of dimensionality in the tanh case. The numerical experiments use one manufactured equation, fixed time grids and short training budgets, and network training does not enforce the theoretical Sobolev constraint but uses a product of layer norms as a proxy, so the experiments illustrate rather than validate the theory; higher-dimensional experiments and broader baselines are left for future work.

Sources