CEEN assigns one network per time subinterval and trains sequentially, cutting relative error to 1e-3–1e-2 on long-time PDE integrations where PINN fails
Synopsis
The work proposes causality-enforced evolutional networks (CEEN), which divide the time domain into nonoverlapping subintervals, assign a unique neural network to each subinterval, build the loss from the integral form of the PDE (trapezoidal rule), and train sequentially from the initial time step, achieving higher accuracy than the original PINN on long-time integration of the KdV, Schrödinger, convection, and Allen-Cahn equations while requiring less computational cost and memory, with a parallelization algorithm for further speedup.
(a)
arXivInterpretation
The paper attributes PINN failure in long-time integration to the lack of temporal causality in training: stochastic gradient descent tends to satisfy the governing equations at later times before learning the initial condition, converging to a low-loss but nonphysical local minimum (e.g., the convection example predicts a near-zero solution at later times). Relative to the original PINN's global spatio-temporal loss, the authors write causality explicitly into the model structure: the time domain is divided into nonoverlapping subintervals, each assigned an independent neural network, with the loss derived from the integral form of the governing equation over each subinterval. Using a one-dimensional convection equation with a fully connected network of 2 hidden layers and 820 neurons each, trained with Adam for 100,000 iterations, the authors show residual loss stalling at an earlier time while continuing to decrease at a later time, as direct evidence of violated causality.
The sequential training strategy reduces the residual loss at each subinterval below a given tolerance in time order, and the already-trained past parameters remain constant during subsequent training, ensuring the past state is unaffected by the current state. Unlike PINN, which minimizes the total loss over the entire time domain at once, CEEN splits the loss by time step and minimizes it in ascending time order, advancing to the next subinterval only when the current subinterval loss falls below tolerance. In the KdV and Schrödinger examples, CEEN reduces residual losses to tolerance at both earlier and later times, whereas PINN's residual loss stalls at an earlier time while continuing to decrease at a later time.
Across the four long-time integration examples (KdV, Schrödinger, convection, Allen-Cahn), CEEN achieves lower final-time relative error than PINN; for convection, PINN's relative error is 1.000e-0 versus CEEN's 5.465e-3. The comparison table shows CEEN consistently achieving higher accuracy across all four problems, whereas PINN fails on these long-time integration problems. Reference solutions are computed with the py-pde package at a time step size of 1e-7; final-time relative errors are KdV 5.314e-3 vs 1.237e-0, Schrödinger 4.793e-3 vs 2.346e-1, convection 5.465e-3 vs 1.000e-0, and Allen-Cahn 4.222e-2 vs 1.394e-0.
The parallelization algorithm trains multiple consecutive subinterval networks simultaneously, reducing computation time while keeping accuracy comparable across parallelization levels M; error analysis yields second-order accuracy in time step size and half-order accuracy in tolerance. Relative to serial per-subinterval training, the parallel version reduces computational cost by computing gradient terms concurrently; the error bound is decomposed into a training-loss term and a trapezoidal-rule consistency-error term, corresponding to tolerance and time step size respectively. With M = 1, 2, 5, 10, accuracy remains comparable across the four examples (e.g., convection around 9.2e-3); numerical experiments on the convection example confirm second-order accuracy in time step size for sufficiently small tolerance and half-order accuracy in tolerance for sufficiently small time step size.
Perspective
The method targets long-time integration of time-dependent PDEs with initial conditions, applicable when the initial condition or its derivative information is known and the solution over a subinterval can be represented by a small network. The authors note that the loss function can be replaced by other linear multistep methods (e.g., Adams-Moulton), and higher-order methods are anticipated to accommodate larger time steps while maintaining high-order accuracy, leaving room for future extension. The parallelization level M is constrained by GPU resources and must be chosen heuristically based on available memory. For high-frequency solutions, the authors suggest considering Fourier feature networks.
The error analysis currently provides a per-subinterval error bound and numerically verified accuracy orders; the authors explicitly state that a more comprehensive error analysis is left for future work. Unlike classical numerical methods, PINNs lack a convergence analysis analogous to the Lax Equivalence Theorem, and questions such as generalization error remain open. Training the initial condition relies on precise knowledge of the initial condition or its derivative information, and the applicability of this assumption in practical data scenarios deserves attention. The examples are concentrated on one-dimensional periodic-boundary problems (except Allen-Cahn), and the effectiveness of extension to higher dimensions and complex boundary conditions is unclear. The optimal choice of parallelization level M depends on GPU resources and lacks a systematic criterion.
