Skip to main content
Back to timeline
arXivSource publication:

SpatialRNN replaces direct neuron-to-neuron links with a PDE medium, holding near-100% copy-task accuracy with 6.8k parameters while HORNNs degrade to about 13%

Synopsis

The authors introduce SpatialRNN, in which recurrent neuron-to-neuron communication is replaced by a spatially evolving medium governed by discretized PDEs, and prove that the medium state equals a convolution over the entire history of hidden states, i.e. a structured infinite-order RNN with a fixed parameter count; constructive conditions confine gradient modes to circles of prescribed radii, yielding better long-horizon performance than other recurrent models with fewer parameters.

Source-provided article image: Learning infinite context windows in recurrent architectures via spatial neural computing
Figure 1 ·

Figure 1: Schematic of the model architectures. ( a ) SpatialRNN: The neuron states h t h_{t} (represented by white nodes) are coupled through a shared spatiotemporal communication medium ψ t \psi_{t} (color coded), where information propagates locally over space and time. Each node writes to and reads from the medium, inducing wave-like interactions that encode past activity. ( b ) HORNN: Nodes correspond to neuron states h t h_{t} , and interactions are explicitly represented by edges connecting the current state to a finite set of past states h t − k h_{t-k} , for k = 1 , … , d k=1,\ldots,d . For d = 1 d=1 , the model reduces to a standard RNN. Block diagrams of the corresponding discretized computations are shown on the right of each panel.

arXiv

Interpretation

SpatialRNN replaces neuron-to-neuron communication with a spatial medium governed by discretized PDEs, and the medium state is proven equivalent to a convolution over the entire history of hidden states, making the model a structured infinite-order RNN with an unbounded receptive field and a fixed number of parameters. Finite-order HORNNs have a receptive field bounded by their order, while an infinite-order RNN is intractable in parameters and memory; this work implicitly encodes the infinite sequence of convolution kernels in four fixed-size matrices while keeping constant memory. Theorem 2.3 proves the equivalence between medium state and history convolution, with kernels defined recursively as the impulse response of the second-order linear difference equation.

The authors prove that finite-order HORNNs, including standard RNNs, cannot robustly maintain marginal stability: the input-dependent diagonal factor directly perturbs the coefficients governing modal decay, so they inevitably fall into vanishing or exploding gradient regimes. The result is independent of the specific parameterization and extends to orthogonal, nonnormal and diagonal RNNs previously used to mitigate gradient decay, showing such constraints can delay but not remove it. Propositions 3.1 and 3.2 are established via the characteristic equation and a proof by contradiction that relies only on the input-dependent factor multiplying the entire recurrence operator.

For SpatialRNN, the authors derive constructive conditions under which the companion matrix of the gradient dynamics has exactly simple eigenvalues of prescribed modulus in each block, invariant to any input-dependent perturbation; radius 1 gives memory retention and smaller radii give controlled forgetting. Unlike first-order recurrence where the nonlinearity directly scales each gradient step, the second-order recurrence separates the input-dependent perturbation from the coefficients that set gradient mode magnitudes, leaving those magnitudes invariant. Theorem 3.3 exploits the block upper-triangular structure of the weight matrices to factor the characteristic equation into independent quadratics, with condition 2 ensuring a negative discriminant and hence complex conjugate roots.

On the copy task, psMNIST and npCIFAR10, SpatialRNN achieves competitive results with fewer parameters: a 6.8k-parameter model holds near-100% validation accuracy across all sequence lengths on the copy task while a 64-neuron HORNN degrades to about 13%, and an 88-neuron model reaches 57.2% on npCIFAR10 with performance insensitive to the amount of appended noise. The copy-task comparison also serves as an ablation of state dimension versus PDE-mediated structure: a second-order HORNN already matches the order and state dimension of SpatialRNN yet remains subject to Proposition 3.2 and eventually loses accuracy. Experiments use no dropout or normalization and standard BPTT throughout; psMNIST and npCIFAR10 results are reported over three independent initializations.

Perspective

The result targets streaming sequence problems that need a fixed-size state, long retention and selective forgetting; the authors name sensor monitoring, long-horizon control and spatiotemporal signal prediction, where the locality and geometry of the medium can supply useful inductive biases. Methodologically, the block-triangular weight structure lets each block's gradient mode radius be set directly by a parameter, with radius 1 for retention and smaller values for forgetting; non-dissipative blocks preserve a symplectic form, in principle allowing exact backward reconstruction of the full trajectory and reducing BPTT memory to be independent of sequence length, while dissipative blocks require caching or checkpointing every few steps. The authors also expect the formulation to suit physical computing hardware where waves serve as the computational substrate.

The frozen-time guarantees are necessary but not sufficient for marginal stability under arbitrary switching, since products of matrices with unit-modulus spectra may still grow or decay; the authors report that exact Jacobians of a trained SpatialRNN change by only about 1% over 1000 steps along psMNIST trajectories, but the drift is not shown to be non-exponential. The authors state that necessary and sufficient conditions for LTV marginal stability remain an open problem. In addition, the information capacity of the medium is bounded by its dimension and numerical precision, and an unbounded receptive field does not imply lossless retention of arbitrary history; on npCIFAR10 performance is insensitive to noise length, indicating the bottleneck is classification expressivity rather than memory. In finite precision, the reliability of backward reconstruction depends on the dynamical regime, and dissipative blocks exponentially amplify round-off errors.

Sources