ChronoSpike combines learnable LIF neurons with a lightweight Transformer temporal encoder, gaining 2.0% Macro-F1 and 2.4% Micro-F1 on average across three large dynamic-graph benchmarks while training 3-10 times faster than recurrent methods
Related research and updatesSynopsis
The work proposes ChronoSpike, an adaptive spiking graph neural network that integrates learnable LIF neurons with per-channel membrane dynamics, multi-head spatially-attentive aggregation over continuous features, and a lightweight Transformer temporal encoder; it outperforms twelve state-of-the-art baselines on three large benchmarks by 2.0% Macro-F1 and 2.4% Micro-F1 on average, trains 3-10 times faster than recurrent methods, holds a constant 105K-parameter budget independent of graph size, provides theoretical guarantees for membrane potential boundedness, gradient flow stability under contraction factor ρ<1, and BIBO stability, and its interpretability analyses reveal heterogeneous temporal receptive fields and a learned primacy effect with 83-88% sparsity.
Figure 1: Overview of ChronoSpike. The dynamic graph is represented as a sequence of snapshots. At each time step, node features are aggregated from sampled neighborhoods using a multi-head attentive spatial aggregator and encoded into spike signals via adaptive LIF neurons. Temporal dependencies across snapshots are captured by a lightweight Transformer-based temporal aggregation module with learnable positional encodings. The final node representations are used for prediction.
arXivInterpretation
ChronoSpike integrates learnable LIF neurons with per-channel membrane dynamics, multi-head spatially-attentive aggregation, and a lightweight Transformer temporal encoder into an adaptive spiking graph neural network for jointly capturing structural relations and temporal evolution in dynamic graphs. Existing approaches face a core trade-off: attention-based methods offer expressiveness at O(T^2) complexity, while recurrent architectures suffer from gradient pathologies and dense state storage; spiking neural networks provide event-driven efficiency but are constrained by sequential propagation, binary information loss, and local aggregation lacking global context. The design targets this trade-off with a combined scheme. The abstract states the architectural components and complexity profile: O(T·d) activation/state memory plus an additional O(T^2) per-node attention term that remains small for the horizons evaluated here.
On three large benchmarks, ChronoSpike outperforms twelve state-of-the-art baselines by 2.0% Macro-F1 and 2.4% Micro-F1 on average, while training 3-10 times faster than recurrent methods with a constant 105K-parameter budget independent of graph size. The result reports efficiency and accuracy advantages within the same comparison rather than emphasizing event-driven efficiency alone. Evidence comes from the three large benchmarks and twelve baselines described in the abstract, reported as averages; per-dataset values, variance, and significance tests are not given in the abstract.
The authors provide theoretical guarantees for membrane potential boundedness, gradient flow stability under contraction factor ρ<1, and BIBO stability. This offers a formal characterization of training stability for spiking dynamics on dynamic graphs, beyond empirical observation. The abstract explicitly lists three theoretical guarantees but does not give theorem numbers, proof details, or assumptions.
Interpretability analyses reveal heterogeneous temporal receptive fields and a learned primacy effect, with 83-88% sparsity. The analysis links model behavior to temporal preference, offering an observation window into the interpretability of spiking dynamics. Evidence is the interpretability analysis described in the abstract; the specific analysis protocol and statistical treatment are not expanded in the abstract.
Perspective
The work targets dynamic graph representation learning that must model structural relations and temporal evolution together, suited to settings where graph size varies but the parameter budget must stay constant: the abstract states a constant 105K-parameter budget independent of graph size, O(T·d) activation/state memory, and an additional O(T^2) per-node attention term that remains small for the horizons evaluated here. The efficiency claim is relative to recurrent methods, with training reported as 3-10 times faster. The theoretical guarantees cover membrane potential boundedness, gradient flow stability under contraction factor ρ<1, and BIBO stability. Interpretability analyses report heterogeneous temporal receptive fields and a primacy effect with 83-88% sparsity. Code is publicly available.
Questions still open at the abstract level include: the names, scales, and time horizons of the three large benchmarks; the composition of the twelve baselines and their tuning protocols; the variance and significance of the 2.0% Macro-F1 and 2.4% Micro-F1 average gains; the hardware and batch settings behind the 3-10 times training speedup; whether the O(T^2) attention term remains small over longer horizons; the assumptions and proof details underlying the three theoretical guarantees; and how the 83-88% sparsity and primacy effect are measured. These are scope and open questions rather than grounds for dismissal.
