ResonatorLM: Replacing Attention with Causal Resonant Field Mixing for Efficient Long-Context Language Modeling
Synopsis
This work introduces ResonatorLM, which replaces self-attention with a causal resonant field mixer built from damped resonators, treating a token sequence as a driven latent field, running training and prefill through O(T log T) causal FFT convolution and decoding through a fixed-size recurrent state; in a matched setting near 6M parameters, WikiText-2 character-level test perplexity drops from 4.617 to 3.764 and accuracy rises from 55.32% to 61.31%, decode reaches a 6.47x speedup over an optimized attention block at 32K tokens, and a separate kernel-level benchmark reports 440.29x at 8K and 575.86x at 32K.
Fig. 1. ResonatorLM block overview. A grouped input projection feeds a resonant mixing path that uses FFT convolution during training and prefill and fixed-size recurrent state during autoregressive decoding. A local causal convolution and head coupling complement the global resonant field before the output is passed to the MLP.
· Page 4Interpretation
It proposes a causal mixer that replaces attention dot products with damped-resonator kernels, where the same kernel family runs as causal FFT convolution for training and prefill and as a fixed-size recurrent state for autoregressive decoding. Relative to linear and kernelized attention and to S4, Mamba, and Hyena, which remain derived from attention or state-space formulations, this work grounds sequence modeling directly in a driven latent field with damped oscillatory dynamics and shares one kernel parameterization across both execution modes. The paper gives an explicit kernel parameterization k_h[t]=exp(−α_h t)cos(ω_h t+ϕ_h), the FFT convolution form, and the recurrent update, and states that the decode state has size B×H×d_h×2 and does not grow with decoded length; this is a design-level formalization.
In a matched WikiText-2 character-level setting near 6M parameters, it improves quality over a matched transformer while training throughput is lower. Against the matched baseline, test perplexity falls from 4.617 to 3.764 (about an 18.5% relative reduction) and accuracy rises from 55.32% to 61.31% (+5.99 percentage points), while training throughput is 172,890 tok/s versus 262,610 tok/s. The main comparison uses six seeds per model and reports 95% confidence intervals (perplexity [4.561, 4.673] versus [3.757, 3.772]; accuracy [54.968, 55.674] versus [61.239, 61.390]), which do not overlap.
Long-context efficiency grows with sequence length, and decode crosses over between roughly 4K and 8K tokens to become clearly faster. In the practical block benchmark, train and prefill speedups rise monotonically with length, while decode is 0.40x at 2048 tokens (0.6254 versus 0.2532 ms/tok), turns to 1.45x at 8192, 3.40x at 16384, and 6.47x at 32768 (0.6228 versus 4.0296 ms/tok). This benchmark includes projections, state or cache construction, and decode-step cost, measured on a single NVIDIA L4 with batch size 1, decode length 128, and five timed iterations after two warm-ups; the paper explicitly separates it from the kernel-level benchmark that isolates asymptotic mixer behavior.
The quality advantage holds across datasets, tokenizations, and longer contexts, and does not depend on a single fragile hyperparameter setting. ResonatorLM leads on WikiText-2 char/byte, WikiText-103 char, TinyStories char/byte, WikiText-2 at lengths 512 and 1024, and TinyStories long-context; in ablations, removing head coupling or the local path changes perplexity by only 0.021 and accuracy by only 0.171 points. Breadth and long-context studies use three seeds per setting and ablations use six seeds; the paper also reports physics diagnostics with a maximum prefix error of 7.75×10⁻⁷ and a learned half-life range of 2.0 to 2048.0 tokens.
Perspective
The result concerns the efficiency-quality tradeoff in long-context language modeling and applies to the character- and byte-level tasks tested here (WikiText-2, WikiText-103, TinyStories), under settings near 6M parameters with quality evaluation up to 1024 tokens and efficiency evaluation up to 32K tokens; it enables follow-up work to test the same kernel family at larger scales, on more complex long-context tasks, and in broader sequence-modeling domains such as multimodal settings, and it offers readers concerned with decode memory growing linearly with context a reproducible alternative operator design.
A careful reader would still watch whether the quality and efficiency trends persist well beyond 6M parameters and on more complex long-context tasks; how comparisons against newer architectures and stronger system-level optimizations would look; how the gap between the practical block benchmark and real production environments would affect observed speedups; and how the stated use of AI assistants to help draft, proofread, and formalize some key figures affects the presentation of details. In addition, this reading covered the full text, but where specific values in figures or tables are not fully restated in the prose, individual details may still need to be checked against the original.
