Skip to main content
Back to timeline
arXivSource publication:

FRAC swaps exponential forgetting for power-law long memory in SSMs, beating Mamba and GDN on long-context 1.3B language modeling

Synopsis

The work introduces FRAC, a selective state space model architecture derived from fractional dynamics that approximates a heavy-tailed fractional kernel with a finite-state, log-spaced sum of exponential modes, replacing exponential forgetting with power-law long memory and improving long-context performance over SSM baselines such as Mamba2, GDN, and Mamba3 on synthetic long-tail and recall tasks, 1.3B-parameter language modeling, and DNA modeling, while staying competitive on short-context tasks.

AI-generated editorial illustration: Fractional State Space Transition for Long Sequence Modeling

Interpretation

FRAC turns the long-memory mechanism of fractional differential equations into a finite-state recurrent layer that supports parallel training and autoregressive decoding. Most SSMs rely on ODE-based dynamics that induce exponential forgetting; fractional differential equations capture heavy-tailed hereditary effects but are non-Markovian, with the state depending on the entire past, so they are not directly compatible with finite-state recurrent layers. This work approximates the target power-law kernel with a finite, log-spaced sum of exponential modes, giving fractional memory a finite-state realization. The paper provides diffusive-representation and finite sum-of-exponentials approximation theorems, and Proposition 3 shows the mode bank uniformly approximates the fractional differential equation solution on the horizon of interest; the implementation uses zero-order-hold discretization with a chunked parallel scan and custom Triton kernels.

On controlled synthetic tasks, changing the memory law itself substantially improves length generalization. The heavy-tail probing task requires accumulating sparse events under a fixed power-law decay; models are trained only on length 512 and evaluated up to 128K, where FRAC shows the smallest decay while attention drops to near-random starting around 8K. Single-layer, roughly 200K-parameter, parameter-matched comparisons against Mamba2, GDN, Mamba3, and self-attention; on MADLab, four-layer models of about 500K parameters average 75.4 versus 74.8 for Mamba3 across compression, fuzzy in-context recall, memorization, and selective copying.

On 1.3B-parameter language models pretrained from scratch on 100B tokens, FRAC improves long-context performance over strong SSM baselines while remaining competitive on short context. NIAH evaluation pushes context far beyond the 4K training length, where FRAC degrades more slowly; on LongBench it averages 1.9% above the strongest linear baseline GDN and reports the best result on 8 of 14 tasks. All models are matched at roughly 1.3B parameters under the same training protocol with the Llama-2 tokenizer; on short-context LM Harness, FRAC trails the Transformer and Mamba3-MIMO by only 0.2% on average with comparable perplexity; on recall-retrieval it ranks third, 2.1% behind the Transformer and 0.4% behind Mamba3-MIMO.

The benefit of fractional memory transfers to genomic sequence modeling. Under the HyenaDNA setup, 7M-parameter causal language models are trained across sequence lengths, and FRAC achieves on-par or lower perplexity than Mamba3 and GDN at all lengths, with larger gains at longer contexts. Uses the HG38 human reference genome with a character-level DNA tokenizer, an 8-layer decoder-only core with matched hidden size, convolution size, and heads, and parameter matching to about 7M via model-specific expansion factors.

Perspective

The result targets sequence modeling settings where long-range dependencies matter, including long-context language modeling, genomic sequences, time series, video, and persistent memory for agentic systems; for tasks dominated by short-tailed dependencies requiring rapid forgetting, the advantage may be reduced. The method remains linear in sequence length with bounded state, but computational cost grows with the number of modes, so large-scale use requires care in kernel design and memory layout. A natural next step is combining fractional long-memory transitions with delta-rule style update mechanisms to obtain both a stronger long-range inductive bias and more precise associative retrieval.

The fractional kernel is approximated by a finite sum of exponentials on a fixed geometric timescale grid, so approximation quality depends on the number of modes and on whether the selected timescale range covers the dependencies a task requires; the paper's ablation shows that reducing modes to 8 consistently degrades performance, while increasing to 32 yields only comparable or small improvement at about 5% more parameters and 7% slower computation. Evaluation concentrates on tasks where long-range dependencies are central, leaving long-horizon settings such as video, scientific sequences, and persistent memory for agentic systems to be tested. Recall-retrieval tasks are truncated to 2K tokens following prior work, making them primarily short-context retrieval in this setting. In addition, the fractional theory specifies the per-token memory geometry and the zero-order-hold transition form, while the selective layer adds content-dependent read and write routing on top, an architectural generalization of the fixed-system fractional construction.

Sources