Shared Memory in Looped Transformers: Five Recursions Cut Validation Perplexity by 1.12-1.82 Versus a Same-Size Standard Transformer While Using 76-79% Less Context Memory
Related research and updatesSynopsis
The work pretrains looped language models to share memory—only the first recursion writes a key-value cache, and later recursions read it while keeping a short window of their own—and at 150M-1B parameters its Looped Prediction Transformer (LPT) and hybrid variant lower validation perplexity on FineWeb-Edu by 1.12-1.82 relative to a same-size standard Transformer with five recursions while using 76-79% less context memory, with analysis showing shared and local memory develop different representations, later recursions attend mostly to the shared memory, and the shared memory acts as a gradient highway to the first recursion.
Figure 1: Shared memory not only saves memory, it also improves quality. Left: in our Looped Prediction Transformer ( lpt ), later recursions read the KV cache written by the first recursion and keep only a short window of their own. Right: perplexity and context memory of 1 1 B-parameter models with two to five recursions, against a parameter-matched standard Transformer ( std ). Each of our models adds shared memory to the looped baseline above it: lpt to the looped Transformer lt , and h-lpt to the gated DeltaNet hybrid lt2 .
arXivInterpretation
Introduces a pretraining scheme in which looped Transformers share memory: only the first recursion writes a key-value cache, and later recursions read it while keeping a short window of their own. Looped Transformers write their own key-value cache at each recursion, so memory still grows with compute, and inference-time techniques can shrink this cache at a cost in quality. This work shares memory during pretraining rather than only compressing at inference time. The abstract describes the pretraining scheme and contrasts it with inference-time techniques, reporting validation at 150M-1B parameters.
Sharing memory does not cost quality and instead improves it: with five recursions the hybrid variant lowers validation perplexity on FineWeb-Edu by 1.12-1.82 relative to a same-size standard Transformer while using 76-79% less context memory. Sets a new quality-memory frontier for looped models, showing shared memory can improve both quality and memory efficiency rather than trading quality for memory savings. The abstract reports specific numbers at 150M-1B parameters, five recursions, a 1.12-1.82 validation perplexity reduction on FineWeb-Edu, and 76-79% less context memory.
Explains why memory sharing helps through analysis: shared and local memory develop different representations, later recursions attend mostly to the shared memory, and the shared memory acts as a gradient highway to the first recursion. Provides a mechanism-level account that attributes the quality gain to representation differentiation, attention allocation, and gradient pathways, not only to empirical observation. The abstract states an extensive analysis yielded these mechanistic findings, though the abstract does not list specific experimental details.
Perspective
The result targets looped Transformer language models: at 150M-1B parameters, five recursions, and FineWeb-Edu validation, the shared-memory scheme improves both perplexity and context memory footprint. It applies to pretraining settings that want more recursive compute under a fixed memory budget and to deployment settings that need long-context inference-time compute scaling. The division of labor between shared and local memory, the concentration of later-recursion attention on shared memory, and the shared memory acting as a gradient highway offer testable design directions for follow-up work.
The abstract does not list full experimental settings, ablations, baseline details, or statistical significance, so the robustness of the quality-memory frontier, the behavior of shared memory across different recursion counts and data distributions, and whether the gradient-highway mechanism generalizes to other architectures remain open questions for a careful reader. The abstract also does not state the exact size of the short window or implementation details of the shared cache, which affect reproduction and transfer.
