Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Shared Memory in Looped Transformers: Five Recursions Cut Validation Perplexity by 1.12-1.82 Versus a Same-Size Standard Transformer While Using 76-79% Less Context Memory

The work pretrains looped language models to share memory—only the first recursion writes a key-value cache, and later recursions read it while keeping a short window of their own—and at 150M-1B parameters its Looped Prediction Transformer (LPT) and hybrid variant lower validation perplexity on FineWeb-Edu by 1.12-1.82 relative to a same-size standard Transformer with five recursions while using 76-79% less context memory, with analysis showing shared and local memory develop different representations, later recursions attend mostly to the shared memory, and the shared memory acts as a gradient highway to the first recursion.