Public articles linked to the same research event.
arXiv The work pretrains looped language models to share memory—only the first recursion writes a key-value cache, and later recursions read it while keeping a short window of their own—and at 150M-1B parameters its Looped Prediction Transformer (LPT) and hybrid variant lower validation perplexity on FineWeb-Edu by 1.12-1.82 relative to a same-size standard Transformer with five recursions while using 76-79% less context memory, with analysis showing shared and local memory develop different representations, later recursions attend mostly to the shared memory, and the shared memory acts as a gradient highway to the first recursion.
The work pretrains looped language models to share memory—only the first recursion writes a key-value cache, and later recursions read it while keeping a short window of their own—and at 150M-1B parameters its Looped Prediction Transformer (LPT) and hybrid variant lower validation perplexity on FineWeb-Edu by 1.12-1.82 relative to a same-size standard Transformer with five recursions while using 76-79% less context memory, with analysis showing shared and local memory develop different representations, later recursions attend mostly to the shared memory, and the shared memory acts as a gradient highway to the first recursion.
The work pretrains looped language models to share memory—only the first recursion writes a key-value cache, and later recursions read it while keeping a short window of their own—and at 150M-1B parameters its Looped Prediction Transformer (LPT) and hybrid variant lower validation perplexity on FineWeb-Edu by 1.12-1.82 relative to a same-size standard Transformer with five recursions while using 76-79% less context memory, with analysis showing shared and local memory develop different representations, later recursions attend mostly to the shared memory, and the shared memory acts as a gradient highway to the first recursion.
The work pretrains looped language models to share memory—only the first recursion writes a key-value cache, and later recursions read it while keeping a short window of their own—and at 150M-1B parameters its Looped Prediction Transformer (LPT) and hybrid variant lower validation perplexity on FineWeb-Edu by 1.12-1.82 relative to a same-size standard Transformer with five recursions while using 76-79% less context memory, with analysis showing shared and local memory develop different representations, later recursions attend mostly to the shared memory, and the shared memory acts as a gradient highway to the first recursion.