FactorEngram factorizes n-gram memory over a shared dictionary, lifting 8K long-context retrieval by up to 43.3 points at 340M and 1B backbones
Synopsis
The work proposes FactorEngram, which replaces monolithic per-pattern n-gram embeddings with sparsity-regularized coefficients over a dictionary shared across patterns and reuses that same dictionary for basis-level contextual gating, improving language modeling, downstream tasks, and long-context retrieval on 340M- and 1B-parameter Transformer backbones.
Interpretation
FactorEngram represents the memory of local token patterns as sparsity-regularized coefficients over a shared dictionary instead of assigning each n-gram its own dense embedding, so related patterns can share parameters through common basis vectors rather than only through hash collisions. Where Engram treats each retrieved n-gram embedding as a monolithic unit stored in its own hashed slot and modulated by a single scalar gate, this work moves the smallest unit of memory from the n-gram embedding vector to dictionary coefficients, with a sparsity penalty following sparse coding. The paper trains jointly from scratch at two backbone scales, 30B tokens for 340M and 120B tokens for 1B (FineWeb-Edu, context 8192), comparing against a pure Transformer and the authors' reproduction of Engram; ablations show the full factorization design reduces WikiText perplexity from 22.32 to 21.29 and LAMBADA perplexity from 24.42 to 21.69 versus Engram with unigrams, while raising average downstream accuracy by 0.92 percentage points.
Basis-level contextual gating lets the context modulate each memory component individually: the backbone hidden state is projected into a query and scored against each basis vector, and the resulting weights scale the corresponding coefficients before reconstruction. Engram's scalar gate can only amplify or suppress the embedding as a whole and cannot retain context-relevant components while suppressing irrelevant ones for polysemous patterns such as "the bank"; here the vectors used to assess contextual relevance are exactly those used to reconstruct the output. In ablations, replacing basis-level gating with scalar gating while keeping the coefficient-dictionary representation degrades every metric, with average downstream accuracy dropping from 49.82 to 48.15 and the hardest NIAH-3 from 58.3 to 9.4.
FactorEngram covers both individual tokens and multi-token n-grams, and it systematically studies where the memory branch should be inserted, identifying insertion before the attention sublayer in the middle layers as an effective configuration. The paper notes that STEM retrieves embeddings only for individual tokens and Engram only for 2-grams and 3-grams, and that existing methods insert the memory branch at a fixed set of positions; here pattern coverage and insertion location are both treated as controlled variables. On the 24-layer 340M backbone, a single-module depth sweep favors layer 12 as a balance across evaluations; fixing layer 12 and sweeping the second module gives the 10 & 12 configuration the highest NIAH-3 among tested pairs; within-layer comparison shows before-attention insertion best on all accuracy metrics, before-FFN insertion reaching only 18.8 NIAH-3, and inside-FFN insertion worst across evaluations.
Across both backbone scales, FactorEngram improves overall evaluation performance over the Transformer and Engram, with particularly strong gains in long-context retrieval. Relative to Engram, NIAH-2 and NIAH-3 improve by 41.5 and 43.3 percentage points at 340M; at 1B the retrieval gains range from 9.3 to 27.6 percentage points, while WikiText perplexity is comparable to the Transformer and downstream accuracy is slightly lower than Engram. Results come from the paper's Table 1 and detailed Appendix Table 7, covering WikiText and LAMBADA perplexity, seven downstream accuracies (PIQA, HellaSwag, WinoGrande, ARC-Easy, ARC-Challenge, SocialIQA, BoolQ), and three 8K-context NIAH variants.
Perspective
This work targets research and engineering settings that scale LLM parameters through lookup-based memory: the memory is added as an auxiliary branch while attention and feed-forward modules remain intact, so it can be layered onto existing Transformer backbones. The paper's default configuration is a 24-layer 340M backbone with memory inserted at layers 10 and 12 before the attention sublayer, covering unigram, bigram, and trigram patterns with a moderate sparsity strength. The authors note that 1B-parameter models remain useful for edge computing and on-device deployment, so the result speaks directly to resource-constrained deployment settings; they also list whether these benefits persist at larger scales as future work.
The limitation the paper states is that FactorEngram is evaluated only at 340M and 1B parameters, leaving open whether the benefits persist at larger scales. Ablations show an optimum for sparsity strength: overall performance generally improves as the penalty increases from zero to a moderate value, but at a larger value all evaluation metrics become worse than the non-regularized model, so the sparsity strength remains a hyperparameter to reconfirm per scale and dataset. At 1B, WikiText perplexity is comparable to the Transformer and average downstream accuracy is slightly below Engram, indicating that gains are not uniform across metrics. In addition, basis vectors carry no predefined semantic labels, and the dictionary and coefficients are learned jointly with the backbone, so interpretability is not explored in this paper.
