MoME: Turning Sparse Lookup into Context-Aware Memory Mixtures
Synopsis
The work introduces Mixture of Memory Embeddings (MoME), which replaces each token's single memory row with multiple memory slots and uses a learned gate over the hidden state to sparsely choose which slots to read at each position, giving memory retrieval a context-aware form while keeping cheap token-indexed lookup; in controlled pretraining across nanochat, Llama-3/MobileLLM, and Qwen3 backbones, MoME improves over Value Embedding, Bigram, and STEM baselines in iso-parameter and iso-training-FLOP settings, shows a more favorable memory-size scaling trend at sub-billion scale, remains efficient in training and inference, and routing analyses on polysemous tokens suggest the same surface token is dispatched to distinct memory slots under different senses.
Interpretation
MoME changes the token-indexed memory table from one row per token to multiple memory slots per row, and uses a learned gate reading the hidden state to select which slots are active at each position, so memory addressing is no longer only a deterministic function of the surface form. Existing memory-embedding methods (Per-Layer Embedding, Value Embedding, STEM, Engram, and others) retrieve through deterministic token identity or fixed n-gram hashing, collapsing different contextual senses of the same token into one fixed vector; MoME keeps the token-indexed lookup structure while adding context-dependent slot selection. The paper formalizes the module (row indexer plus slot gate, sigmoid-norm aggregation, gated injection into the value stream) and runs controlled pretraining comparisons across three backbone families; an ablation shows hidden-state routing improves over learned token-only and fixed-random routing.
In controlled pretraining across three backbone families, MoME improves over Base, Value Embedding, STEM, and Bigram baselines in most matched comparisons under iso-parameter and iso-training-FLOP settings, while using fewer memory parameters than STEM. Prior comparisons of memory-embedding methods were often confined to a single backbone or scale; this work transfers the same mechanism to nanochat-style, Llama/MobileLLM-style, and Qwen3-style backbones at scales including 125M, 350M, and 0.6B. Training/validation bpb and CORE are reported, for example MoME-A2/6 on Llama/MobileLLM 125M reaches validation bpb 0.8560 and CORE 0.1686 versus VEmbedding's 0.8620 and 0.1530; on Qwen3 0.6B, MoME-A2/8 reaches CORE 0.2737 versus VEmbedding's 0.2530; training wall time is about 1.04–1.08 times the no-memory base.
In the low-compute sub-billion regime, MoME shows a more favorable validation-bpb trajectory than Bigram as memory capacity grows, and it can be combined with Bigram. Prior work reports less systematic memory-table scaling curves; this work fixes FLOPs on nanochat d12, sweeps memory parameters, and additionally tests using the Bigram indexer as MoME's first-stage indexer. The paper reports that MoME achieves lower validation bpb than the matched-memory Bigram row at every tested memory size, with substantially lower run-to-run variance in training and validation bpb curves; the combined configuration improves over Bigram under the same parameter budget.
Routing analyses indicate that MoME's learned routing is associated with word sense: same-sense prompts tend to route to the same memory slot, while a changed-sense prompt jumps to a different slot, with quantitative evidence on WiC. Prior memory embeddings retrieve deterministically, so there is no analyzable context routing; this work brings MoE-style sparse routing into the memory table and probes its semantic structure qualitatively and quantitatively. Across 144 memory layer–head sites in the d24 checkpoint, evaluated on 670 strictly filtered WiC pairs, 117 sites show greater routing divergence for different-sense pairs than same-sense pairs and 110 sites show greater slot overlap for same-sense pairs; the largest effect appears at layer 11/head 8, but that head has a nearly closed injection gate, and after gate weighting the largest effect moves to layer 7/head 6.
Perspective
The result speaks to researchers and engineering teams studying sparse capacity expansion under controlled pretraining: MoME applies to backbones that keep token-indexed lookup and inject memory into the attention value stream, and it has been validated on nanochat-style, Llama/MobileLLM-style, and Qwen3-style architectures at 125M to 0.6B scales and in a roughly 100B-token ClimbMix run; the mechanism can be combined with a Bigram indexer and stacked with sparse-FFN MoE. Its design goal is to let memory retrieval vary with context while keeping added latency modest, which makes it relevant to deployment settings concerned with inference latency and memory-parameter budgets.
The paper itself notes that the main limitation is computational scale: the extended 100B-token experiment reaches roughly billion-scale total parameters, the underlying dense backbone is about a billion parameters, and this scale check uses a single seed; the CORE activation and WiC analyses provide descriptive evidence of selective memory use and sense-sensitive routing, and causal interventions are still needed to establish whether these behaviors improve downstream accuracy or robustness under domain shift. In the WiC analysis only 81 target-token groups contain both sense labels, so the pooled comparison may reflect differences in lexical composition, and held-out generalization is not estimated; bootstrap intervals are pointwise and the largest effects are selected on the same examples without correction across the 144 sites. In addition, some table values in the loaded text appear as blanks, so specific per-task numbers should be checked against the original tables.
