EngramEdit Decouples Knowledge Updates via Conditional Memory, Reaching 93.9 Utility on CounterFact While Retaining Over 96% of General Ability
Synopsis
EngramEdit proposes decoupled knowledge updates through conditional memory: it first computes target memory representations that make the model predict an updated fact across multiple expressions, then jointly solves updates to shared n-gram embeddings while penalizing changes to frequently reused embeddings more strongly; on LongCat-Flash-Lite, updating only conditional memory yields near-perfect editing success on CounterFact and ZsRE, revised facts transfer to unseen expressions and multi-hop reasoning, and unrelated knowledge and general capabilities are largely preserved.
Figure 1 : Conditional memory and knowledge updating with EngramEdit . (a) Input n n -grams look up learned embeddings for backbone computation. (b) Disabling memory in DeepSeek Engram causes larger relative performance drops on factual-knowledge benchmarks [ 3 ] . (c) EngramEdit computes target memory representations and jointly updates the corresponding n n -gram embeddings, with stronger penalties on frequently reused embeddings.
arXivInterpretation
Updating only n-gram embeddings in conditional memory while keeping the Transformer backbone and MoE experts fixed suffices for factual editing, reaching 93.9 Utility on CounterFact and 76.4 on ZsRE, above FT, FT-L, AdaLoRA, UnKE, MoEEdit, and the direct memory fine-tuning baselines MFT-S and MFT-A. Conditional memory had mainly served model scaling, personalization, and domain adaptation; this work treats it as an editable knowledge interface that separates factual storage from general-purpose computation without touching backbone parameters. 2,000 sequential batch edits (batches of 100) on LongCat-Flash-Lite, a 68.5B-parameter MoE with 31.4B parameters in n-gram embeddings, reporting Efficacy, Generalization, Specificity, and their mean Utility with 95% confidence intervals.
Revised knowledge transfers to paraphrases unseen during editing and is used in multi-hop reasoning: on MQuAKE with chain-of-thought prompting, accuracy is nearly 3 times the strongest baseline and leads in every hop group. Prior editors can succeed on the edit prompt without recalling the fact under other wordings; generating multiple expressions per fact and jointly matching their targets expands the set of updated n-grams. 3,000 MQuAKE counterfactual cases (1,135 two-hop, 1,136 three-hop, 729 four-hop) under standard and CoT prompting with Wilson 95% confidence intervals; ablating joint allocation drops CounterFact Utility from 93.9 to 77.4.
Unrelated knowledge and general capabilities are largely preserved during sustained editing: over 96% of pre-edit mean F1 remains after 5,000 sequential CounterFact edits, and on ZsRE 93.83% of correctness states and about 87.4% of initially correct predictions are retained. The method uses n-gram length and corpus frequency as reuse indicators, penalizing updates to shorter or more frequent embeddings more strongly, and proves that this weighting bounds the average squared change in memory representations on unrelated inputs. Six general-ability tasks with 100 examples each, evaluated before editing and every 500 edits; correctness transition analysis over 11,379 neighborhood prompts; ablating reuse-based regularization lowers Efficacy and Specificity.
The model relies on these memory updates to recall revised facts: disabling fact-related updates turns 5,228 of 5,850 successful prompts into failures, nearly matching disabling all 9,253 updates, whereas matched random disabling causes no successful prompt to fail. This provides causal evidence of use rather than inferring it from editing success alone, showing revised knowledge enters the forward computation through a small set of activated n-gram embeddings. Four disabling conditions on the same model after 2,000 CounterFact edits, with edit margins computed from length-normalized log-likelihood; 92.24% of unrelated-query regressions accompany activation of updated embeddings.
Perspective
The results target LLM deployments that use n-gram conditional memory, such as services needing to revise outdated facts while keeping general capabilities and unrelated facts stable; at inference an embedding update enters the computation only when its n-gram is active in the current prefix, so the method suits edit requests whose subject position can be localized. The authors note future work on other conditional memory architectures and model scales, and on efficient, stable updates to more facts at higher frequencies.
Readers should note that experiments center on a single model, LongCat-Flash-Lite, at 2,000 to 5,000 sequential edits, so behavior at larger scales and higher edit frequencies remains to be tested; among general-ability tasks, MRPC shows a relatively larger F1 decline, which the authors attribute to comparing two sentences that may activate different n-grams; cross-edit sharing involves only 1.25% of updated n-grams yet concentrates 97.96% of Efficacy failures, leaving conflicting targets on shared embeddings an open question; and evaluation uses existing benchmarks whose counterfactual edit targets are experimental test cases, not verified real-world facts.
