MemTrim indexes memory at the evidence level, cutting memory-induced errors under partial overlap from 11.6% to 0.6%
Related research and updatesSynopsis
The authors identify and study memory overreliance, an inference-time failure in which benign, correctly stored, and appropriately retrieved memory misleads reasoning when it only partially overlaps the current query, and propose MemTrim, a plug-and-play framework that decomposes memories into evidence units indexed in a trie at write time and, at read time, removes evidence repeated in or conflicting with the query, retains memory-only information, and suppresses a previous answer when its supporting evidence has changed, reducing memory-induced errors across benchmarks, models, and memory architectures while preserving the benefits of useful memory.
Figure 1: Illustration of memory overreliance. A retrieved memory can appear highly relevant because it shares substantial evidence with the current query, while still carrying context-specific information that does not transfer to the present task. Reusing the memory without separating these pieces of evidence can misguide the model toward an unsupported conclusion.
arXivInterpretation
The paper identifies memory overreliance as an inference-time failure mode in which memory that is benign, correctly stored, and appropriately retrieved still harms inference on the current task. Prior discussion of memory failure attributes it to flawed memory, flawed retrieval, or an attacker, whereas this work shows that a correct, well-retrieved memory can fail simply because its answer-relevant condition has changed. Across RippleEdits, CounterLogic, and CLADDER-Derived, four memory architectures (Mem0, A-Mem, MemGPT, Graphiti), and two models (Qwen3.8-27B, GPT-5.6 Luna), memory with and without retrieval is compared; on task-divergent with overlapping evidence (TD-Dup) cases, memory-induced correct-to-wrong changes far outnumber wrong-to-correct ones, e.g., 19.4 versus 2.1 for Qwen3.8-27B with Graphiti on CLADDER-Derived.
Failures concentrate under partial query-memory overlap rather than at the extremes of little or very high overlap. The paper measures query-memory overlap as content-word Jaccard overlap and compares its distribution between successful and failed cases, finding that the two groups occupy different parts of the overlap range with a statistically significant separation across all four memory systems. Based on the distribution of naturally occurring TD-Dup cases, plus controlled experiments that hold the remaining memory content fixed and vary only the fraction of repeated evidence, tested with a quadratic term and the two-lines test for a non-monotonic trend.
Controlled experiments show a non-monotonic effect of overlap: with little overlap memory has limited influence, its harmful effect grows with overlap and is strongest at intermediate levels, and the effect reverses once the interaction becomes sufficiently aligned. This turns the correlational observation of RQ2 into direct manipulation of the amount of overlapping evidence, showing that overlap increases memory influence without guaranteeing that the evidence supporting its conclusion transfers to the current task. A consistent non-monotonic trend is observed across models and memory systems, assessed with a quadratic term and the two-lines test.
MemTrim controls memory reuse at the evidence level, reducing memory overreliance while preserving useful memory and without weakening performance on task-aligned cases. Unlike baselines that remove entries, discard stored answers, or compress context, MemTrim removes only repeated and conflicting evidence, retains memory-only information, suppresses a previous answer when its supporting evidence changes, and relabels it as task-tagged evidence when the task changes. Compared under the same retrieved memories against Facts-only, Top-1, MMR, Similarity Filter, RECOMP, LLMLingua-2, Context-faithful, Answer-first, and Micro-Act; with Mem0 as backbone, correct-to-wrong on RippleEdits drops from 11.6 to 0.6 for Qwen3.8-27B and from 9.7 to 0.5 for GPT-5.6 Luna, with TA performance close to or better than the original memory system; ablations show each component contributes in its intended setting; degradation is slower under PoisonedRAG, MINJA, and ShadowMerge attacks; runtime changes are small.
Perspective
The work targets LLM agents that use agentic memory, especially long-horizon settings where memory entries are repeatedly retrieved and reused. MemTrim is designed as a plug-and-play layer on top of existing memory systems: the underlying system still stores and retrieves by its original mechanism, while MemTrim builds an evidence index at write time and processes retrieved memories at read time, so it applies to embedding-based and structured memory systems without retraining the underlying model. Its benefits are evaluated separately for task-aligned (TA) and task-divergent with overlapping evidence (TD-Dup) cases, and stress-tested under text-based and graph-based memory attacks. The method relies on parsing interactions into key-value evidence units, task signatures, and answer support sets, so its applicability is tied to parsing quality.
Several open questions remain: how stable evidence parsing and key canonicalization are in more open domains, and how the conservative policy on parsing failures (keeping the raw memory but not reusing its stored outcome through MemTrim) affects overall gains; task-signature and answer-support decisions are conservative under uncertainty, which may suppress answers that could have been reused; when several retrieved memories give different values for the same key and the current query does not resolve the conflict, all values are withheld by default rather than chosen by retrieval similarity, and behavior with trusted ordering signals such as timestamps deserves further observation; the controlled overlap experiments hold the remaining memory content fixed, while in natural settings factors beyond overlap also distinguish successful from failed cases; attack stress tests increase injected entries under a fixed retrieval budget, whereas real attacks may inject in more complex ways. In addition, some table values in the provided text are not fully rendered, so references to specific accuracy numbers are limited to figures explicitly stated in the text.
