Gated Memory places an admission gate before extraction, gaining +2.6% relative accuracy on LoCoMo-10 with identical retrieval and generation
Synopsis
The work proposes Gated Memory, a formation layer interposed between conversation and storage: an utterance-level admission gate evaluates candidate facts before extraction, and admitted facts undergo entity scope classification and conditional enrichment, with only the current exchange treated as a fact source and prior turns used read-only for reference resolution. On LoCoMo-10, changing only what enters the memory store while holding retrieval, generation, and evaluation fixed yields a +2.6% overall relative gain in LLM-judge accuracy, with Single-Hop +4.4%, Open Domain +3.3%, Temporal +1.7%, and a 4.2% Multi-Hop regression.
Figure 1: The Gated Memory pipeline. Every candidate fact passes through an utterance-level admission gate. Admitted facts are scope-classified and conditionally enriched before being passed to the memory backend.
arXivInterpretation
The paper identifies the formation stage of long-term conversational memory as a structurally distinct and previously under-addressed failure point, characterizing three failure modes: reactive admission, absent enrichment, and irrecoverable loss of context. Existing systems such as MemGPT, mem0, and A-MEM invest in post-storage operations including retrieval, deduplication, lifecycle management, and episodic/semantic categorization, assuming that facts entering the store are worth storing; this work moves the intervention to the moment a fact is first written to storage. The argument is primarily structural, illustrated by a rented SUV and an owned SUV producing the same triple, and by entries an existing framework produced on LoCoMo such as "Has favorite memories," which pass deduplication, survive pruning, and occupy top retrieval slots on life-related queries.
It proposes an utterance-level admission gate that judges whether a candidate fact deserves storage against the complete original utterance before extraction, and decomposes the dialogue into a current exchange versus a read-only reference history so that formation is idempotent across turns. Lifecycle reasoning asks whether a fact conflicts with stored knowledge and operates on the extracted triple; admission control asks whether a fact deserves storage at all and operates while the utterance context is still intact, preserving signals such as ownership versus rental that no downstream component can recover. Admission criteria are three: encoding a personalization relation, a semi-stable user attribute, or cross-session reusability; transient states, generic emotional reactions, and vague non-factual content are rejected. The gate processed 5,882 utterances across ten conversations, admitting 5,527 (94.0%) and rejecting 355 (6.0%), with rejections mainly vague emotional expressions, procedural turns, and transient situational references.
It formalizes formation as structured extraction rather than binary filtering: admitted content is split into atomic facts, each assigned one or more of six formation categories, a provenance label distinguishing directly stated from inferred, and a conditional scope, with relative temporal and spatial references grounded at formation time. The episodic-semantic distinction is typically applied post hoc to already-stored content; this work determines category, provenance, and applicability scope while the utterance is intact, because the signals separating a directly stated attribute from one inferred in passing belong to the original utterance rather than the extracted triple. Examples include "gift ideas for my son" yielding both a persona fact and a procedural fact, and "find a hotel with an onsen, and remember I'm vegetarian" yielding a scoped preference and a durable persona attribute; provenance governs backend confidence, with directly stated attributes candidates for durable semantic storage and inferred attributes retained provisionally pending corroboration.
It introduces a four-class entity scope taxonomy (USER_PII, INTERNAL, EXTERNAL, GENERIC) governing conditional enrichment, and enforces a grounding constraint forbidding assertion of any entity or relation absent from the evaluated window, while preserving the language of the user utterance. Uniform enrichment risks sending personal information to external services without consent while zero enrichment forfeits semantic depth; scope classification decides at formation time whether an external knowledge source is permitted, and the grounding constraint converts confident assertion of unsupported entities into a bounded, auditable decision. USER_PII receives no enrichment, INTERNAL uses internal metadata only with no external calls, EXTERNAL is grounded externally when semantically necessary, and GENERIC uses parametric knowledge or none; a window containing only "he is seven" with no antecedent naming a child yields the user's child is seven rather than a fabricated son.
Perspective
The framework targets the memory write stage of personalized conversational assistants and is meant to be used as a formation layer inserted into existing memory backends, with retrieval and generation unchanged, so it suits deployments that want better memory quality without modifying downstream components. It explicitly addresses the moment a fact is first written to storage, with admission criteria built around personalization relations, semi-stable user attributes, and cross-session reusability, rejecting transient states, generic emotional reactions, and vague non-factual content. Scope classification and conditional enrichment target settings that need privacy constraints: USER_PII is not enriched, INTERNAL uses internal metadata only, and EXTERNAL is grounded externally only when semantically necessary. The authors note that in production environments with a high volume of irrelevant context the gate could be more beneficial than on LoCoMo-10, and list domain-adaptive thresholds and knowledge-graph grounding as natural extensions.
The admission gate's effectiveness depends on the quality of the underlying LLM prompt, and adversarial or ambiguous utterances may be misclassified at formation time; unlike downstream deduplication, formation errors are undetectable once the fact is formed. Evaluation is conducted on LoCoMo-10, whose utterance data carries atypical emotional density, so generalization across domains, languages, and task-oriented deployments still needs empirical demonstration; the 6.0% rejection rate is treated as a conservative lower bound, with task-oriented deployments expected to show substantially higher filtering rates. The 4.2% Multi-Hop regression is attributed to a precision-recall tradeoff inherent to independent fact evaluation, and chain-aware admission modeling inter-fact dependencies is left as future work. The enrichment stage introduces external knowledge calls for EXTERNAL-scoped facts, adding latency and potential staleness risk for time-sensitive information. In addition, the individual contributions of admission and enrichment are not disentangled and are left as future work.
