This Perspective proposes three principles—substrate proportionality, temporal defeasibility, and auditable continuity—so memory-enabled LLM agents can decide which substrate a new piece of evidence belongs in, or whether to write nothing at all.
Synopsis
This Perspective addresses LLM agents that accumulate state across sessions, tools, users, and changing environments, and proposes a framework of controlled knowledge updating: treating an update as a decision to select the least invasive substrate sufficient for the claim's scope—among context, external memory, model parameters, activation states, and tool or workflow definitions—including the option not to write persistent state, organized by three principles (substrate proportionality, temporal defeasibility, auditable continuity) and motivating evaluation criteria covering update selection, temporal consistency, interference, reversibility, efficiency, and robustness, together with a four-phase minimal benchmark unit.
FIGURE 1 Controlled knowledge updating as a substrate-selection loop. The core decision is which substrate should change for a given evidence source, scope, persistence requirement, and risk profile. The branches indicate intervention options, including mechanisms acting on the selected substrate; forgetting or attenuation can apply across substrates.
· Page 3Interpretation
The paper reframes knowledge updating as a substrate-selection problem: controlled knowledge updating is selecting the least invasive substrate sufficient for the intended scope, applying the change with provenance, and evaluating its downstream effects over time, where the correct action may be not to write persistent state. Relative to prior work centered on model editing and recall, this shifts the question from whether the system can update to what should change and how we can know the change remains appropriate, and it admits no-update as a legitimate decision. This is a Perspective, so the argument is conceptual, built on a synthesis of existing memory systems, model editing, retrieval augmentation, and benchmark work; no new experimental data, sample sizes, or effect sizes are reported.
The paper proposes three operational principles: substrate proportionality (update persistence and blast radius should match the claim's scope, evidence confidence, and expected half-life), temporal defeasibility (persistent knowledge stays revisable, with each update carrying validity conditions and a path for revision, attenuation, or rollback), and auditable continuity (future behavior remains traceable to update decisions, sources, and affected substrates). These principles translate least invasive substrate into concrete decision variables—scope, evidence confidence, expected duration, reversibility requirement, interference tolerance, and audit requirement—and describe a control loop from update candidate through evidence classification to substrate choice, forgetting, or attenuation. The principles and decision variables are defined by the authors and illustrated with examples (such as routing a one-off request for short answers differently from a changed national capital), making this framework-level argument without controlled experiments.
The paper recasts evaluation as longitudinal control, offering eight criteria with failure signals—scope fit, evidence quality, temporal validity, interference, reversibility, auditability, efficiency, and robustness—and a four-phase minimal benchmark unit: an update event, a substrate choice or refusal with justification, immediate probes for correctness and interference, and delayed probes for temporal consistency, supersession, abstention, and audit trace. Relative to existing measures of memory effectiveness, editing reliability, and task success, this design separates selection error, execution error, and maintenance error, and argues for the update event as the unit of analysis with multiple admissible substrates and reported expert disagreement. The criteria and benchmark design are presented as a table of operational tests and failure signals, forming a research agenda; no implemented benchmark results are reported.
The paper folds forgetting into the same control problem, treating it as a control action that limits the future influence of obsolete, contradicted, unsafe, or low-value state, with forms including deletion, decay, masking, reduced retrieval priority, provenance-preserving archival, or replacement by a superseding memory, and notes these choices differ in privacy, reliability, and audit consequences. Rather than treating forgetting as a memory failure or a metric alongside recall, this view argues the choice of forgetting mechanism itself needs design and evaluation, and distinguishes privacy revocation in runtime memory from training-data machine unlearning. The argument cites MemoryBank's selective preservation and forgetting and a graph-based perspective on affective conversational agents modeling reinforcement, decay, pruning, and structural reorganization as illustrations, making this a conceptual synthesis rather than new experiments.
Perspective
The framework targets agentic deployments governed by a defined update policy, where a global update is meant to apply across relevant future contexts within that deployment. It applies to settings that must decide whether new evidence belongs in context, external memory, parameters, activations, tools, or workflows, or should be refused or forgotten—for example, a cross-session assistant routing a one-off preference, a repeatedly confirmed user-level preference, a single-turn override, and an unverified instruction in a retrieved document to different substrates. The authors note that substrate selection is not a universal mapping from evidence type to mechanism but depends on domain risk, privacy constraints, deployment policy, computational cost, and acceptable expert disagreement; modifying a shared knowledge base or model may affect other deployments, and cross-system updates require authority over the shared resource, provenance, conflict resolution, and coordinated rollback, while a protocol for coordinating independently managed systems is explicitly outside the present scope.
As a Perspective, the paper reports no experimental validation, so whether the three principles and eight criteria can distinguish selection, execution, and maintenance errors in real deployments remains an open question. The reference policy the authors sketch stays qualitative; how thresholds are set, by whom, and how rule-based, learned, or hybrid triage policies perform are posed as a research agenda rather than settled results. The evaluation design allows multiple admissible substrates and calls for reporting expert disagreement, but how that disagreement should affect scoring and deployment decisions is not yet answered. In addition, the text available here is an incomplete scope: the content of Figure 1 and some reference entries are missing, so the figure's control-loop details and some citation correspondences cannot be confirmed from this reading.
