Jev-Mem puts a System-One control plane over agentic memory, lifting LoCoMo overall score to 0.777 while cutting memory build time to 158 seconds
Synopsis
Jev-Mem introduces an agentic memory architecture inspired by System-One/System-Two cognition, in which a dedicated System-One control plane handles memory typing, relation construction, query routing, retrieval-budget allocation, graph traversal, candidate scoring, and adaptive stopping, while System Two is reserved for complex reasoning and answer synthesis; on LoCoMo it reaches an overall LLM-as-a-Judge score of 0.777 (an 11.0% relative gain over the strongest baseline), a memory construction time of 158 seconds (a 6.6 speedup over the fastest competing memory system), and an average query latency of 0.93 seconds (a 36.7% reduction).
Interpretation
The paper treats memory control itself as a first-class systems layer, proposing a three-part architecture of a System-One control plane, a shared multi-relational memory data plane, and a System-Two reasoning plane: the write path performs memory typing, redundancy filtering, and construction of semantic, temporal, causal, and entity relations, while the read path performs query routing, retrieval-budget allocation, graph traversal, candidate scoring, evidence assessment, and adaptive stopping. Existing agentic memory systems rely either on fixed heuristics or on repeatedly invoking autoregressive LLMs for free-form decisions during construction and retrieval; Jev-Mem replaces these high-frequency but bounded-output judgments with typed probabilistic decisions and lets one control abstraction span both write and read paths, rather than one set of heuristics for construction and another for retrieval. The paper provides a full algorithm design, including the typed controller interface, separation of candidate discovery from relation judgment, relation-edge insertion thresholds, budget allocation, and stopping criteria, and reproduces the typed prompts used for writing, consolidation, routing, candidate scoring, and stopping in the appendix; the authors state that returned values are model-reported and not assumed to be calibrated probabilities.
On the LoCoMo long-term multi-session conversation benchmark, Jev-Mem reaches an overall LLM-as-a-Judge score of 0.777 versus 0.700 for the strongest baseline MAGMA, an 11.0% relative improvement, performing best in five of six categories and matching the best temporal result. Gains concentrate where evidence must be combined across memories or distractors must be rejected: Multi-Hop 0.623 versus 0.569 for the strongest baseline, Open-Domain 0.618 versus 0.517, and Adversarial 0.962 versus 0.742; the paper attributes this to query-adaptive selection of relational views, budget allocation, and evaluating candidates and evidence sufficiency during traversal. Results come from Table 1, comparing against representative long-term memory approaches including Full Context, A-MEM, MemoryOS, Nemori, and MAGMA, using the same backbone answer model where applicable, with LLM-as-a-Judge correctness scoring as the metric.
On efficiency, Jev-Mem's total memory build time is 158 seconds, an 84.9% reduction (6.6 speedup) over the fastest competing memory system at 1,044 seconds, and its average query latency is 0.93 seconds, 36.7% lower than the fastest memory-based baseline at 1.47 seconds and 46.6% lower than directly processing the full context at 1.74 seconds. The paper attributes the efficiency to moving memory control off the autoregressive path: construction batches lightweight typed decisions over a bounded candidate set, and retrieval uses System-One for routing, candidate evaluation, and evidence checking while explicit search budgets and adaptive stopping constrain graph expansion; by contrast, MemoryOS requires 32.68 seconds per query. Results come from Table 2, reporting end-to-end memory construction time and average per-query latency; the paper also states that control overhead is explicitly bounded by independent limits on graph expansions, inspected edges, visited nodes, depth, controller invocations, and elapsed time.
The paper applies the System-One/System-Two distinction to the memory subsystem itself, explicitly framing it as an analogy for how computation is allocated rather than a claim that the underlying mechanisms correspond to human dual-process reasoning. Earlier analogous separations were mostly used for extra deliberation within a single reasoning episode (such as System 2 Attention, Tree of Thoughts, and Graph of Thoughts) or for routing complete requests between models; Jev-Mem applies the principle at a finer granularity, separating frequent bounded memory-control decisions from open-ended generative reasoning. The paper presents this distinction in its related work and design motivation, and notes that Jev is one concrete realization of the System-One controller while the key architectural novelty is separating lightweight memory control from deliberative reasoning across both construction and retrieval.
Perspective
This work targets persistent agents that must retain and reuse experience across sessions, such as coding assistants, personal agents, research agents, and autonomous workflows; its conclusions apply to settings evaluated as long-horizon multi-session conversational memory, namely benchmarks like LoCoMo that require temporal, causal, and cross-session reasoning. For practitioners who want to cut memory construction and retrieval overhead without sacrificing answer quality, the reusable takeaway is to make memory typing, relation construction, query routing, budget allocation, candidate scoring, and stopping into structured decisions with bounded outputs, and to let write and read paths share one control abstraction, while constraining the control process itself with explicit budgets and stopping criteria. The paper also notes that when a higher-level textual abstraction is needed, the control plane can explicitly escalate an approved merge or promotion to System Two, so generation is an optional consequence of a structured control decision rather than the default mechanism for memory maintenance.
The tables presented in the paper's main text report LoCoMo results; the abstract mentions two benchmarks, but the verifiable numbers concentrate on LoCoMo, so consistency across benchmarks remains a question readers will watch. The appendix states that the typed controller returns model-reported values not assumed to be calibrated probabilities, so the stability of threshold-based criteria (such as the 0.60 relation-edge threshold and the stopping requirement that evidence_sufficient be at least 0.95) across models or data distributions is worth continued observation. Retrieval has independent limits (depth 8, visited nodes 60, examined edges 2400, Jev attempts 16, and a 15-second time budget), and the paper notes these limits can terminate retrieval even when evidence is incomplete, with the time budget checked between operations rather than as a strict preemption guarantee. In addition, this reading is full text, but the concrete presentation of tables and figures is not fully expanded in the text, so verifying item-by-item numbers still requires returning to the original.
