Skip to main content
Back to timeline
arXivSource publication:

DyadMem benchmark released: 20 models stay strong on Gold-Memory QA but drop sharply on Full-Pipeline QA, exposing low Capture recall, incomplete Recall, and unsafe deletion

Related research and updates

Synopsis

The work introduces the DyadMem benchmark and a new definition, User-conditioned Relational Agent Memory (URAM), jointly annotating user-side memory and URAM along the same multi-session trajectories, yielding 6 memory categories, 3,065 episodes, 50,961 sessions, and 61,210 QA instances with Gold-Memory and Full-Pipeline QA settings; across 16 open-weight and 4 proprietary models, Gold-Memory QA is consistently strong while Full-Pipeline QA drops sharply, quantitative results further reveal low Capture recall, incomplete Recall, and unsafe deletion even in frontier LLMs, and a validation experiment on URAM shows positive effects for all 20 models.

Source-provided article image: DyadMem: A Long-Term Memory Benchmark of How Agents Work with Users
Figure 1 ·

Figure 1: DyadMem is a dual-domain, full-pipeline memory benchmark. Across multi-session user–agent interactions, Capture extracts current-session memories, Update reconciles them with the prior bank, Recall selects valid terminal- bank entries, and QA uses the selected evidence. The evolving bank contains shared episodic memory, user-side memory, and User-conditioned Relational Agent Memory (URAM); unlike general agent memory that transfers across users, URAM records how this particular agent should work with this particular user.

arXiv · Page 1

Interpretation

Introduces User-conditioned Relational Agent Memory (URAM), a new definition that makes explicit the relationship-specific memory of how a particular agent should work with that user, jointly annotated with user-side memory along the same multi-session trajectories to form 6 memory categories. Existing benchmarks primarily supervise user facts and preferences or experience reusable across users, leaving relationship-specific agent memory implicit; DyadMem annotates it as a distinct object within the same trajectories. The abstract states that joint annotation proceeds along the same multi-session trajectories and yields 6 memory categories, which is evidence at the level of benchmark definition and annotation design.

Builds a benchmark with 3,065 episodes, 50,961 sessions, and 61,210 QA instances, featuring session-level Capture and Update gold annotations, query-level Recall support, and two QA settings: Gold-Memory and Full-Pipeline. Most prior works measure the model solely with final-answer QA over long interaction histories, making assessment incomplete and unreliable; DyadMem replaces single final-answer evaluation with dual-domain, full-pipeline annotation and two QA settings. The abstract provides the scale figures and annotation types and explicitly names the two QA settings, which is evidence at the level of benchmark resources and evaluation protocol.

Across 16 open-weight and 4 proprietary models, Gold-Memory QA is consistently strong while Full-Pipeline QA drops sharply, and this gap supports the fine-grained evaluation design. The result indicates that looking only at final answers can mask problems in the memory stages of the full pipeline, providing empirical grounds for staged evaluation. The abstract reports comparative results across 20 models, which is evidence at the level of multi-model benchmark evaluation.

Quantitative results reveal low Capture recall, incomplete Recall, and unsafe-deletion issues even in frontier LLMs; a further experiment validating URAM observes positive effects for all 20 models. It localizes memory failures to specific stages such as Capture, Recall, and deletion, and provides cross-model positive evidence for the effectiveness of URAM. The abstract describes these as 'several quantitative results' and a 'rigorous experiment', which is evidence at the level of benchmark diagnosis and validation.

Perspective

The benchmark targets research and development on long-term agent memory and applies to evaluation settings where agents collaborate with users across many sessions; its dual-domain, full-pipeline design with Gold-Memory and Full-Pipeline QA settings can separate memory-stage performance from final-answer performance, and the session-level Capture and Update gold annotations plus query-level Recall support help localize specific failure stages. The positive URAM validation results cover all 20 models, suggesting the definition can serve as an evaluation and training signal for relationship-specific memory.

The abstract does not give the concrete scores for Gold-Memory and Full-Pipeline, the magnitude of the drop, the per-model Capture recall and Recall completeness values, the criteria for unsafe deletion, or the specific design of the URAM validation experiment; the basis for the 6 memory categories, the data sources, and the domain coverage are also not expanded in the abstract. These are information boundaries at the abstract level, and readers who need to reproduce or compare results should consult the tables and experimental setup in the full text.

Sources