Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

DyadMem benchmark released: 20 models stay strong on Gold-Memory QA but drop sharply on Full-Pipeline QA, exposing low Capture recall, incomplete Recall, and unsafe deletion

The work introduces the DyadMem benchmark and a new definition, User-conditioned Relational Agent Memory (URAM), jointly annotating user-side memory and URAM along the same multi-session trajectories, yielding 6 memory categories, 3,065 episodes, 50,961 sessions, and 61,210 QA instances with Gold-Memory and Full-Pipeline QA settings; across 16 open-weight and 4 proprietary models, Gold-Memory QA is consistently strong while Full-Pipeline QA drops sharply, quantitative results further reveal low Capture recall, incomplete Recall, and unsafe deletion even in frontier LLMs, and a validation experiment on URAM shows positive effects for all 20 models.