Public articles linked to the same research event.
arXiv The work introduces the DyadMem benchmark and a new definition, User-conditioned Relational Agent Memory (URAM), jointly annotating user-side memory and URAM along the same multi-session trajectories, yielding 6 memory categories, 3,065 episodes, 50,961 sessions, and 61,210 QA instances with Gold-Memory and Full-Pipeline QA settings; across 16 open-weight and 4 proprietary models, Gold-Memory QA is consistently strong while Full-Pipeline QA drops sharply, quantitative results further reveal low Capture recall, incomplete Recall, and unsafe deletion even in frontier LLMs, and a validation experiment on URAM shows positive effects for all 20 models.
The work introduces the DyadMem benchmark and a new definition, User-conditioned Relational Agent Memory (URAM), jointly annotating user-side memory and URAM along the same multi-session trajectories, yielding 6 memory categories, 3,065 episodes, 50,961 sessions, and 61,210 QA instances with Gold-Memory and Full-Pipeline QA settings; across 16 open-weight and 4 proprietary models, Gold-Memory QA is consistently strong while Full-Pipeline QA drops sharply, quantitative results further reveal low Capture recall, incomplete Recall, and unsafe deletion even in frontier LLMs, and a validation experiment on URAM shows positive effects for all 20 models.
The work introduces the DyadMem benchmark and a new definition, User-conditioned Relational Agent Memory (URAM), jointly annotating user-side memory and URAM along the same multi-session trajectories, yielding 6 memory categories, 3,065 episodes, 50,961 sessions, and 61,210 QA instances with Gold-Memory and Full-Pipeline QA settings; across 16 open-weight and 4 proprietary models, Gold-Memory QA is consistently strong while Full-Pipeline QA drops sharply, quantitative results further reveal low Capture recall, incomplete Recall, and unsafe deletion even in frontier LLMs, and a validation experiment on URAM shows positive effects for all 20 models.
The work introduces the DyadMem benchmark and a new definition, User-conditioned Relational Agent Memory (URAM), jointly annotating user-side memory and URAM along the same multi-session trajectories, yielding 6 memory categories, 3,065 episodes, 50,961 sessions, and 61,210 QA instances with Gold-Memory and Full-Pipeline QA settings; across 16 open-weight and 4 proprietary models, Gold-Memory QA is consistently strong while Full-Pipeline QA drops sharply, quantitative results further reveal low Capture recall, incomplete Recall, and unsafe deletion even in frontier LLMs, and a validation experiment on URAM shows positive effects for all 20 models.