Skip to main content
Back to timeline
arXivSource publication:

CogMem reconstructs long-term dialogue memory with a PEC2F graph and agentic retrieval, reaching top accuracy on LoCoMo and LongMemEval

Synopsis

The work introduces CogMem, a training-free cognitive memory architecture built on the PEC2F (Person-Event-Concept-Claim-Fact) graph schema: dedicated Claim nodes preserve the source and target of subjective statements, dialogue turns are incrementally converted into provenance-aware graph records and consolidated into higher-level facts, and a rule-based controller driven by LLM intent parsing composes four deterministic graph operators—anchoring, traversal, intersection, and evidence grounding—to reconstruct query-relevant context; on LoCoMo and LongMemEval it performs strongly on multi-hop, temporal, and knowledge-update tasks, with ablations and a semantic-collapse probe supporting complementary contributions from epistemic separation, consolidation, and agentic retrieval.

Source-provided article image: From Retrieval to Reconstruction: Constructing Evolvable Cognitive Memory for Long-Term Dialogue
Figure 1 ·

Figure 1: The Paradigm Shift from Passive Retrieval to Active Reconstruction. Left (traditional RAG and flat memory): passive retrieval over coarse text chunks introduces noise and may miss relevant counter-evidence. Even when a related statement is retrieved, the flat representation obscures its source and can cause the LLM to treat an attributed opinion as an unqualified record. Right (CogMem): a PEC 2 F cognitive graph supports active reconstruction, allowing the agent to attribute conflicting viewpoints to their sources and synthesize an evidence-grounded answer.

arXiv

Interpretation

The PEC2F schema structurally separates source-attributed subjective beliefs from unattributed event/fact records via Claim nodes, mitigating Semantic Collapse. Prior RAG and graph-retrieval systems such as LightRAG and HippoRAG use task-agnostic entity-relation graphs that do not explicitly distinguish source-attributed claims from event and fact records; CogMem models 'Jon was fired' and 'Gina thinks Jon being fired was a mistake' as different node types and a two-hop path. In 150 synthetic fact-claim pairs, mean BGE-M3 cosine similarity is 0.8231 and flat RAG misattributes in 64/150 cases (42.7%), while type-constrained PEC2F retrieval reduces this to 9/150 (6.0%); removing Claim nodes drops LongMemEval Knowledge Update accuracy from 73.08% to 56.67%.

Offline memory consolidation abstracts sparse episodic traces into higher-level Fact nodes while Evidence Mounting preserves provenance links, enabling multi-resolution retrieval. Unlike memory systems that only store or prune, CogMem triggers an LLM abstraction when an event cluster reaches a threshold (3 events, 30-day window), producing semantic facts with a temporal envelope while retaining original episodes via SUPPORTED_BY edges. Consolidation improves Multi-Hop F1 by 4.86 points (43.56% to 48.42%) and reduces average reasoning steps from 3.42 to 2.67; the graph gains 540 nodes and 4,708 edges, an intentional additive evidence-mounting design.

A ReAct-based Cognitive Search Agent uses LLM intent parsing to drive a rule-based controller that dynamically composes four deterministic graph operators for active recall. Shifts retrieval from a static pipeline to agentic active recall: comparison words trigger intersection, attribution questions trigger Claim relations, time expressions trigger temporal gating, and verification questions trigger evidence grounding. Replacing the agent with a fixed pipeline drops Multi-Hop F1 from 48.42% to 22.81% (a 25.61-point fall) and reduces weighted overall F1 by 65.4% relative; the agent averages 2.65 tool calls versus 3.00 for the fixed flow.

On the LoCoMo and LongMemEval long-term memory benchmarks, CogMem achieves the highest or leading results under both backbones, especially on multi-hop, temporal, and knowledge-update tasks. Relative to Naive RAG, Mem0, A-MEM, LightMem, GAM, LightRAG, HippoRAG, and zero-shot RoG, CogMem shows clear gains on categories requiring structured reasoning. With Qwen2.5-14B, Multi-Hop F1 is 48.42% (GAM 42.96%, RoG 41.80%, HippoRAG 34.78%); LongMemEval accuracy is 68.40% (GPT-4o-mini) and 66.50% (Qwen2.5-14B); in a 50-example blind human preference study CogMem wins 68%, ties 20%, and loses 12%.

Perspective

The architecture targets long-term dialogue agents that need to reconstruct evidence across sessions, applies to the settings represented by English multi-session dialogue benchmarks, and is validated with GPT-4o-mini and Qwen2.5-14B backbones. It makes retrieval results traceable to original utterances, and Claim reconciliation occurs only when source, target, and predicate match, with claims from different speakers never merged, so it suits settings that must preserve disagreement and provenance. For readers, this means that if a system must distinguish 'who said what' from 'what happened,' the PEC2F node separation and Evidence Mounting offer a directly reusable design; for deployment cost, construction uses about 1,578k tokens and inference averages 6,610 tokens per query, which can serve as a resource-planning reference.

Evaluation covers two English long-term dialogue benchmarks and two backbones and does not establish generalization to multilingual dialogue, noisy real-world conversations, or continuously deployed assistants; the semantic-collapse probe is synthetic, and the human preference study has only 50 examples, two primary annotators, and one adjudicator, so it complements rather than replaces larger independent human evaluations. LongMemEval relies on an LLM judge, which may inherit model-specific biases. The error analysis shows temporal normalization accounts for 44% of temporal errors, vector semantic drift for 23% overall, consolidation-related granularity loss for 11%, and incomplete search scope for 15%, which are open questions to watch when transferring the approach. Consolidation currently runs periodically and offline, leaving real-time incremental consolidation and neural tool learning as future directions.

Sources