Mem++ replaces write-time distillation with non-destructive memory, beating the strongest memory baseline by 8.0 to 13.1 points on OrgMemBench
Related research and updatesSynopsis
The authors propose Mem++, a non-destructive memory framework that stores every document whole with its date and author, calls no generative model at write time, and at read time retrieves only documents dated up to the time a question asks about while fusing lexical and semantic rankings; on the organizational benchmark OrgMemBench it surpasses the strongest memory system baseline by 8.0 to 13.1 points across two answering models, achieves the best overall score with gpt-4.1-mini at 2.6 points above RAG, and obtains the best average LLM-judge score on LoCoMo while ranking second on LongMemEval-S behind only its entity-graph variant.
Figure 1: Conversational vs. organizational memory. Top: a single narrator restates their own facts, and LLM extraction at ingest keeps only the surviving fact. Bottom: different authors in an organization write the same fact, and an earlier decision may still hold.
arXivInterpretation
Mem++ shifts memory design from write-time distillation to read-time selection: each document is stored whole with its date and author, and no generative model is called during writing. Most memory systems compress documents at write time into facts, notes, or graph edges, fixing what can be answered before any question is asked; Mem++ performs no such compression. The abstract presents this design through framework description and contrast, without implementation details or ablation data.
At read time Mem++ retrieves only documents dated up to the time the question asks about, and fuses lexical and semantic rankings. Rather than overwriting older versions, the system keeps them and leaves version choice to the answering model, in contrast to systems that overwrite older versions. The abstract states the retrieval and ranking-fusion mechanism but reports no separate evaluation of retrieval precision or temporal filtering.
On the organizational benchmark OrgMemBench, Mem++ surpasses the strongest memory system baseline by 8.0 to 13.1 points across two answering models. The gain is measured in long-term organizational document settings against existing memory system baselines, not merely against naive retrieval. The abstract reports a score range across two answering models but gives no per-model scores, sample sizes, or statistical tests.
With gpt-4.1-mini, Mem++ achieves the best overall score, 2.6 points above RAG; it achieves the best average LLM-judge score on LoCoMo and ranks second on LongMemEval-S, behind only its entity-graph variant. Results span the organizational benchmark and two general long-term memory benchmarks, indicating the design is not confined to a single evaluation set. The abstract gives the 2.6-point margin over RAG and rankings on two benchmarks, without absolute scores, judge settings, or confidence intervals.
Perspective
The work targets long-term records that accumulate over months in organizational settings, where revised decisions arrive as new documents, and it applies where documents carry dates and authors and questions carry a temporal reference. It lets the answering model decide at read time which version held then, rather than accepting an answer scope fixed at write time; teams that need to audit how decisions evolved, or to distinguish versions of the same topic across periods, can use this orientation directly. The abstract notes that benchmark evaluation code is available, supporting reproduction and extension on OrgMemBench and comparable long-term memory benchmarks.
The abstract gives no per-model scores, no size or composition of OrgMemBench, no separate contribution of temporal filtering versus ranking fusion, and no LLM-judge scoring setup. Ranking second on LongMemEval-S behind its own entity-graph variant suggests different memory organizations suit different conditions, and which settings favor the entity-graph variant remains open. Storing without write-time generation also means storage and read costs grow with document volume, a trade-off the abstract does not quantify.
