MemLife builds entity-grounded first-person text memories read by a time-indexed agentic reader, beating the strongest training-free baseline by 4.6–12.0% on four long-horizon egocentric video benchmarks, with MemOpt adding 2.7–5.0% by training only the memory writer under the FIRM reward
Synopsis
The work introduces MemLife, a multimodal memory system that compacts egocentric video into time- and entity-anchored first-person text episodes retrieved by a time-indexed agentic reader, improving over the strongest training-free baseline by 4.6–12.0% across four long-horizon benchmarks without training or query-time video access, and further proposes MemOpt, which applies reinforcement learning only to the memory writer under the faithful, informative, and retrievable FIRM reward, adding 2.7–5.0% with gains that transfer across writer and reader backbones, memory systems, and out-of-domain video and question distributions.
Interpretation
MemLife writes each 30-second video segment independently into a time- and entity-anchored first-person text episode, and an agentic reader answers by rewriting queries, searching over time intervals, and organizing retrieved evidence chronologically. Earlier memory systems either compress too aggressively and lose evidence or fail to locate relevant entries because of retrieval competition in a growing search space; MemLife separates question-independent writing from question-dependent access, combining semantic search with time-scoped fetching. On SuperMemory-VQA, EgoLifeQA, SuperMemory-LVQA, and EgoLife-EQA, MemLife reaches 56.50, 52.80, 56.18, and 52.00 accuracy, above every training-free baseline in the table; ablations show entity grounding and first-person narration each add gains, with first-person narration also consistently raising EgoLifeQA recall.
MemOpt applies reinforcement learning only to the memory writer, using the FIRM reward to jointly enforce faithfulness, informativeness, and retrievability, thereby improving the memory itself. Prior work either applies reinforcement learning to final answer quality or distills memory writing from stronger models via supervised fine-tuning; MemOpt decomposes supervision into three verifiable dimensions and uses a token-level faithfulness reward to localize spans unsupported by the source video. MemOpt raises MemLife to 60.35, 56.60, 58.91, and 57.00 across the four benchmarks, above every trained baseline in the table; ablations show retrievability alone raises recall but not consistently accuracy, adding informativeness strongly benefits SuperMemory-VQA, and the complete objective with token-level faithfulness and multiplicative aggregation achieves the highest accuracy on both benchmarks.
The writer trained by MemOpt transfers across backbones, memory systems, and distributions. This addresses whether learning what to remember is merely an artifact of one backbone or one dataset. Across combinations of Qwen3.5-9B and Qwen3.6-27B writers and readers, MemOpt improves accuracy in every setting; feeding optimized EgoRAG memories to the MemLife reader yields further gains, while the best performance comes from the trained MemLife writer; training only on SuperMemory-VQA still improves out-of-domain benchmarks such as EgoLifeQA.
MemLife offers advantages in storage and reading latency, and its default reader needs no access to raw video at question time. One practical bottleneck of long-horizon memory QA is re-watching raw clips for every query, which MemLife replaces with text memory. In the efficiency analysis MemLife uses 0.68 and 0.67 MB/h of storage and reads at 13.1 and 20.7 seconds per question, the smallest footprint and fastest reading among the listed systems; MemLife-V lets the reader sample up to 50 frames but changes results only modestly relative to default MemLife.
Perspective
The results target long-horizon memory QA over egocentric video captured by wearable devices, where questions concern the user's own past; default MemLife does not access raw video at question time, while MemLife-V stores low-resolution redacted video and samples limited frames for storage and privacy reasons. MemOpt's training supervision comes from SuperMemory-VQA subjects S1–S6, with S7–S8 for validation and S9–S10 for testing, and all other benchmarks are test-only. For readers building personal memory assistants, this means hand-designed writing principles can deliver training-free gains first, training only the writer can improve them further, and the trained writer can be reused in existing memory systems.
The faithfulness reward uses the same frozen model as evaluator, and its agreement with the source video is not independently human-verified in the text; the retrievability reward relies on reader actions cached at the start of each epoch, so how the gap between training-time and deployment-time reader behavior affects final memory quality remains an open question; EgoLife-EQA was constructed by the authors from official captions and verified by two annotators, and boundary cases beyond their 92.1% agreement are not elaborated; additionally, reader training improves in-domain results markedly but does not consistently raise accuracy on out-of-domain EgoLifeQA, and making reader gains consistent under distribution shift is left as future work.
