Skip to main content
Back to timeline
arXivSource publication:

ZJU team releases EMem-Bench: across 2,554 long-horizon embodied episodes the strongest model Gemini-3-Flash reaches only 64.2% average success, while its EMem memory system lifts GPT-5.4-mini by 23.6 points

Synopsis

The OmniAI Group of ZJU ACES Lab introduces EMem-Bench, a long-horizon embodied memory benchmark of 2,554 executable episodes spanning Passive Observation, Dynamic Tracking, Interaction Failure, and Experience Generalization, together with an external memory system EMem and an 8B policy EMem-8B; evaluating 16 open-source and proprietary MLLMs plus representative multimodal memory systems shows the strongest proprietary model Gemini-3-Flash reaches only 64.2% average success rate and most open-source models fall below 45%, while EMem achieves the best overall performance among matched backbones, improving Mistral-Small-3.1-24B by 16.4 points and GPT-5.4-mini by 23.6 points, with EMem-8B improving over Qwen3-VL-8B by 20.6 points.

AI-generated editorial illustration: EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks

Interpretation

By manually inspecting 100 failed trajectories of Gemini-3-Flash and Qwen3-VL-32B on EmbodiedBench's EB-ALF and EB-Hab, the paper identifies four embodied-memory deficiencies: forgetting fine-grained visual cues (35%), overlooking environmental changes (18%), failing to record world state revealed by interaction outcomes (31%), and failing to generalize from prior experience (7%), totaling about 91%. Prior work often attributed long-horizon embodied failures to planning or perception; this work maps failure modes directly onto four separately measurable memory capabilities and designs benchmark task families accordingly. Based on manual inspection and percentage statistics of 100 failed trajectories, a diagnostic form of evidence with a limited sample but clear category definitions.

EMem-Bench contains 2,554 episodes split into Passive Observation (1,036), Dynamic Tracking (1,052), Interaction Failure (263), and Experience Generalization (203), spanning 1,118 scenes, 125 visible object types, 83 target types, and 33 receptacle types; each episode consists of a multimodal interaction history, a target task, and a feasible action space, with success judged by the simulator's terminal state. Compared with MemoryAgentBench, Evo-Memory, WorldMemArena, EmbodiedBench, FindingDory, LMEE-Bench, SpaMEM, and WorldLines, EMem-Bench is the first to jointly cover visual, dynamic, interaction-outcome, and experience-generalization memory demands within embodied interaction. The construction pipeline combines automated checks with manual screening by five domain experts; 143 candidates were rejected, and all 2,554 retained episodes pass both automated execution checks and manual review.

Evaluating 16 open-source and proprietary MLLMs plus representative multimodal memory systems shows the strongest proprietary model Gemini-3-Flash reaches only 64.2% average success rate, all but one open-source model score below 45%, and among 400 sampled failures 61.3% begin when the model acts on a memory contradicted by its interaction history. The results turn long-horizon embodied memory from a qualitative description into a quantifiable capability profile, and show that embodied-specialized models such as RynnBrain-8B (3.6%) and Cambrian-S-7B (0.8%) do not necessarily outperform general-purpose models. Based on success rate measured through unified simulator execution and terminal-state judgment, complemented by Error Recurrence Rate and failure-type analysis.

EMem organizes embodied experience into spatial, event, and scene memories and achieves the best overall performance among the evaluated memory systems under matched backbones: 58.9% average success rate, 21.1 points above the strongest alternative TeleMem, with ERR reduced by 11.4 points; it improves Mistral-Small-3.1-24B by 16.4 points and GPT-5.4-mini by 23.6 points, while EMem-8B improves over Qwen3-VL-8B by 20.6 points. Compared with MIRIX, MemVerse, TeleMem, and MMA, EMem's three memories map onto different task families, and ablation shows removing any memory lowers average success rate, with each memory playing a task-specific role. Compared under fixed backbones and a unified evaluation protocol, with ablation studies and cross-benchmark results on EB-ALF, MMMU-Pro, BLINK, and HallusionBench.

Perspective

This work targets researchers and system developers building agents for long-horizon embodied tasks in simulation: EMem-Bench provides 2,554 executable episodes and a unified evaluation protocol, EMem provides an external memory module pluggable into different backbones, and EMem-8B provides an 8B-scale open policy baseline. The paper explicitly states that the benchmark and cross-benchmark evaluations are conducted primarily in simulation, so the results apply to controlled simulated embodied interaction settings rather than real-robot deployment.

The Limitations section notes that simulation does not capture sensor noise, actuation error, open-world changes, or ambiguous human feedback, so the reported results do not yet establish long-term memory performance on physical robots. In addition, although the loaded text is the full paper, some appendix figures (such as Figure 2, Figure 3, Figure 7, and Figures 10 through 14) exist as images whose numerical details require consulting the original figures; the four task families also differ substantially in size (Interaction Failure 263, Experience Generalization 203), and while equal-family averaging prevents the larger families from dominating, the stability of results for the smaller families warrants continued observation in future work.

Sources