Skip to main content
Back to timeline
arXivSource publication:

MemCo pairs local and global memory so LLM agents gain 43.47% held-out ALFWorld accuracy and 32.10% fewer steps in unseen environments

Related research and updates

Synopsis

The authors propose MemCo, a memory-centric collaboration framework in which each agent keeps an environment-specific local memory graph while evidence-aggregated, Wilson-lower-bound-filtered transferable workflows are promoted to a shared global memory, and at decision time local and global memories are retrieved and grounded by current state, task phase, and goal; across ALFWorld, PDDL, FEVER, and ScienceWorld with multiple LLM backbones, MemCo attains the highest accuracy in 10 of 12 evaluation settings and the fewest interaction steps in 9, e.g., on Qwen3-4B it improves ALFWorld held-out layout accuracy by 43.47% and reduces steps by 32.10% relative to the runner-up baseline.

Source-provided article image: MemCo: Memory-Centric Collaboration for Generalizing LLM Agents to Unseen Environments
Figure 1 ·

Figure 1: (a) Memory-centric collaboration among independently acting agents: each agent preserves environment-specific local experience while contributing reusable experience to global memory for use across environments. (b) and (c) ALFWorld performance on familiar and held-out layouts, respectively, with Qwen3-32B. Markers and error bars denote the mean and standard deviation.

arXiv

Interpretation

MemCo splits cross-environment experience sharing into complementary local and global memory: the local memory graph preserves environment-specific context, while global memory keeps only role-abstracted transferable workflows and preconditions. Existing shared-memory methods pool episodic memories across tasks or agents, so retrieval is either too specific to preserve current grounding or too coarse to support the next action; MemCo keeps environment-specific details local and promotes only transferable relations. On ALFWorld, Local-only reaches 89.64% held-out layout accuracy, Global-only 97.14%, and the full system 97.86%, with the full system using fewer steps than either isolated source on both splits; these are controlled comparisons reusing the same frozen local-memory snapshot.

Global-memory admission uses a one-sided Wilson lower confidence bound that accounts for both empirical supporting rate and evidence size, with deterministic exact-key merging of abstract records. Using the empirical support ratio as the admission score yields 736 global records, whereas Wilson promotion reduces the bank to 426 records while improving familiar- and held-out-layout accuracy by 2.27 and 1.43 percentage points, respectively, with fewer steps. The two variants inject nearly the same number of entries per decision, so the difference is not due to prompt size alone; raising the threshold from 0.25 to 0.50 shrinks the global bank from 735 to 294 records while accuracy and trajectory length remain insensitive.

Query-guided memory support scores entries along state, task phase, and goal facets, applies applicability filtering before ranking, and binds global record roles to the current environment during composition. Fixed concatenation of more memory performs worse: L5+G5 trails full MemCo by 10.71 and 16.07 percentage points on the familiar and held-out layout splits, indicating that selection and composition, not memory volume, drive the benefit. Removing any one relevance facet or retaining only a single facet lowers accuracy and increases steps on both splits; replacing explicit matching with cosine similarity or BM25 also underperforms the full method.

Across four heterogeneous benchmarks and multiple LLM backbones, MemCo improves task success while generally reducing redundant exploration. Relative to the runner-up baseline, Qwen3-4B improves ALFWorld held-out layout accuracy by 43.47% with 32.10% fewer steps; GPT-4o-mini improves FEVER accuracy by 12.54% with 58.45% fewer steps; on ScienceWorld MemCo achieves the highest final success and score. MemCo achieves the highest accuracy in 10 of 12 evaluation settings and the fewest steps in 9; exact two-sided Wilcoxon signed-rank tests on the 12 matched configuration means, after Holm correction, indicate improvements over each baseline in accuracy and steps, under the stated block-level assumptions.

Perspective

The result targets interactive agents that make multi-step decisions from text or symbolic observations, in a setting where each agent accumulates experience in its own source environment and is then evaluated on a merged test set covering both its acquisition domain and environments it has not interacted with; ALFWorld's Familiar and Held-out correspond to the official valid_seen and valid_unseen layout splits. For researchers and engineers aiming to reduce redundant exploration and deploy agents under limited interaction budgets, MemCo provides a reusable layered-memory and query-guided composition pipeline, with code released. During evaluation both local and global memories are frozen and test trajectories are not written back, so it describes inference-time experience reuse rather than online continual learning.

Global memory relies on initial interaction traces to bootstrap, so the cold-start phase offers limited global guidance; the authors note this characteristic is common to all retrieval-based experience-sharing systems. Evaluation covers only text-based and symbolic interactive benchmarks, and extending to multimodal environments with visual observations remains an open direction. The controlled ablations reuse the same frozen local-memory snapshot, so they are within-snapshot comparisons rather than independent repetitions of the main experiments; the cross-setting significance analysis treats matched backbone-task configurations as blocks, and because configurations reuse backbones and task families, dependence between blocks may limit the inferential interpretation. The one-sided confidence interpretation of the Wilson bound relies on a binomial sampling assumption, and episode-level reduction does not establish independence or identical distribution across episodes.

Sources