AECG tops 11 of 12 framework-benchmark settings across three multi-agent frameworks and four benchmarks, with scope-preservation ablation dropping accuracy by up to 16.89 points
Related research and updatesSynopsis
The authors introduce AECG, a multi-agent memory governance framework that combines scope preservation, asymmetric evidence admission, dual-timescale reliability estimation, and downstream-exposure-aware budgeted review to turn memory from static storage into a reliability-governance loop; across three frameworks (AutoGen, MacNet, AgentNet) and four benchmarks (PDDL, FEVER, ScienceWorld, MBPP-Plus), AECG achieves the best score in 11 of 12 settings, improving over the strongest memory baseline by up to 10.23 percentage points, while removing scope preservation reduces accuracy by up to 16.89 points.
Figure 1: Conceptual comparison of memory maintenance: coarse accumulation or reactive revision risks persistent errors and utility loss, whereas AECG targets governance for reliable reuse.
arXivInterpretation
AECG reframes multi-agent procedural memory from passive accumulation into a lifecycle-governance loop comprising scope preservation, asymmetric evidence admission, dual-timescale reliability estimation, exposure-aware budgeted review, evidence-routed intervention, and paired-replay activation. Prior work emphasizes experience representation, skill abstraction, and retrieval; AECG instead governs reliability decay under continued reuse and cross-level propagation. The method is instantiated across three frameworks and four benchmarks with ablations and mechanism experiments; execution uses gpt-4.1-mini, memory-layer calls use gpt-4o-mini, and results average three independently seeded runs.
Scope preservation is the most consequential ablation component: removing it drops ScienceWorld success by 16.89 points and MBPP-Plus by 8.34 points; flattening scope causes 36.5% of selected uses to cross their evidence-supported coordination scope, and these cross-scope uses have average failed-task exposure of 14.8. The result associates scope collapse with lower success and broader structural reach, and shows Flat-scope reaches only 51.29 on ScienceWorld, below the No-Memory baseline of 55.68. Ablations run on AgentNet–ScienceWorld and MBPP-Plus with a shared post-warm-up memory snapshot and matched conditions; exposure is explicitly defined as a structural reachability measure, not a causal propagation estimate.
Dual-timescale beliefs overcome historical-success inertia: AECG reviews a corrupted skill after 3.8 uses on average versus 9.7 for the cumulative-only variant, a 60.8% reduction in detection delay; polluted-skill failed uses drop from 17.9 to 6.4 and post-corruption success rises from 58.2% to 64.5%. Because both receive the same historical evidence, the delay reduction isolates the role of the fast belief, preventing recent contradictory outcomes from being overwhelmed by accumulated success. Controlled-corruption experiments run on AgentNet–ScienceWorld with four injected corruptions validated on three matched tasks, plus a No-Governance reference point (31.2 polluted-skill failed uses, 49.3% success).
Downstream exposure improves high-impact prioritization under a fixed review budget: exposure-aware ranking covers 81.5% of polluted-skill predicted exposure in the first two reviews versus 53.2% for risk-only and 49.8% for random; post-corruption success is higher by 4.3 and 8.4 points respectively, and realized failed-task exposure drops to 22.4 versus 48.6 and 61.2. All methods share the same confidence gate, corruption set, intervention rule, and review budget, so the gain isolates the value of prioritization itself. Budgeted-triage experiments allow at most one review per round with qualified candidates exceeding the budget; Coverage@2 is a ranking diagnostic and post-corruption success is the primary outcome.
Perspective
The work targets multi-agent task streams characterized by execution graphs, suited to deployments that reuse procedural knowledge across tasks under limited review compute. Scope preservation, dual-timescale beliefs, and exposure-aware triage can respectively constrain reuse boundaries, detect degradation early, and prioritize high-propagation-risk skills under a fixed budget. The method depends on calibration of localized feedback, and exposure is defined as a structural reachability proxy rather than a causal influence estimate, so its conclusions apply to structural risk ranking and governance scheduling rather than causal attribution.
Readers should still watch: failure attribution relies on a dual-LLM consistency gate whose calibration quality across task domains remains to be observed; how well exposure as a structural proxy approximates true causal propagation is an open question; controlled-corruption and budgeted-triage experiments concentrate on AgentNet–ScienceWorld, leaving mechanism behavior under other frameworks and task families to further validation; and although hyperparameters are fixed, their stability at larger scale or over longer task streams still needs testing.
