EpiCon builds a shared multimodal memory bank with two 2B models, letting different agent systems reuse each other's experience and lifting macro-average scores by 1.7 to 4.9 points
Synopsis
EpiCon introduces a shared multimodal memory framework in which two independently trained 2B models, a memory controller and a tree self-organizer, co-evolve question-level textual guidance with visual evidence and link it to a persistent experience bank, so different multi-agent systems can reuse and contribute experience without updating host model parameters; across eleven benchmarks, four multimodal task domains, two harnesses and multiple backbones, a frozen bank improves other systems with a single solving attempt, a second harness raises the original system's macro-average by 2.6 points, the 2B variant improves macro-average scores by 1.7 to 4.9 points over No Memory across four host configurations, and memory-operation time drops 67% to 74% relative to backbone-sized memory models.
Interpretation
EpiCon connects question-level multimodal memory evolution to a persistent cross-question experience bank, forming collective learning without updating host parameters: within a question the memory controller jointly revises actionable textual guidance and the associated image regions in response to successive attempts and feedback, while the tree self-organizer hierarchically consolidates lessons, abstracts reusable rules, and retrieves experience for new problems. Prior work largely treats memory as textual histories, hierarchical notes, or cross-model shared experience structures, whereas this work handles joint revision of textual guidance and visual evidence together with accumulation across harnesses and backbones inside one shared bank. The paper provides the framework description, formulated update and retrieval procedures, and two independently trained 2B models; supervision comes from roughly 25K replay-screened memory transformations and roughly 6.5K structurally validated tree-operation demonstrations.
A shared experience bank keeps working beyond its builder: a bank built by Codex with Qwen3.8-27B, when reused by Codex with Gemma4-31B and by DeepSeek-Harness with Qwen3.8-27B, changes only the backbone or the harness and raises macro-average scores by 3.2 and 4.2 points respectively, with a single solving attempt and question-level updates disabled. This tests the transferability of historical experience itself rather than gains bought by extra attempts, separating experience reuse from multi-round reflection. Cross-backbone and cross-harness transfer are compared across all eleven benchmarks; for example backbone transfer improves MATH-Vision from 25.8 to 33.0 and harness transfer improves ParseBench from 52.4 to 65.7.
Experience contributed by another harness can in turn improve the original builder: after Codex and DeepSeek-Harness exchange construction and evolution roles in both directions, Codex rises from 53.5 to 55.7 and DeepSeek-Harness from 53.4 to 56.6 in macro-average, with evolved banks improving eight and nine of eleven benchmarks respectively. This gives direct evidence of collective learning, since a system can both benefit from another's experience and contribute new experience that improves the original contributor's later solving. Both directions are compared on the same 4,388 questions with one retrieval, a single solving attempt, and question-level updates disabled; the paper also notes the effect depends on task and contributing system, for example Codex on ParseBench drops from 63.4 to 58.9 after DeepSeek-Harness evolves its bank.
The compact 2B memory models trade a little accuracy for large cost savings: relative to backbone-sized memory models, memory-operation time falls about 67% to 74% and total time 25% to 35%, macro-average scores are 0.9 to 3.6 points lower, yet still 1.7 to 4.9 points above No Memory; ablations show tree organization beats flat storage (ParseBench 52.2 to 67.3) and adaptive visual injection improves seven of eight scores with either frozen or evolving visual memory. Memory management is offloaded from the backbone to separate small models, and the contributions of hierarchical organization and visual injection are quantified separately. Results span four host configurations and eleven benchmarks with normalized time and token costs; ablations fix the controller, disable historical retrieval, and share a budget of up to five attempts.
Perspective
The framework targets multi-agent systems that use external memory: the host keeps its own orchestration and backbone, while memory is carried by two 2B models and an independent experience bank, suited to tasks such as document understanding, visual-to-code generation, vision-grounded mathematics, and general visual-language reasoning that need image evidence. It lets systems reuse others' experience without updating host parameters, keeps the bank usable after a harness or backbone change, and supports different systems taking turns building and evolving the same bank.
The benefits of cross-harness evolution depend on the task and the contributing system: the paper reports Codex dropping on ParseBench after another harness evolves its bank, and DeepSeek-Harness declining from 56.6 to 53.4 on MATH-Vision, so which lessons deserve a place in a shared bank remains open. Tree-operation demonstrations pass structural checks but are not individually verified through downstream replay, and controller supervision comes from replay on source tasks, so generalization to new tasks needs more observation. Changes on general visual-language reasoning are smaller and less consistent (ReasonMap and BabyVision), marking a boundary beyond the training domains. In addition, this reading is of the full text, but tables and figures arrive as text, so the exact endpoints of some reported ranges cannot be confirmed from the text.
