SEDIMA adds cross-run hierarchical insight memory to evolutionary search agents, raising average final performance by 5.5% on AlgoTune and 6.6% on ALE-Bench LITE
Related research and updatesSynopsis
The work introduces SEDIMA, a persistent hierarchical insight memory for LLM-driven evolutionary search agents that distills raw traces into natural-language insights, clusters them by semantic similarity using attention-weighted centroids, and retrieves relevant guidance to condition future mutations, accumulating transferable knowledge across runs and problems; as a drop-in module that leaves search operators unmodified, it improves average final performance by 5.5% on AlgoTune and 6.6% on ALE-Bench LITE under a fixed budget of 100 evaluated candidates, and under OpenEvolve requires 32.3% fewer iterations on average to reach baseline-best performance across the five evaluated backbones.
Figure 1: Overview of Sedima . A persistent three-level memory (right) augments an evolutionary search loop (left, dashed): evaluated children are distilled into hierarchical insights on save , and relevant insights are injected into the mutation prompt on retrieval .
arXivInterpretation
SEDIMA distills raw search traces into natural-language insights and clusters them by semantic similarity using attention-weighted centroids, forming a persistent hierarchical memory. Existing LLM-driven evolutionary search systems are largely memoryless, exploring from scratch each run so agents repeatedly rediscover the same improvements and re-encounter the same dead ends; SEDIMA moves knowledge accumulation from within a single trajectory to across runs and problems. The abstract describes the mechanism (distillation, attention-weighted centroid clustering, retrieval of guidance to condition mutations) and supports it with quantitative results on AlgoTune, ALE-Bench LITE, and OpenEvolve.
SEDIMA is a drop-in module that leaves search operators unmodified and improves average final performance under a fixed budget of 100 evaluated candidates: 5.5% on AlgoTune and 6.6% on ALE-Bench LITE. The gains come from the memory and retrieval layer rather than from changing the search operators themselves, indicating that cross-run knowledge reuse can pay off independently of the specific search algorithm. Average final performance improvement percentages on two benchmarks, explicitly bounded by a fixed budget of 100 evaluated candidates.
Under OpenEvolve, SEDIMA requires 32.3% fewer iterations on average to reach baseline-best performance across the five evaluated backbones. The benefit extends from higher final performance to fewer iterations needed for the same level, i.e., an improvement in sample efficiency. An average iteration-reduction ratio across five evaluated backbones, an efficiency metric rather than a final-score metric alone.
Perspective
The result targets settings that use LLM-driven evolutionary search for automated program and algorithm discovery, especially agent systems that need to reuse experience across multiple runs and multiple problems. Its gains are measured under a fixed budget of 100 evaluated candidates and are reported on AlgoTune, ALE-Bench LITE, and five backbones under OpenEvolve; it therefore applies most directly to search pipelines of this kind, where the budget constraint is the number of candidate evaluations. As a drop-in module, it suits teams that want cross-run memory without modifying their search operators.
Verifiable information is limited to what the abstract states: the concrete implementation of insight distillation and clustering, how retrieval interacts with mutation operators, and the variance and per-task distribution behind the 5.5%, 6.6%, and 32.3% figures are not elaborated in this text. A careful reader would still watch whether cross-run memory remains effective when the problem distribution shifts or transfers to new problem families, how retrieval quality and cost change as the memory grows, and whether gains move in the same direction under different candidate-evaluation budgets. These are open questions for follow-up testing.
