EvoDuet lets LLMs retrieve the web by knowledge gap during evolutionary search, raising OpenEvolve's normalized discovery gain from 74.1% to 78.0% (GPT-5.6-Luna) and from 61.3% to 82.3% (Gemini-3.8-Flash) across 21 optimization tasks and surpassing prior best scores on eight
Synopsis
The work introduces EvoDuet, a bi-level method that co-evolves solutions and web search queries with fixed model parameters: at each iteration a knowledge-gap-based retrieval gate decides whether to retrieve new documents, reuse stored ones, or proceed without them, an inner loop refines queries and ranks documents by the solution scores they are predicted to yield, and an outer loop generates candidates in parallel and records evaluated outcomes for later searches; across 21 optimization tasks with one candidate per iteration it raises OpenEvolve's normalized discovery gain from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash, Qwen3.
Interpretation
Web search is recast from a tool inside the solution loop into a bi-level optimization that co-evolves with the solution: the inner loop optimizes queries, the outer loop optimizes solutions, and the two are coupled by a knowledge-gap-based retrieval gate, with model parameters and prompt templates fixed throughout. Earlier scaffolds such as EvoX add a strategy loop but still draw only on the run's evolutionary history and the model's parametric knowledge, so the search is closed; DeepEvolve adds web search but couples it sequentially rather than in a bi-level formulation. EvoDuet lets the model's self-assessed knowledge state decide when to search and what to ask, and ranks documents by predicted candidate scores. The paper formalizes the bi-level objective (Eq. 1) and gives pseudocode, and compares random gating, a stagnation heuristic, and the knowledge-gap gate on Denoising and Erdős, where the knowledge-gap gate achieves the highest NDG (84.5% and 100.0%); on Denoising and Sums/Diffs it compares no search, joint-level, and bi-level coupling, with bi-level reaching 84.5% and 76.0%.
Across 21 optimization tasks with one candidate per iteration, EvoDuet improves OpenEvolve's overall normalized discovery gain from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash, with parallel generation adding further gains. The benefit is not universal: Qwen3.5-9B declines by 14.4% at one candidate per iteration and still loses 4.7% with parallel generation, even though it gains 10.3% on average when given oracle documents. The paper attributes this contrast to EvoDuet's additional demand that the model decide when and what to search for, and reports that in 50 sampled revisions 6% left methods unused and 20% implemented them incorrectly. Results cover 21 tasks, three models, and two candidate budgets (1 and 8); under parallel generation the average NDG on eight tasks rises from 89.6% to 95.7% for GPT-5.6-Luna and from 84.5% to 89.6% for Gemini-3.8-Flash.
The best runs surpass previously reported best scores on eight tasks and match them on three more across five domains, with a mean cost of $45.56 across the eleven reported runs; on Denoising, EvoDuet reaches 100.27% NDG at $38.45, below SimpleTES's estimated API-equivalent cost of about $265.39. The paper also explains the mechanism: retrieved documents are most often used for method transfer (55 of 82 runs), followed by using published best scores as reference targets (40 runs); public artifact reuse occurs in only 6 runs, and only two of the eleven SOTA results reuse public artifacts in their best programs. Table 1(a) lists prior SOTA and EvoDuet scores with costs per task, and Appendix K provides the complete source and design notes for eleven best programs, noting that eight improve on the reference and three match it within tolerance.
The link between retrieved documents and solution improvement is quantified: when the retained set contains at least one document whose hypothetical evidence score exceeds the parent's actual score, the selected candidate improves on its parent in 67.1% of iterations and sets a new run best in 14.1%, versus 52.1% and 9.6% without documents; yet 20% of retrievals predicted to help still produce a worse child. This supports using predicted scores as a surrogate for the inner objective while showing that predicted gains must still be validated by the evaluator; the paper also reports that across 4,324 retrievals hypothetical evidence scores correlate strongly with evaluated scores, with positive correlations on all 21 tasks. The statistics come from a behavior audit of 82 GPT-5.6-Luna runs and 8,200 iterations, plus a failure analysis of 33,784 iterations from 362 EvoDuet runs; the Galileo case gives a concrete counterexample with a predicted score of 0.63701, a parent at 0.63692, and a child at 0.37713.
Perspective
The result targets computational optimization tasks scored by deterministic evaluators and is meant for researchers and engineering teams who want to attach web retrieval to existing evolutionary scaffolds; the paper shows it can be packed with OpenEvolve, Top-K, and EvoX and that it yields its largest average gain on mathematics tasks (7.1% across six model/budget settings). It relies on the model to assess its own knowledge gap and to implement retrieved methods, so it suits backbones that can both find and apply evidence; the paper suggests training and evaluating evidence selection and implementation as separate skills. Gains on algorithm engineering tasks (AHC039, AHC058) are not consistent, with an average decline of 1.6% across six settings.
The paper itself notes that EvoDuet's benefit varies with the backbone model, that Qwen3.5-9B falls below OpenEvolve alone at both candidate budgets with failure modes including unused retrieved methods and incorrect implementations, and that gains on algorithm engineering tasks are inconsistent. Some best programs also depend on public artifacts (for example, the Erdős row retains a runtime download of a published witness), which the paper flags in Appendix K and accompanies with recommendations for independent validation and source attribution. Retrieved information may be unreliable, and reuse of published solutions can raise novelty concerns; both warrant continued attention in downstream use.
