DAGO uses a contextual bandit to pick multi-parent workflow fusions, raising the macro-average from 80.3 to 81.7 across six benchmarks while cutting search spend by 11.2%
Related research and updatesSynopsis
DAGO models each parent-workflow combination as a bandit arm represented by pretrained embeddings of its constituent workflows' code and prompts, uses a diagonal LinUCB policy to trade off predicted offspring quality against uncertainty-driven exploration, has an LLM generate a child through summary-guided fusion with the child's validation score as the reward, and across six benchmarks in mathematical reasoning, code generation, and question answering achieves the highest macro-average score among the evaluated baselines, improving over AFlow from 80.3 to 81.7 under matched validation-evaluation budgets while reducing aggregate search expenditure by 11.2%.
Figure 1: Overview of the DAGO framework. (a) The DAG stores workflows, validation scores, and parent lineage. (b) Annealed score sampling proposes parent combinations. (c) Diagonal LinUCB selects a combination from code-and-prompt embeddings. (d) Summary-guided fusion generates a child. Validation feedback updates LinUCB and the DAG; the highest-scoring workflow is returned.
arXivInterpretation
It formulates parent selection for workflow fusion explicitly as a contextual bandit: each candidate parent combination is an arm whose reward is the generated child's validation score rather than an aggregate of the parents' existing scores. Earlier workflow-recombination approaches such as EvoFlow's LLM-based crossover and MermaidFlow's constraint-preserving crossover supply a fusion operator but do not decide which parents to fuse; DAGO makes that selection itself the object of learning. The paper gives a formal definition of the setting (arms, potential rewards, partial observability) and compares against AFlow, MaAS, and other baselines on six benchmarks, reporting a macro-average of 81.7 versus 80.3 and 79.3.
A shared diagonal LinUCB reward model reuses sparse fusion feedback across arms, paired with annealed score-based sampling that proposes candidate arms without enumerating the combination space. Instead of learning an independent reward estimate per combination, it uses concatenated pretrained code-and-prompt embeddings as arm features so feedback from evaluated arms informs the ranking of newly proposed arms. Ablations show the default LinUCB exceeds both random selection and an exploration-free variant on all three tested benchmarks; removing parent summaries or switching to random selection lowers the three-task average from 84.9 to 82.9 and 83.0.
Arm selection is embedded in a generation loop where a DAG records multi-parent lineage and an LLM summarizes parent strengths and constraints before generating a child by targeted modification starting from the strongest parent. The DAG records search lineage rather than the execution graph of an individual workflow, providing an expanding pool of parents for subsequent arm proposals. The appendix reports that all best fusion workflows on the six benchmarks come from 5-parent fusion, that best nodes appear between iterations 11 and 29, and that best workflows for MBPP, HotpotQA, and DROP have ancestor depths of 3 to 7.
Under matched validation-evaluation budgets it reports performance, search expenditure, and cross-task transfer together with a decision analysis that does not assume linear realizability. The analysis separates candidate-proposal loss from arm-selection loss and notes that the diagonal approximation changes both the fitted coefficients and the exploration term, so classical linear-bandit guarantees do not transfer directly. Total search cost falls from $86.45 for AFlow to $76.78, though it is higher on HotpotQA and DROP; average transfer regret across three task pairs falls from 3.43 to 2.80.
Perspective
The result targets LLM workflow search driven by validation feedback where evaluation is costly, applies to the three benchmark families of mathematical reasoning, code generation, and question answering, and can be paired with different execution models. The default configuration uses arm size 5, an initial population of 10, up to 30 fusion rounds, and a candidate-arm cap of 200, and it requires at least 10 initial parents; deployments with fewer initial parents would need expanded initialization or an explicit padding-and-mask convention. The lightweight diagonal LinUCB update keeps arm scoring and online updates cheap, which suits settings with a constrained validation-execution budget; the paper also specifies falsifiable diagnostics such as prequential prediction records, frozen-archive multi-arm probes, and matched-budget end-to-end controls for later checks under the same protocol.
Several numbers appear in tables while their placeholders in the main text are empty, for example the macro-average scores for AFlow and DAGO, the search-expenditure reduction, and some ablation deltas in Section 5, so the tables should be treated as authoritative. The paper itself notes that the diagonal approximation changes both fitted coefficients and the exploration term, that the exploration term is a proxy rather than a calibrated confidence interval, and that the additive representation does not model interactions between parents; Appendix A explicitly does not claim sublinear regret and notes the numerical bound can be vacuous relative to the elementary bound in the reported configuration. Test scores are means over three executions of one validation-selected workflow rather than independent optimization runs, so run-to-run variance is not characterized. MBPP and HotpotQA each contain a 0.0-scoring node, indicating that generation and runtime format checks remain open. Cross-task transfer and inference cost comparisons cover only some task pairs and three test sets, leaving behavior on more task types and larger parent pools to be observed.
