Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

DAGO uses a contextual bandit to pick multi-parent workflow fusions, raising the macro-average from 80.3 to 81.7 across six benchmarks while cutting search spend by 11.2%

DAGO models each parent-workflow combination as a bandit arm represented by pretrained embeddings of its constituent workflows' code and prompts, uses a diagonal LinUCB policy to trade off predicted offspring quality against uncertainty-driven exploration, has an LLM generate a child through summary-guided fusion with the child's validation score as the reward, and across six benchmarks in mathematical reasoning, code generation, and question answering achieves the highest macro-average score among the evaluated baselines, improving over AFlow from 80.3 to 81.7 under matched validation-evaluation budgets while reducing aggregate search expenditure by 11.2%.