Skip to main content
Back to timeline
arXivSource publication:

DAGent grows its task DAG batch by batch from node-level evidence, beating the strongest open-source baseline by 5.3/5.8/2.0 points on BrowseComp-Plus, GAIA and xbench-DeepSearch

Synopsis

DAGent introduces Evaluate-then-Grow incremental planning, in which an Orchestrator grows the task DAG one batch at a time conditioned on confidence and uncertainty signals from completed nodes, paired with a hierarchical context layer that propagates compact QueryDocs by default while preserving full execution traces for on-demand recall, and with DAGRPO, a DAG-conditioned RL adaptation combining topology-conditioned credit with structural compliance regularization; it surpasses the strongest open-source baseline by 5.3/5.8/2.0 points on BrowseComp-Plus, GAIA and xbench-DeepSearch at the Qwen3-235B-A22B scale, improves over a same-budget outcome-only GRPO baseline by 3.

AI-generated editorial illustration: DAGent: Evaluate-then-Grow Planning for Deep Research Agents

Interpretation

It proposes Evaluate-then-Grow incremental planning: the task DAG starts empty and the Orchestrator expands it one batch of nodes at a time, conditioning each expansion on the structured QueryDoc emitted by completed nodes (answer, explanation, key steps, confidence, uncertainties); nodes are evaluated as Success, Uncertain or Not Found, the latter two spawn refine nodes with each line capped at three attempts (one initial execution plus two refinements), and termination is triggered by scheduling an answer-type node that must appear alone in its batch. Earlier DAG-based deep research systems (Flash-Searcher produces its complete plan in a single decomposition call and then periodically updates it; FlowSearch builds its knowledge flow over several planner iterations and then rewrites it through six edit operations) are Plan-then-Patch: a task-covering plan is committed before any node executes and later edits patch that standing plan. DAGent creates every node, including the answer node, only after the nodes it depends on have executed and been evaluated, and no node is ever deleted or rewired. A same-architecture ablation isolates the factor: replacing Evaluate-then-Grow with Plan-then-Patch costs 5.3/4.8/4.0 points on the three benchmarks, and collapsing further to a single upfront plan with no replanning costs another 8.7/7.8/6.0 points, totaling 14.0/12.6/10.0; the Plan-then-Patch variant reuses DAGent's Orchestrator, ReAct Executor, backbone, tool stack and context components, changing only cross-iteration planning behavior.

It proposes a hierarchical context layer: each node propagates only a compact QueryDoc to downstream nodes by default, bounding each node's context by its fan-in rather than graph depth or total node count (Selective Propagation), while preserving the full InteractionTranscript that a downstream Executor can retrieve through RecallTool for goal-conditioned evidence snippets. Long-horizon deep research produces InteractionTranscripts of arbitrary length while downstream nodes typically need only their dependencies' conclusions; the design separates what is propagated by default from what is recalled on demand, making long-horizon evidence usable without overloading each sub-task. Ablations show removing Selective Propagation costs 12.6/10.6/8.0 points, removing QueryDoc costs 8.0/7.7/5.0 points, and removing InteractionTranscript (which also disables recall) costs 4.0/2.9/2.0 points; in the Qwen3-32B runs RecallTool is invoked in only 22.0/13.6/11.0% of tasks, at a mean of 0.33/0.21/0.18 calls per task, confirming it is a fallback rather than a default channel.

It proposes DAGRPO: a GRPO-style objective augmented with two structural signals uniquely determined by the recorded topology, namely topology-conditioned credit on ReAct Executor rollouts (full credit inside the answer-inclusive closure of the answer-type node, attenuated outside it) and a structural compliance regularization on Orchestrator plans that penalizes failing to assign an explicit status to every node in the previous batch, failing to refine a weak node within the three-attempt budget, scheduling the answer node outside a singleton batch, duplicate sub-task prompts, and unparseable JSON or invalid dependency references. Outcome-only GRPO assigns the trajectory-level reward uniformly across all rollouts and ignores DAG-induced contribution differences; append-only growth makes a node's ancestry in the final graph coincide with the context it saw when it ran, a property that does not hold once a refiner can delete or rewire nodes after execution, as in FlowSearch and Flash-Searcher. At Qwen3-8B, DAGRPO raises DAGent's BrowseComp-Plus/GAIA/xbench-DeepSearch Pass@1 from 40.0/46.6/60.0 to 49.6/53.4/65.7 over three seeds, exceeding the same-budget outcome-only GRPO baseline by 3.6/2.6/2.7 points; removing topology-conditioned credit costs 2.3/1.9/1.7 points and removing the compliance regularization costs 1.6/1.0/0.7 points, with the off-chain credit coefficient peaking at 0.5 on every benchmark.

It reports broad empirical advantage and efficiency evidence: a 5.3/5.8/2.0-point lead over the strongest open-source baseline at the Qwen3-235B-A22B scale, 47.3/55.3/65.0 at Qwen3-32B against FlowSearch's 4.0/3.8/2.0-point gap, replication across four open-source backbones from four vendors, extension to GPT-5 at 327K context, and higher accuracy at lower per-task token, tool-call and step footprints than its Plan-then-Patch counterpart. The advantage is not concentrated in one setting: the paper reports that DAGent matches or exceeds every baseline on every difficulty split (with two ties at 235B-A22B) and that relative gains grow with difficulty at 8B and 32B; on efficiency, the Plan-then-Patch variant raises the off-chain Executor ratio from 0.25/0.20/0.18 to 0.40/0.30/0.25 and spends 40/29/18% more tokens, 28/21/21% more steps, 35/27/21% more external tool calls and 35/27/23% more wall-clock time while still trailing by 5.3/4.8/4.0 points. The main table covers three benchmarks, three Qwen3 scales and both training-free and training-based settings; the cross-backbone table covers Qwen3-32B, Seed-OSS-36B, GLM-4-32B and Nemotron-3-Nano-30B; the efficiency table reports per-task means in Qwen3-32B training-free mode and notes that 32KN is a per-sub-task cap rather than a per-task budget.

Perspective

The result targets deep research agents that need multi-step retrieval and evidence synthesis, in settings with a fixed retrieval backend and closed-form answers verifiable by an LLM judge (the 150-task BrowseComp-Plus evaluation split, GAIA's 103-task text-only validation subset and 165-task full set, and the 100 Chinese-language tasks of the xbench-DeepSearch 2505 release). It lets a planner decide which branches to expand only after evidence arrives, yielding higher accuracy at equal or lower token, tool-call, step and wall-clock cost; for training pipelines that want to use DAG topology directly as an RL credit signal, append-only growth supplies the dependable property that a node's final ancestry is the context it saw when it ran. The hierarchical context layer and RecallTool combine default propagation with on-demand recall, suited to multi-agent deployments with a bounded per-sub-task context (an 8K prompt plus a 32K response).

The paper states that DAGent is text-native, and the extension used for the full GAIA set converts attachments into text before planning begins; handling multi-modal evidence inside the loop would require changes to how the Orchestrator evaluates evidence and how the Executor represents its transcript. Its accuracy also depends on search-engine, retrieval-embedding and page-extraction quality, and it cannot create evidence the retriever never retrieves. The answer-inclusive closure is a structural proxy for contribution rather than a causal attribution, and the authors explicitly describe the association between structural compliance and accuracy as correlational rather than a causal estimate, leaving per-node causal attribution in DAG-based multi-agent systems an open question. All three benchmarks have closed-form answers an LLM judge can verify, so open-ended report generation lies outside the evidence presented. The RL experiments train one backbone (Qwen3-8B) with LoRA for 21 update steps over three seeds, and whether the DAGRPO gains persist at larger training scales, longer schedules or full fine-tuning is not established; moreover, some single-seed ablation rows differ from the three-seed means within seed noise, and the authors rest support for both structural signals on the consistent sign of every per-benchmark difference.

Sources