Skip to main content
Back to timeline
arXivSource publication:

DeReAct externalizes action validation and completion certification, lifting Pass@1 by 6.5–7.0 points for Qwen3-Coder-480B on GAIA and SWE-bench Verified

Related research and updates

Synopsis

DeReAct introduces a modular agent architecture that separates action authorization and completion control from a single LLM policy, delegating them to a Critic that validates proposed actions before execution and a Context Manager that reconstructs an environment-supported State and certifies task completion; across GAIA and SWE-bench Verified it improves Pass@1 most for weaker Brain models (6.5–7.0 points for Qwen3-Coder-480B and 4.2–5.2 points for Claude Sonnet 4.5), with gains diminishing as Brain capability increases, while on Claude Opus 4.5 Pass@1 remains comparable to ReAct but trajectories are more evidence-complete and constraint-satisfying.

Source-provided article image: DeReAct: Decomposed Reasoning and Acting for Reliable AI Agents
Figure 2 ·

Figure 2: Pass@3 lift versus Pass@1 lift over ReAct for all configurations across both benchmarks, Brain tiers, CM models, window sizes, and prompts. The vertical distance from the y = x y{=}x diagonal is Diff. Nearly every point lies above the diagonal. In the shaded region, DeReAct trails ReAct per attempt yet solves more distinct tasks at least once.

arXiv

Interpretation

DeReAct splits the action authorization and completion control that ReAct-style agents couple inside a single LLM policy into two externalized gating policies: a Critic that validates proposed actions before execution, and a Context Manager that reconstructs an environment-supported State and certifies task completion. Relative to a ReAct baseline where one policy proposes actions, interacts with the environment, and decides when a task is complete, this work externalizes authorization and completion judgment so each can be enforced independently. The abstract states this mechanism through architecture description and design motivation, framing it against error propagation and unsupported completion claims terminating execution.

On GAIA and SWE-bench Verified, DeReAct improves Pass@1 most for weaker Brain models: 6.5–7.0 points for Qwen3-Coder-480B and 4.2–5.2 points for Claude Sonnet 4.5, with gains diminishing as Brain capability increases. The result states an empirical regularity about how external gating benefits vary with base-model capability, rather than an isolated gain on one model. The abstract reports Pass@1 gain ranges on two benchmarks and explicitly notes that gains shrink as Brain capability rises.

Trajectory and ablation analyses show that external gating is effective when targeted failures are sufficiently prevalent and the gating policy is itself sufficient. The analysis makes the conditions for external gating effectiveness explicit, tying benefit to the failure distribution and to gating-policy quality. The abstract cites trajectory and ablation analyses as the basis, without reporting specific ablation values.

With Claude Opus 4.5, DeReAct's Pass@1 remains comparable to ReAct while producing more evidence-complete and constraint-satisfying trajectories, indicating completion control can trade earlier termination for stronger grounding. This observation extends evaluation beyond Pass@1 to trajectory evidence completeness and constraint satisfaction, suggesting a trade-off between termination timing and grounding. The abstract reports comparable Pass@1 and uses trajectory attribute differences as evidence of grounding benefit.

Perspective

The work targets LLM agent systems that need reliable action authorization and completion certification, in the task settings represented by GAIA and SWE-bench Verified, with Brain model capability as a moderating condition: gains are largest for weaker Brain models and diminish as capability increases. Its design intent is to make action validation and completion judgment independently enforceable, reducing error propagation and unsupported completion claims terminating execution; for settings using strong models such as Claude Opus 4.5, the benefit appears as trajectory evidence completeness and constraint satisfaction rather than Pass@1.

At the abstract level, sample sizes, variance, statistical tests, and specific ablation values are not given, nor are the implementation details of the Critic and Context Manager, how the gating policies are constructed, or per-benchmark breakdowns on GAIA and SWE-bench Verified; therefore which failure distributions maximize external gating benefit, how sufficient a gating policy must be, and how trajectory evidence completeness and constraint satisfaction are measured remain open questions that require the full text.

Sources