Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

DeReAct externalizes action validation and completion certification, lifting Pass@1 by 6.5–7.0 points for Qwen3-Coder-480B on GAIA and SWE-bench Verified

DeReAct introduces a modular agent architecture that separates action authorization and completion control from a single LLM policy, delegating them to a Critic that validates proposed actions before execution and a Context Manager that reconstructs an environment-supported State and certifies task completion; across GAIA and SWE-bench Verified it improves Pass@1 most for weaker Brain models (6.5–7.0 points for Qwen3-Coder-480B and 4.2–5.2 points for Claude Sonnet 4.5), with gains diminishing as Brain capability increases, while on Claude Opus 4.5 Pass@1 remains comparable to ReAct but trajectories are more evidence-complete and constraint-satisfying.