DeReAct: Decomposed Reasoning and Acting for Reliable AI Agents
Source
arxiv.org
Author
Ajay Vohra, Tao Chen, Neeti Narayan, Caron Zhang
Date
Why it matters
Separating action validation and completion certification from the acting model gives measurable gains for weaker models (up to 7 points) but little for frontier ones. This tells you when an external gating layer is worth building.
Key takeaways · AI-distilled
DeReAct splits the usual single ReAct policy into gated roles: a Critic validates each proposed action before it runs, and a context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → Manager rebuilds an environment-supported state and certifies when the task is actually complete.
The motivation: when one LLMA large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.Full definition → proposes actions and also decides it is done, unsupported completion claims can end a run early and errors propagate, because authorization and completion cannot be enforced independently.
On GAIA and SWE-benchThe standard benchmark for AI coding agents: real GitHub issues from real repositories, scored by whether the agent's patch passes the project's own tests.Full definition →, gains were largest for weaker base models, about 6.5 to 7.0 Pass@1 points for Qwen3-Coder-480B and 4.2 to 5.2 for Claude Sonnet 4.5, and shrank as the base model got stronger.
With Claude Opus 4.5, Pass@1 stayed roughly level with ReAct, but the authors report more evidence-complete, constraint-satisfying trajectories, trading earlier termination for stronger groundingTying a model's answers to checkable sources — retrieved documents, live data, tool results — instead of letting it answer from memory alone.Full definition →.
Terms in this piece · Glossary
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
SWE-bench — The standard benchmark for AI coding agents: real GitHub issues from real repositories, scored by whether the agent's patch passes the project's own tests.
LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
grounding — Tying a model's answers to checkable sources — retrieved documents, live data, tool results — instead of letting it answer from memory alone.