Contents
Figure 1: Think, act, observe, think again — the loop under most agents
Give a language model a hard question and two things can go wrong. If it only reasons, it can talk itself into confident nonsense because it has no way to check facts. If it only acts, it fires off tool calls with no stated plan, and when one fails it just tries something random. ReAct, short for "reasoning plus acting," fixes both by making the model alternate between the two: think, act, look at the result, think again. That simple loop sits underneath most tool-using agents today, even when nobody uses the name. For the full agent loop built from scratch around this pattern, see the from-scratch agent guide in this series.
The Idea and Where It Came From
ReAct comes from a paper by Shunyu Yao and colleagues, posted in October 2022 and presented at ICLR 2023. The setup is deliberately simple: a few-shot prompting framework in which a frozen LLM interleaves free-form reasoning traces (Thought) with task actions (Action) and environment feedback (Observation) in a loop. No training is involved in the core method, only prompting.
The paper's central claim is a division of labor: reasoning guides acting, by planning, tracking progress, and recovering from problems, while acting grounds reasoning, by retrieving facts that cut down on hallucination. The authors reported that interleaving beat chain-of-thought alone on question answering and fact verification, and lifted success rates on the interactive ALFWorld and WebShop benchmarks. They also argued that the visible trace is more interpretable, because you can read what the model believed before each action.
What a Trace Looks Like
Here's a trace for a support agent with two tools, using a made-up order system:
Question: Has order A100 shipped, and where is it now?
Thought 1: I need the order record first to see its status.
Action 1: lookup_order[A100]
Observation 1: {"status": "shipped", "tracking_id": "TRK-5531"}
Thought 2: It has shipped. The record has a tracking ID, so I can look up location.
Action 2: get_tracking[TRK-5531]
Observation 2: Error: tracking service timed out
Thought 3: The tracking lookup failed. I shouldn't invent a location. I'll retry once.
Action 3: get_tracking[TRK-5531]
Observation 3: {"location": "Regional hub, Denver", "eta": "2 days"}
Thought 4: I have everything I need.
Action 4: finish[Order A100 has shipped. It's at the Denver regional hub, arriving in about 2 days.]
The valuable moment is Thought 3. The tool failed, and the model's visible reasoning decided to retry instead of fabricating an answer. Without a thought step, an acting-only agent has nothing that distinguishes a principled retry from a random one. The recovery-from-bad-tool-calls ability is largely where the advantage comes from.
Implementing the Classic Text Version
Before native tool-calling APIs existed, ReAct was a prompt format plus a controller that parsed the model's output. Building it once teaches you what every agent loop does. The prompt gives the format and a worked example:
import re
PROMPT = """Answer the question by interleaving Thought, Action and Observation steps.
Available actions:
lookup_order[order_id] - returns status and total for an order
finish[answer] - give the final answer
Use exactly one Action per step. Never write an Observation yourself.
Example:
Question: What is order A100's total?
Thought 1: I need the order record.
Action 1: lookup_order[A100]
Observation 1: {{"status": "shipped", "total": 59.97}}
Thought 2: The record gives the total.
Action 2: finish[Order A100's total is $59.97.]
Question: {question}
"""
The controller runs the loop and enforces the rules the model can't be trusted to follow on its own:
ACTION_RE = re.compile(r"Action\s*\d*:\s*(\w+)\[(.*)\]\s*$", re.S)
def react(llm, tools, question, max_steps=8, obs_chars=1500):
scratch = PROMPT.format(question=question)
seen = set()
for step in range(1, max_steps + 1):
out = llm(scratch + f"Thought {step}:", stop=["\nObservation"])
scratch += f"Thought {step}:{out}\n"
m = ACTION_RE.search(out.strip())
if not m:
scratch += f"Observation {step}: Reply with 'Action {step}: name[input]'.\n"
continue
name, arg = m.group(1), m.group(2).strip()
if name == "finish":
return arg
if (name, arg) in seen:
obs = "You already ran this exact action. Try something different or finish."
elif name not in tools:
obs = f"Unknown action '{name}'. Available: {', '.join(tools)}."
else:
seen.add((name, arg))
try:
obs = str(tools[name](arg))[:obs_chars]
except Exception as e:
obs = f"Error: {e}"
scratch += f"Observation {step}: {obs}\n"
return "Stopped: step limit reached."
tools = {"lookup_order": lambda oid: {"A100": {"status": "shipped", "total": 59.97}}
.get(oid, "order not found")}
Each guard exists for a reason:
- The stop sequence (
"\nObservation") halts generation before the model can write its own fake observation, which models will happily do. The controller supplies real ones. - One action per step, parsed with a regex, keeps control flow deterministic.
- Feedback on parse errors and unknown actions lets the model recover instead of crashing.
- Duplicate detection breaks the repetitive loops that the original work noted as a weakness.
- Observation truncation protects the context window.
- A step limit caps cost.
ReAct in the Era of Native Tool Calling
Modern APIs absorbed this machinery. The mapping is direct: a typed tool_use block is the Action, a tool_result message is the Observation, and the loop is the one from the from-scratch agent guide. Native calling gives you validated arguments, parallel calls, and no regex parsing, which is why it's the default — the same machinery the function calling guide details.
What happened to the Thought? It moves around. Some models write brief text before a tool call, some reasoning models think internally, and some can reason between tool calls within a single task, though providers differ on how thinking must be preserved across tool-use turns, so check your provider's docs. A common misconception worth correcting: internal reasoning within one model turn is not the same as ReAct. A reasoning model's internal thinking doesn't inherently interleave with external tool calls and their results across multiple turns, whereas ReAct is specifically about structuring reasoning around actions and their observed outcomes.
Some 2026 write-ups go further and claim every frontier model now implements the ReAct loop automatically through tool use, so you rarely need to hand-format Thought, Action, and Observation. That's a fair description of the mechanics, but the design questions remain yours: which tools exist, how they describe themselves, what limits apply, and when a human must approve.
When Explicit Thoughts Still Help
The text format is still worth reaching for when:
- You need an auditable reasoning trail and want the model's stated rationale logged next to each action. (Treat it as a clue, not proof, since stated reasoning isn't guaranteed to be faithful to what actually drove the model.)
- You're using a model without native tool calling, such as some small or open-weight models.
- Debugging: reading the trace is often the fastest way to see why an agent went astray.
- Teaching or prototyping, where seeing the loop makes the concept click.
Related Patterns
ReAct isn't the only way to structure an agent, and the alternatives trade differently:
- Plan-and-execute writes a full plan first, then carries it out, reducing model calls but adapting more slowly when observations contradict the plan.
- ReWOO plans with placeholders for tool results and fills them in later, cutting the number of reasoning calls, at the cost of flexibility.
- Reflexion adds self-critique across attempts, storing lessons for the next try.
- Tree of Thoughts explores multiple reasoning branches, mostly for problems where search over alternatives matters more than tool access.
A reasonable default is ReAct-style adaptivity for open-ended tasks and a fixed workflow or plan-first approach when the steps are known in advance, which is also cheaper and more predictable.
Where ReAct Struggles
- Cost. Every step re-sends the growing scratchpad, so token usage climbs with each iteration.
- Latency. Steps are sequential by nature, which is why parallel tool calls and plan-first approaches exist.
- Error compounding. A wrong early observation or assumption can steer everything after it.
- Loops. Models can repeat the same action expecting a different result. Dedupe and step limits help.
- Brittle parsing in the text version, solved largely by native tool calling.
- Hallucinated observations if you forget stop sequences.
- Unfaithful reasoning. The Thought can sound sensible while the real decision came from somewhere else.
- Injection through observations. Whatever a tool returns, whether a web page, email, or document, can contain instructions aimed at your agent. Treat observations as data, never as commands, and apply the approval and sandboxing practices from the function calling security model.
Prompting Tips That Help
- Keep thoughts short and purposeful: what's known, what's missing, what to do next.
- Allow one action per step in text-format ReAct.
- Include a worked example that shows a failure and a recovery, not only a smooth success.
- Provide an explicit finish action so the model knows how to stop.
- Tell the model never to invent observations and to say so when it lacks information.
- Describe each tool's inputs, outputs, and limits precisely.
Evaluating a ReAct Agent
Read traces, not just final answers. Track task success rate across a test set, number of steps, cost, and latency, and run each case several times because outputs vary. Categorize failures: wrong tool, wrong arguments, ignored observation, loop, premature finish, and so on. A check that an agent reached the right answer through a dangerous or wasteful path is part of the evaluation, not an optional extra.
Common Pitfalls
- Skipping the stop sequence and letting the model fabricate observations.
- No step limit or duplicate-action guard.
- Returning huge raw tool outputs that crowd out the reasoning.
- Trusting the visible thought as a faithful explanation of the model's process.
- Treating observations as trusted instructions.
- Using an adaptive ReAct loop for a task with fixed steps, paying extra cost and variance for flexibility you don't need.
Recommended Books
| Cover | Book | Description | Get it |
|---|---|---|---|
![]() |
AI Agents in Action | - agent builds where the reasoning-acting loop is the backbone of every chapter. | View on Amazon |
![]() |
Designing Agentic AI Systems | - architecture patterns for choosing ReAct vs plan-first per workload. | View on Amazon |
![]() |
Build a Large Language Model (From Scratch) | - the model side: how prompting shapes what comes out of a frozen LLM. | View on Amazon |
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
What is ReAct in one paragraph?
ReAct (reasoning plus acting) is a prompting framework where a frozen LLM interleaves free-form reasoning traces (Thought) with task actions (Action) and environment feedback (Observation) in a loop. No training is involved in the core method, only prompting. Reasoning guides acting — planning, tracking progress, recovering from problems — while acting grounds reasoning by retrieving facts that cut down on hallucination. The loop sits underneath most tool-using agents today, even when nobody uses the name.
Does ReAct still matter now that models have native tool calling?
The mechanics moved into the APIs — a typed tool_use block is the Action, a tool_result message is the Observation — but the discipline is unchanged. Internal reasoning within one model turn is not the same as ReAct: a reasoning model's internal thinking doesn't inherently interleave with external tool calls and their results across multiple turns, whereas ReAct is specifically about structuring reasoning around actions and their observed outcomes. Native calling gives you validated arguments, parallel calls, and no regex parsing, which is why it's the default.
When should I still use the explicit Thought-Action-Observation text format?
When you need an auditable reasoning trail logged next to each action (treat it as a clue, not proof); when you're using a model without native tool calling, such as some small or open-weight models; for debugging, where reading the trace is often the fastest way to see why an agent went astray; and for teaching or prototyping, where seeing the loop makes the concept click.
Where does ReAct struggle?
Cost (every step re-sends the growing scratchpad), latency (steps are sequential by nature), error compounding (a wrong early observation can steer everything after it), loops (models can repeat the same action expecting a different result), brittle regex parsing in the text version, hallucinated observations if you forget stop sequences, unfaithful reasoning (the Thought can sound sensible while the real decision came from elsewhere), and injection through observations (tool returns can contain instructions aimed at your agent — treat observations as data, never as commands).
Wrapping This Up
ReAct's insight is that reasoning and acting do their best work together: thoughts decide what to do with each observation, and observations keep the thoughts tied to reality. The original text-based format is now mostly absorbed into native tool-calling APIs, but the loop, the guards, and the design discipline remain the core of agent building.
Do you still need to write Thought, Action, Observation prompts by hand? Usually not. But build the classic loop once, watch a trace where the model recovers from a failed tool call, and you'll understand what your framework is doing and where its failures come from.


