Contents
Figure 1: Decide the steps once, then run them — with a gate for revision
A ReAct agent decides one step at a time, which is great for exploration and bad for long, structured jobs. On a 15-step task it can lose the thread, repeat itself, or burn money re-reading a growing scratchpad. Planning agents take a different approach: break the goal into steps first, then carry the steps out, revising the plan only when something goes wrong. This guide covers how to decompose tasks well, the main plan-and-execute patterns, a working implementation with parallel execution and replanning, and the failure modes to watch for. For the step-by-step alternative this pattern trades against, see the ReAct guide in this series.
Why Plan at All?
One engineering comparison states ReAct's limitation bluntly: the agent never sees the big picture, since it decides one step at a time. Planning addresses that, and brings several other benefits:
- Cost. A 10-step task that needs 10 model calls in ReAct might need only one or two with plan-and-execute, since a cheaper model can run each step.
- Parallelism. If the plan shows which steps are independent, they can run at once.
- Reviewability. A plan is something a human can read and approve before anything happens, which is far better than auditing actions after the fact.
- Control and safety. Fixing the plan before the agent reads untrusted content limits how much injected text can redirect its behavior, though it doesn't stop poisoned data from corrupting arguments within a step.
- Focus. Each executor step gets a small, scoped context instead of the whole history.
The cost is rigidity. A plan made without seeing the world can be wrong, which is why replanning matters.
What Good Decomposition Looks Like
Decomposition is the heart of planning, and a plan is only as good as its steps. Good steps are:
- Atomic: one clear action or tool call, not "research the market."
- Verifiable: each has a checkable outcome, so you know whether it worked.
- Explicit about dependencies: which steps need which results.
- Tool-bound: each maps to a tool or capability the agent actually has.
- Scoped: each needs only the context relevant to it.
The right granularity is a tradeoff. Too coarse and steps become mini-agents that fail unpredictably, while too fine and you pay overhead and lock in brittle detail. Surveys frame two styles: proactive decomposition, which generates a complete plan before any action, and progressive decomposition, which interleaves planning with execution and refines the plan based on observations. Most practical systems blend them.
The Main Patterns
Plan-and-Execute. A capable model writes the full plan, then an executor, often a cheaper model, works through the steps, and a replanner revises when needed. The planner typically generates a DAG of subtasks with dependencies. A cost caution applies: if replanning fires on most tasks, you're paying the planning cost and the adaptation cost.
ReWOO (Reasoning Without Observation). Plans once with placeholders for tool results, runs the tools, then synthesizes at the end, using two LLM calls in total. It's the cheapest pattern, but the plan is locked in during the first call, so it can't plan for tools you can't anticipate upfront.
LLMCompiler. Borrows from classical compilers. It has a planner that builds a DAG showing which tools to call and in what order, a task fetching unit that schedules ready tasks, and an executor that runs tools in parallel where possible. It also supports replanning, which distinguishes it from ReWOO.
Planner/Executor separation. Plan-and-Act, for example, trains a dedicated planner to generate structured high-level plans and an executor to translate them into actions, reporting strong results on web navigation benchmarks. Other work reports that structured planning scripts improved multi-step tool-calling accuracy in enterprise settings from 41% to 96%, and that confining replanning to the active sub-task cut token use by up to 82%. These are individual research claims, so treat them as indicators of direction, not guarantees for your system.
Hierarchical orchestration. A manager model plans and delegates to workers. One computer-use design has the manager dispatch parallel subagents to nodes on the ready frontier of the DAG, and continuously revise the DAG — adding, canceling, or rewriting nodes as findings arrive. This is the orchestrator-and-workers pattern common in research agents.
Search-based planning. Tree-of-Thoughts and tree-search variants explore alternative plans and backtrack, which suits problems where evaluating options matters more than tool access.
A common hybrid: plan first, then run each step as a small ReAct agent. You get the global structure of a plan and the local adaptability of a loop — the best of the ReAct pattern applied per-step.
A Working Implementation
This version plans a DAG, validates it, executes independent steps in parallel, and replans on failure. It uses two placeholder functions you supply: llm_json(prompt), which returns parsed JSON, and llm_text(prompt), which returns text. Step results can be referenced in later arguments as "$s1".
import json
from concurrent.futures import ThreadPoolExecutor, as_completed
MAX_STEPS, MAX_REPLANS = 12, 2
PLANNER_PROMPT = """You are a planner. Break the task into steps.
Tools: {tool_docs}
Return ONLY JSON:
{{"steps": [{{"id": "s1", "goal": "...", "tool": "<tool name>",
"args": {{...}}, "depends_on": []}}]}}
Rules: at most {max_steps} steps. Reference an earlier step's output in args as "$s1".
List prerequisites in depends_on. Independent steps must not depend on each other.
Do not reuse these ids or repeat completed work: {done}
Task: {task}"""
def validate_plan(plan, tools, done=frozenset()):
steps = plan.get("steps", [])
if not steps or len(steps) > MAX_STEPS:
raise ValueError("plan size out of bounds")
ids = [s["id"] for s in steps]
if len(set(ids)) != len(ids) or set(ids) & set(done):
raise ValueError("duplicate or reused step ids")
for s in steps:
if s["tool"] not in tools:
raise ValueError(f"unknown tool {s['tool']}")
if not set(s.get("depends_on", [])) <= (set(ids) | set(done)):
raise ValueError("dependency on unknown step")
indeg = {s["id"]: len([d for d in s.get("depends_on", []) if d not in done])
for s in steps}
ready, seen = [i for i, n in indeg.items() if n == 0], 0
while ready: # cycle check (Kahn's algorithm)
node = ready.pop()
seen += 1
for s in steps:
if node in s.get("depends_on", []):
indeg[s["id"]] -= 1
if indeg[s["id"]] == 0:
ready.append(s["id"])
if seen != len(steps):
raise ValueError("plan has a cycle")
def call_tool(step, tools, results):
args = {k: results[v[1:]] if isinstance(v, str) and v.startswith("$") and v[1:] in results
else v for k, v in step.get("args", {}).items()}
return str(tools[step["tool"]](**args))[:2000]
def run_plan(plan, tools, results):
steps = {s["id"]: s for s in plan["steps"]}
pending, failed = set(steps), None
with ThreadPoolExecutor(max_workers=4) as pool:
while pending and not failed:
ready = [i for i in pending
if set(steps[i].get("depends_on", [])) <= set(results)]
if not ready:
break
futures = {pool.submit(call_tool, steps[i], tools, results): i for i in ready}
for f in as_completed(futures):
i = futures[f]
try:
results[i] = f.result()
except Exception as e:
failed = (i, str(e))
pending.discard(i)
return failed
def plan_and_execute(task, tools, llm_json, llm_text, approve=None):
tool_docs = "; ".join(f"{n}: {(f.__doc__ or '').strip()}" for n, f in tools.items())
results, done = {}, "none"
for _ in range(MAX_REPLANS + 1):
plan = llm_json(PLANNER_PROMPT.format(
tool_docs=tool_docs, max_steps=MAX_STEPS, done=done, task=task))
validate_plan(plan, tools, done=set(results))
if approve and not approve(plan): # optional human review of the plan
return "Plan rejected by reviewer."
failed = run_plan(plan, tools, results)
if not failed:
summary = json.dumps({k: v[:500] for k, v in results.items()})
return llm_text(f"Task: {task}\nStep results: {summary}\n"
"Answer using only these results.")
done = json.dumps({"completed": {k: v[:200] for k, v in results.items()},
"failed_step": failed})
return "Stopped: replan limit reached."
Several design choices here deserve a look:
- Validation before execution. The model is an untrusted planner — the same principle as the function calling security model. The code checks plan size, duplicate IDs, unknown tools, missing dependencies, and cycles before running anything, which catches hallucinated tools and malformed graphs.
- Dependency-driven scheduling. Steps run as soon as their prerequisites finish, and independent steps run concurrently, the same idea behind build tools like make.
- An approval hook. Passing
approvelets a human review the plan before any action, which is one of planning's biggest safety advantages. - Replanning with context. On failure, the replanner receives completed results and the failed step, and is told not to repeat work or reuse IDs. It only needs to produce the remaining steps.
- Limits everywhere: step cap, replan cap, output truncation, and bounded workers.
- Grounded synthesis. The final answer is told to use only recorded results.
Replanning Policy
Replanning is where plan-and-execute either earns its keep or quietly doubles your costs. Decide deliberately:
- Replan when a step fails after limited retries, an observation contradicts a plan assumption, required information is missing, or the user changes the goal.
- Retry first for transient errors like timeouts, since replanning for a flaky API wastes a planner call.
- Replan locally where possible, revising only the affected branch instead of regenerating everything.
- Cap replans and fail loudly when the cap is hit.
- Track the replan rate. If most runs replan, your up-front plans aren't good enough, or a ReAct-style loop fits the task better.
Choosing a Pattern
- Open-ended exploration where each step depends on the last: ReAct.
- Long, structured tasks where cost matters: plan-and-execute with a strong planner and cheaper executor.
- Predictable tool chains, minimum cost: ReWOO, accepting that it can't adapt.
- Many independent subtasks and latency pressure: a DAG planner with parallel execution, LLMCompiler-style.
- Large research or computer-use jobs: hierarchical manager and subagents, with careful state passing.
- Anything with irreversible actions: a plan-then-approve gate, regardless of pattern.
Where Planning Fails
- Plausible but wrong plans. Models produce confident plans with missing steps, wrong order, or impossible assumptions. Verify plans with code: dependency checks, tool availability, constraint validators, or simulation where you can.
- Stale plans. The world changes mid-execution, and a rigid plan keeps marching.
- Error propagation. An early wrong result silently feeds every dependent step. Add verification steps for critical intermediate results.
- Over-decomposition. Dozens of tiny steps add overhead and failure points, and surveys note error cascades in fixed decomposition schemes as a core weakness.
- Context loss between steps. If executors only see their own slice, important findings can fail to carry forward. Pass forward exactly what downstream steps need.
- Injection and data poisoning. A fixed plan limits control-flow hijacking, but retrieved content can still corrupt arguments or the final synthesis. Treat tool outputs as data.
- Hidden costs. Planner calls, replanner calls, and synthesis add up, so measure total cost against a simple ReAct baseline before committing.
Evaluating a Planning Agent
Evaluate the plan and the execution separately. For plans: completeness, correct ordering, dependency accuracy, valid tool use, and whether independent steps were actually parallelized. For execution: task success, step failure rate, replan rate, tokens, latency, and parallel speedup. Test the planner with fixed tasks and a scripted fake model, test validators with deliberately broken plans (cycles, unknown tools, oversized plans), and compare end-to-end results against a ReAct baseline on the same task set.
Practical Tips
- Make plans machine-readable (JSON), and validate them.
- Give each step a goal and a success condition, not just an action.
- Use a stronger model to plan and a cheaper one to execute where tasks are simple.
- Show plans to users, or reviewers, before high-impact steps.
- Pass compact summaries between steps, not raw dumps.
- Log the plan, every replan, and every step result.
Common Pitfalls
- Trusting the plan without validating tools, IDs, and dependencies.
- Skipping replanning limits, then watching a loop of plan, fail, replan run up costs.
- Decomposing into steps too vague to verify.
- Using plan-and-execute for tasks where the next step truly depends on discoveries, when ReAct would adapt better.
- Assuming parallel execution is safe without checking that steps don't conflict.
- Not comparing against a simple baseline.
Recommended Books
| Cover | Book | Description | Get it |
|---|---|---|---|
![]() |
Designing Agentic AI Systems | - architecture patterns for planners, executors, and orchestration boundaries. | View on Amazon |
![]() |
AI Agents in Action | - hands-on agent builds where plan structure shows up in real tasks. | View on Amazon |
![]() |
AI Engineering | - cost, latency, and evaluation tradeoffs behind multi-step systems. | View on Amazon |
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
Why plan at all if ReAct agents already work?
A ReAct agent never sees the big picture — it decides one step at a time, which on a long task can lose the thread, repeat itself, or burn money re-reading a growing scratchpad. Planning brings cost savings (a 10-step task that needs 10 model calls in ReAct might need only one or two with plan-and-execute), parallelism (independent steps run at once), reviewability (a human can read and approve the plan before anything happens), control (fixing the plan before the agent reads untrusted content limits injected-text redirection), and focus (each executor step gets a small, scoped context).
What are the main plan-and-execute patterns?
Plan-and-Execute: a capable model writes the full plan, a cheaper executor works the steps, a replanner revises when needed. ReWOO plans once with placeholders for tool results, runs tools, synthesizes at the end — cheapest, but locked in. LLMCompiler builds a DAG of tool calls and schedules ready tasks in parallel with replanning support. Planner/Executor separation trains dedicated models for each role. Hierarchical orchestration has a manager delegate to parallel subagents while revising the DAG. Search-based planning (Tree-of-Thoughts) explores alternatives and backtracks.
When should replanning fire?
Replan when a step fails after limited retries, an observation contradicts a plan assumption, required information is missing, or the user changes the goal. Retry first for transient errors like timeouts, since replanning for a flaky API wastes a planner call. Replan locally where possible — revise only the affected branch instead of regenerating everything. Cap replans and fail loudly when the cap is hit. Track the replan rate: if most runs replan, your up-front plans aren't good enough, or a ReAct-style loop fits the task better.
How do you evaluate a planning agent?
Evaluate the plan and the execution separately. For plans: completeness, correct ordering, dependency accuracy, valid tool use, and whether independent steps were actually parallelized. For execution: task success, step failure rate, replan rate, tokens, latency, and parallel speedup. Test the planner with fixed tasks and a scripted fake model, test validators with deliberately broken plans (cycles, unknown tools, oversized plans), and compare end-to-end results against a ReAct baseline on the same task set.
Wrapping This Up
Planning agents trade ReAct's step-by-step adaptivity for global structure, lower cost, parallelism, and reviewable plans. Plan-and-execute separates a planner from executors, ReWOO plans once for minimum cost, LLMCompiler turns plans into parallel DAGs, and hierarchical orchestrators delegate to subagents. All of them depend on good decomposition, validated plans, bounded replanning, and verification.
Which pattern should you start with? Begin with a simple loop, add a plan when tasks get long or costly, and add parallelism and replanning only when measurements show they pay off. Then keep the human gate on any plan that does something you can't undo.

