Sam Austin on October 10, 2026

LangGraph Tutorial: Build Stateful Multi-Step Agents

LangGraph Tutorial: Build Stateful Multi-Step Agents
Contents

Figure 1: State, nodes, edges, and a checkpointer — the graph is the agent

A plain agent loop works until you need things it can't give you cheaply: pausing for human approval and resuming tomorrow, surviving a crash mid-run, rewinding to an earlier step, or routing work down different paths based on state. LangGraph exists for those problems. It models an agent as a graph of steps that read and write shared state, with a checkpointer saving that state after every step. This tutorial builds a support agent with tool calling, conditional routing, persistent memory, and a human approval gate for refunds, then covers testing and production concerns.

If you haven't built a raw agent loop first, the from-scratch guide earlier in this series is worth reading, because LangGraph is that same loop made explicit.

Core Concepts in Five Minutes

  • State is a shared data structure (usually a TypedDict) that every step reads and updates.
  • Nodes are plain Python functions that take the state and return updates to it.
  • Edges define what runs next, either always or conditionally based on state.
  • Reducers control how updates merge into state. The common one, add_messages, appends new messages instead of overwriting the list.
  • Checkpointers save state after each step, keyed by a thread_id, which is what enables memory, pausing, and recovery.

Nodes return partial updates, and LangGraph merges them into state using each field's reducer. That's the main mental shift from a plain loop.

Setup

pip install -U langgraph langchain langchain-anthropic
export ANTHROPIC_API_KEY="your-key-here"

LangGraph's API has evolved quickly. Older tutorials import MemorySaver, while current docs use InMemorySaver, and the newest docs show updated streaming methods. Pin your versions and check the current documentation if an import fails.

Step 1: State and Tools

We extend the prebuilt MessagesState with one extra field and define two tools, one safe and one irreversible:

from typing import Literal
from langchain_core.messages import AIMessage, SystemMessage, ToolMessage
from langchain_core.tools import tool
from langgraph.graph import END, START, MessagesState, StateGraph

class State(MessagesState):
    pass  # messages with the add_messages reducer; add fields as needed

ORDERS = {
    "A100": {"status": "shipped", "total": 59.97},
    "A101": {"status": "processing", "total": 19.99},
}

@tool
def lookup_order(order_id: str) -> dict:
    """Look up an order's status and total by order ID."""
    return ORDERS.get(order_id, {"error": "order not found"})

@tool
def issue_refund(order_id: str, amount: float) -> str:
    """Issue a refund for an order. Irreversible. Requires human approval."""
    return f"Refunded ${amount:.2f} on order {order_id}"

ALL_TOOLS = [lookup_order, issue_refund]

Tool docstrings are what the model reads to decide when to call them, so write them like instructions to a new colleague, and say plainly which tool is risky.

Step 2: The Model Node

from langchain.chat_models import init_chat_model

model = init_chat_model("anthropic:claude-sonnet-5-5")   # check current model names
llm = model.bind_tools(ALL_TOOLS)

SYSTEM = SystemMessage(content=(
    "You are a support agent. Use tools for facts; never guess order data. "
    "Only request a refund when the customer asks for one."
))

def agent(state: State):
    return {"messages": [llm.invoke([SYSTEM, *state["messages"]])]}

The node returns a one-item message list, and the add_messages reducer appends it to history. It doesn't mutate state directly.

Step 3: Routing and the Approval Gate

Two functions do the interesting work. The first decides what happens after the model responds, and the second pauses for a human when a refund is requested:

from langgraph.prebuilt import ToolNode
from langgraph.types import Command, interrupt

def route(state: State):
    last = state["messages"][-1]
    if not getattr(last, "tool_calls", None):
        return END                                   # no tools requested: done
    if any(tc["name"] == "issue_refund" for tc in last.tool_calls):
        return "approval"                            # risky: ask a human first
    return "tools"

def approval(state: State) -> Command[Literal["tools", "agent"]]:
    last = state["messages"][-1]
    risky = [tc for tc in last.tool_calls if tc["name"] == "issue_refund"]

    decision = interrupt({"action": "approve_refund", "calls": risky})  # pauses here

    if decision.get("approve"):
        return Command(goto="tools")
    denials = [
        ToolMessage(content="The user denied this action.", tool_call_id=tc["id"])
        for tc in last.tool_calls
    ]
    return Command(goto="agent", update={"messages": denials})

interrupt() is the key primitive. The value you pass to Command(resume=...) becomes the return value of the interrupt() call inside the paused node. So decision above is whatever the human sends back. If they deny, we answer every pending tool call with a denial message (the API requires a result for each call) and send control back to the model, which can explain or offer alternatives.

Step 4: Assemble the Graph

from langgraph.checkpoint.memory import InMemorySaver

builder = StateGraph(State)
builder.add_node("agent", agent)
builder.add_node("tools", ToolNode(ALL_TOOLS))
builder.add_node("approval", approval)

builder.add_edge(START, "agent")
builder.add_conditional_edges("agent", route, ["tools", "approval", END])
builder.add_edge("tools", "agent")          # after tools run, the model sees the results

graph = builder.compile(checkpointer=InMemorySaver())

Read the wiring as a flow: start at the model; if it wants a safe tool, run it and loop back; if it wants a refund, pause for approval; otherwise finish. Every agent loop you've seen is a graph like this one. ToolNode executes whichever tool calls are on the last message and returns the results as tool messages.

Step 5: Run, Pause, Resume

The thread_id identifies a conversation, and the checkpointer stores its state:

from langgraph.types import Command

config = {"configurable": {"thread_id": "ticket-42"}}

graph.invoke(
    {"messages": [("user", "Order A100 arrived damaged. Please refund $20.")]},
    config,
)

snapshot = graph.get_state(config)
print(snapshot.next)    # ('approval',) means the graph is paused at the gate

At this point nothing has been refunded. The graph is suspended with its full state saved, and it could stay that way for hours or days. A reviewer inspects the pending request (the interrupt payload is available on the snapshot's tasks, with details varying by version) and decides:

result = graph.invoke(Command(resume={"approve": True}), config)
print(result["messages"][-1].content)

Resuming with {"approve": False} instead sends the denial path, and the model responds accordingly. Docs also describe a static alternative, compiling with interrupt_before=["tools"] to pause before a node runs, but interrupt() inside a node is more flexible because the pause can be conditional and carry a custom payload.

Memory Comes for Free

Because state is checkpointed per thread, a follow-up on the same thread_id continues the conversation with full history:

graph.invoke({"messages": [("user", "What was the status of that order again?")]}, config)

A different thread_id starts fresh. That's short-term, per-conversation memory. For facts that should persist across conversations, such as user preferences, you'd add a separate long-term store and expose it through tools or a retrieval step.

Time Travel and Debugging

Every step is a checkpoint, so you can inspect or rewind a run:

for snap in graph.get_state_history(config):
    print(snap.next, len(snap.values["messages"]))

Pick an earlier snapshot, and you can resume from it (optionally after editing state with graph.update_state), which creates a fork. This makes debugging far less guessy: instead of re-running a whole task and hoping it fails the same way, you replay from the step just before the bad decision.

Test Without a Model

Nodes and routers are ordinary functions, so the logic that matters most can be tested deterministically:

def test_refund_requires_approval():
    msg = AIMessage(content="", tool_calls=[
        {"name": "issue_refund", "args": {"order_id": "A100", "amount": 20}, "id": "1"}
    ])
    assert route({"messages": [msg]}) == "approval"

def test_lookup_runs_without_approval():
    msg = AIMessage(content="", tool_calls=[
        {"name": "lookup_order", "args": {"order_id": "A100"}, "id": "1"}
    ])
    assert route({"messages": [msg]}) == "tools"

def test_plain_answer_ends():
    assert route({"messages": [AIMessage(content="Done")]}) == END

For end-to-end tests, substitute a scripted fake model for the real one (the same trick as in the from-scratch guide), and keep a separate evaluation suite with the real model for quality measurement.

Production Notes

  • Use a durable checkpointer. InMemorySaver disappears when the process restarts, which defeats the point of pausing for a day. Production setups use a database-backed checkpointer such as the SQLite or Postgres packages. Durable execution is LangGraph's headline feature: agents persist through failures and can run for extended periods, resuming from where they left off.
  • Beware side effects before interrupt(). When you resume, the paused node re-runs from its start, and interrupt() returns the human's value instead of pausing. Any non-idempotent work done earlier in that node, such as sending an email or writing a record, will happen twice. Keep nodes with interrupts clean: decide first, act after, or put side effects in a separate node.
  • Make tools idempotent. Retries and resumes mean a tool may occasionally run more than once. Use idempotency keys for payments and writes.
  • Set a recursion limit. Graphs with loops can run forever if the model keeps calling tools. Pass "recursion_limit" in the config to cap steps, and handle the error it raises:
config = {"configurable": {"thread_id": "ticket-42"}, "recursion_limit": 12}
  • Stream for responsiveness. Use the streaming methods so users see progress, and note that the streaming API has changed between versions, so follow the current docs.
  • Trace and observe. LangSmith is the integrated option (commercial), and OpenTelemetry-based tracing is a vendor-neutral alternative. Either way, you want a trace of every node, tool call, and state change, since "why did the agent do that?" is the question you'll ask most.
  • Version churn is real. LangChain 1.0 introduced create_agent as the high-level way to build a standard tool-calling agent, which wraps much of what we built by hand. If you only need a standard loop, use it. Build the graph yourself when you need custom routing, approvals, or multi-step state, as here.

When LangGraph Is Worth It

It earns its structure when you need pausing and resuming, durable state, human approval, branching workflows, or replayable runs. The graph model costs more upfront structure but pays off when you need branching, checkpointed state, human-in-the-loop interrupts, and replayable execution. For a simple tool-using chatbot that never pauses, a plain loop or the higher-level helper is less code — and the framework comparison in this series maps where LangGraph sits against the alternatives.

Common Pitfalls

  • Mutating state inside nodes instead of returning updates, which bypasses reducers.
  • Forgetting a thread_id in the config, which means no checkpointing and an error when you try to interrupt.
  • Putting side effects before interrupt() in the same node.
  • Using InMemorySaver in production and losing paused runs on deploy.
  • Skipping the denial path, so a rejected action leaves tool calls without matching results.
  • Following an old tutorial's imports without checking your installed version.
  • No recursion limit, so a confused loop becomes a cost problem.
CoverBookDescriptionGet it
LangGraph: Agentic Applications the framework this tutorial uses, covered in depth. View on Amazon
Cover of “AI Agents in Action” AI Agents in Action hands-on coverage of the loops and tool-calling patterns LangGraph makes explicit. View on Amazon
Cover of “Designing Agentic AI Systems” Designing Agentic AI Systems architecture and design decisions for approval gates, state, and multi-step agents. View on Amazon

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

What does LangGraph give you that a plain agent loop doesn't?

Cheap versions of four things: pausing for human approval and resuming later, surviving a crash mid-run, rewinding to an earlier step, and routing work down different paths based on state. A raw loop can do all of these, but you build the plumbing yourself — serialization, resume logic, branch tracking. LangGraph models the agent as a graph of steps over shared state with a checkpointer saving that state after every step, so those behaviors come from the framework instead of your own code.

How do interrupt() and Command(resume=...) work together?

interrupt() pauses the graph inside a node and saves its state; the value you pass to graph.invoke(Command(resume=...), config) becomes the return value of that interrupt() call. So the node continues exactly where it left off, with the human's decision in hand — approve the refund and goto the tools node, deny it and send denial ToolMessages back to the model. The pause can carry a custom payload and be conditional, which is more flexible than compiling with interrupt_before.

Which checkpointer should I use in production?

Not InMemorySaver — it disappears when the process restarts, which defeats the point of pausing for a day. Production setups use a database-backed checkpointer such as the SQLite or Postgres packages. Durable execution is LangGraph's headline feature: agents persist through failures and can run for extended periods, resuming from where they left off.

When is a plain agent loop or create_agent enough instead of a graph?

When you don't need pausing, durable state, human approval, branching workflows, or replayable runs. For a simple tool-using chatbot that never pauses, a plain loop — or LangChain 1.0's create_agent, which wraps a standard tool-calling loop — is less code. Build the graph yourself when you need custom routing, approvals, or multi-step state, as in this support agent.

Wrapping This Up

LangGraph turns an agent into explicit state, nodes, and edges with a checkpointer underneath. That gives you resumable runs, human approval via interrupt() and Command(resume=...), per-thread memory, and time travel for debugging, at the price of some up-front structure. Start from the support agent above, then adapt it: add another tool, add a node that classifies requests before the agent runs, or add a second approval step for large amounts.

Should every agent be a graph? No, a simple loop is often enough. But once your agent needs to wait for a person, survive a restart, or take different paths depending on what it finds, expressing it as a graph stops being overhead and starts being the simplest way to get those behaviors reliably. If you're still mapping the ecosystem, the courses guide in this series covers where to learn next.

What are You Looking For?

esc