Contents
Figure 1: The loop is the whole game — learn it before you adopt a framework
A chatbot answers a question. An agent does something about it: it looks up the order, calculates the refund, checks the policy, and reports back. That shift, from generating text to taking steps toward a goal, is what people mean by "AI agent," and it's simpler to build than the hype suggests. This guide explains what an agent actually is, walks through a working example in plain Python, then covers workflows, tool design, memory, the framework landscape, MCP, safety, and evaluation — plus a beginner's roadmap and the mistakes that sink most first projects.
What an Agent Actually Is
Strip away the marketing and an agent is four things working together:
- A model that decides what to do next.
- Tools: functions the model can call, such as searching, querying a database, or sending a message.
- A loop: call the model, run any tools it requests, feed the results back, and repeat until it finishes.
- State: the running record of the conversation and tool results, which gives the model context for its next decision.
A useful distinction from the practitioner literature is workflow versus agent. In a workflow, your code defines the path — step one, then step two, then step three — and the model fills in individual steps. In an agent, the model chooses the path itself, deciding which tool to call and when to stop. Workflows are more predictable, agents are more flexible, and many successful systems are workflows with a small agentic piece inside.
Build One From Scratch First
Frameworks hide the loop, and understanding the loop is the whole game, so build one before adopting one. This example uses Anthropic's Python SDK, but the structure is the same with any provider's tool-calling API. It gives a support agent two tools: order lookup and a safe calculator.
import ast
import json
import operator as op
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY
MODEL = "claude-sonnet-5-5" # check current model names before running
TOOLS = [
{
"name": "lookup_order",
"description": "Look up an order's status and total by order ID.",
"input_schema": {
"type": "object",
"properties": {"order_id": {"type": "string"}},
"required": ["order_id"],
},
},
{
"name": "calculate",
"description": "Evaluate basic arithmetic, e.g. '59.97 * 0.8'.",
"input_schema": {
"type": "object",
"properties": {"expression": {"type": "string"}},
"required": ["expression"],
},
},
]
ORDERS = {
"A100": {"status": "shipped", "total": 59.97},
"A101": {"status": "processing", "total": 19.99},
}
OPS = {ast.Add: op.add, ast.Sub: op.sub, ast.Mult: op.mul,
ast.Div: op.truediv, ast.USub: op.neg}
def safe_eval(node):
if isinstance(node, ast.Expression):
return safe_eval(node.body)
if isinstance(node, ast.Constant) and isinstance(node.value, (int, float)):
return node.value
if isinstance(node, ast.BinOp) and type(node.op) in OPS:
return OPS[type(node.op)](safe_eval(node.left), safe_eval(node.right))
if isinstance(node, ast.UnaryOp) and type(node.op) in OPS:
return OPS[type(node.op)](safe_eval(node.operand))
raise ValueError("unsupported expression")
def run_tool(name, args):
if name == "lookup_order":
return json.dumps(ORDERS.get(args["order_id"], {"error": "order not found"}))
if name == "calculate":
return str(safe_eval(ast.parse(args["expression"], mode="eval")))
return f"error: unknown tool {name}"
def run_agent(user_message, max_steps=8):
messages = [{"role": "user", "content": user_message}]
for _ in range(max_steps):
response = client.messages.create(
model=MODEL,
max_tokens=1024,
system="You are a support agent. Use tools for facts; never guess order data.",
tools=TOOLS,
messages=messages,
)
messages.append({"role": "assistant", "content": response.content})
if response.stop_reason != "tool_use": # model is done
return "".join(b.text for b in response.content if b.type == "text")
results = []
for block in response.content:
if block.type == "tool_use":
try:
output = run_tool(block.name, block.input)
except Exception as e: # report errors to the model
output = f"error: {e}"
results.append({"type": "tool_result",
"tool_use_id": block.id,
"content": output})
messages.append({"role": "user", "content": results})
return "Stopped: step limit reached."
print(run_agent("What's the status of order A100, and what is 20% off its total?"))
Read the loop slowly, because every agent framework is a variation on it. The model gets the messages and tool descriptions. If it asks for tools, your code runs them and appends the results. If it doesn't, it's finished. Set ANTHROPIC_API_KEY in your environment before running, and expect output like Order A100 is shipped; its total is $59.97, so 20% off is $47.98 — the model has to call both tools to answer.
Three details are worth copying into everything you build:
- A step limit (
max_steps) so a confused agent can't loop forever and burn money. - Errors returned as tool results, so the model can recover or explain instead of your program crashing.
- A safe calculator built on
astinstead ofeval, since tool inputs come from a model that can be manipulated by whatever text it reads.
Why Tool Descriptions Matter So Much
The model chooses tools based entirely on their names, descriptions, and parameter schemas — vague descriptions produce wrong calls. Write them like documentation for a new colleague: say what the tool does, when to use it, what each argument means, and what it returns. Keep the toolset small and distinct; ten overlapping tools confuse a model far more than three clear ones. If calls go wrong after a feature change, suspect the descriptions before you suspect the model.
Workflow Patterns Before "Real" Agents
When you know the shape of the task, a workflow is usually cheaper, faster, and easier to test than a free-roaming agent. Common patterns:
- Prompt chaining: a fixed sequence, each step feeding the next, such as extract, then classify, then draft.
- Routing: a first model call classifies the request and sends it down the right path.
- Parallelization: run independent subtasks at once, then combine the results.
- Orchestrator and workers: one model breaks a task into pieces and delegates them.
- Evaluator and optimizer: one call produces output and another critiques it, looping until it's good enough.
Reach for a fully autonomous loop only when the number and order of steps genuinely can't be known in advance.
Memory and Context
An agent's "memory" is mostly context management:
- Short-term: the message history itself. It grows with every step, so long tasks need trimming or summarizing to stay within limits and keep costs down.
- Long-term: facts that persist across sessions, stored externally (a database or vector store) and retrieved when relevant, often exposed as a search tool.
- Retrieval (RAG): giving the agent a tool that searches your documents, so it answers from sources instead of from memory — the same retrieval plumbing the agentic RAG guide in this series covers, wired into the loop above.
More context isn't always better. Stuffing everything into the prompt degrades focus and raises cost, so retrieve what's relevant and summarize the rest.
The Framework Landscape
Once you understand the loop, frameworks save real work: tracing, persistence, retries, handoffs, and integrations. The field moved fast in 2026, and several 2026 comparisons agree on the broad shape of the landscape:
- LangGraph models agents as graphs with explicit state and durable execution, which makes it a common default for complex, auditable workflows and human approval steps.
- OpenAI Agents SDK offers lightweight primitives, handoffs, guardrails, sessions, and built-in tracing, with sandbox execution added in April 2026.
- Claude Agent SDK, renamed from the Claude Code SDK in February 2026, packages the loop behind Claude's coding tools as a library, with subagents and deep MCP integration — the MCP fundamentals piece in this series shows how the protocol plugs tools in.
- Google ADK supports multiple languages, hierarchical agents, and Google Cloud deployment.
- Microsoft Agent Framework reached 1.0 GA on April 3, 2026, merging AutoGen and Semantic Kernel into one .NET and Python SDK. Per comparisons, new projects should start there rather than on AutoGen.
- CrewAI uses role-based "crews" and is often the fastest route to a multi-agent prototype.
- Pydantic AI and smolagents are lighter options, the former emphasizing type safety.
The trade-offs generally split two ways: graph-based tools like LangGraph give precise control, while model-driven loops like the OpenAI Agents SDK trade some control for simplicity. Provider SDKs make sense when you've committed to one model ecosystem. Much of the public comparison content comes from blogs and vendors with their own angles, and versions change quickly, so test a small prototype instead of trusting a ranking.
MCP and A2A in One Paragraph
MCP (Model Context Protocol) is a standard for connecting agents to tools and data sources, so one integration works across clients. Comparisons suggest all ten major frameworks now support MCP, seven of them natively. A2A (Agent2Agent) is a protocol for agents to talk to other agents, natively supported in Microsoft Agent Framework and Google ADK. As a beginner you can skip A2A and treat MCP as "how I plug in existing tools without writing custom glue."
Safety and Reliability Basics
Agents act, so mistakes cost more than a wrong chatbot answer. Build these in from the first version:
- Least privilege. Give the agent only the tools and permissions the task needs. A read-only lookup tool is far safer than a general database connection.
- Human approval for irreversible actions. Refunds, deletions, outbound emails, and payments should pause for confirmation.
- Assume prompt injection. Text from web pages, documents, emails, and tool outputs can contain instructions aimed at your agent. Never let retrieved content carry the authority of your system prompt — and treat tool output with the same skepticism the hallucination-reduction guide applies to retrieved documents.
- Budgets and timeouts. Cap steps, tokens, wall-clock time, and spend per run.
- Idempotent actions. Design tools so a retried call doesn't double-charge or double-send.
- Sandbox code execution. If the agent runs code, do it in an isolated environment.
- Log everything. Record every model call, tool call, and result, so you can reconstruct what happened.
Evaluating Agents
Agents are non-deterministic, so "it worked when I tried it" means little. Build a test set of realistic tasks with known good outcomes, then measure task success rate, number of steps, cost, and latency across repeated runs. Check trajectories, not just final answers: an agent that reaches the right answer by calling a dangerous tool is still a failure. Add regression tests whenever you fix a bug, and run adversarial tests too — craft inputs designed to distract, inject instructions, or force a wrong tool choice, and make them part of the suite.
A Beginner's Roadmap
Seven steps, in order — resist skipping ahead:
- Write the bare loop above and get one tool working.
- Add a second and third tool, and refine their descriptions until the model chooses correctly.
- Add retrieval over your own documents.
- Add tracing and a small evaluation set.
- Add guardrails: step limits, approvals, permission checks.
- Adopt a framework if you need persistence, handoffs, or observability you don't want to build yourself.
- Consider multiple agents only when one agent with well-chosen tools genuinely can't handle the job.
Good first projects include a documentation Q&A agent, a data-lookup assistant over a read-only database, or a triage agent that classifies and routes tickets.
Common Pitfalls
- Starting with a multi-agent system when a single agent or fixed workflow would do, which multiplies cost and failure modes.
- Writing vague tool descriptions, then blaming the model for choosing badly.
- Giving an agent broad permissions "for convenience."
- Forgetting step limits, then discovering a runaway loop on the invoice.
- Skipping evaluation, so you can't tell whether a change helped.
- Trusting tool output and retrieved text as if it were instructions from you.
- Adopting a framework before understanding the loop it wraps.
Recommended Books
| Cover | Book | Description | Get it |
|---|---|---|---|
![]() |
AI Agents in Action | hands-on coverage of the same loop, tool-calling, and agent patterns this guide introduces. | View on Amazon |
![]() |
Designing Agentic AI Systems | architecture and design decisions for agents that act in production, not just demos. | View on Amazon |
| LangGraph: Agentic Applications | the graph-based approach, once the bare loop in this guide is second nature. | View on Amazon |
Partner offer
Educative Unlimited: every AI course, one subscription
Interactive, browser-based courses with no local setup — a natural fit for the learning path above. One subscription unlocks the full generative AI, ML, and data catalog, so you can work through it at your own pace.
Affiliate link — we may earn a commission at no extra cost to you.
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
What is an AI agent, and how is it different from a chatbot?
A chatbot answers a question; an agent takes steps toward a goal. Concretely, an agent is four things working together: a model that decides what to do next, tools the model can call (lookups, calculations, searches, messages), a loop that runs the model, executes any requested tools, and feeds results back until the model stops, and state — the running message history that gives the model context. The key difference from a workflow is who chooses the path: in a workflow your code defines the sequence and the model fills in steps; in an agent the model picks which tool to call and when to stop.
Do I need a framework to build my first AI agent?
No — and you shouldn't for the first one. Frameworks hide the loop, and understanding the loop is the whole game. Build a bare agent in plain Python (or with your provider's SDK) with two or three tools, a step limit, and errors returned as tool results. Once that works, adopt a framework only when you need something you don't want to build yourself: durable state, human approval steps, handoffs between agents, or production tracing and observability.
Which AI agent framework should beginners use in 2026?
It depends on your ecosystem and how much control you want. LangGraph is the common default for complex, auditable workflows with explicit state and human-in-the-loop steps. The OpenAI Agents SDK offers lightweight primitives with handoffs, guardrails, sessions, and built-in tracing. The Claude Agent SDK (renamed from the Claude Code SDK in February 2026) packages a production loop with subagents and deep MCP integration, and Google ADK covers multiple languages with hierarchical agents on Google Cloud. Microsoft Agent Framework hit 1.0 GA on April 3, 2026, merging AutoGen and Semantic Kernel — comparisons suggest new .NET projects should start there. CrewAI is often the fastest route to a multi-agent prototype, while Pydantic AI and smolagents are lighter, type-safe options. Test a small prototype instead of trusting a ranking, since versions and vendor claims change fast.
How do I keep an AI agent from running away or doing something dangerous?
Treat safety as part of the design, not a later add-on. Cap steps, tokens, wall-clock time, and spend per run. Give the agent only the tools the task needs — a read-only lookup is far safer than a general database connection. Pause for human approval on irreversible actions like refunds, deletions, payments, and outbound emails. Assume prompt injection: text from web pages, documents, emails, and tool outputs can contain instructions aimed at your agent, so never let retrieved content carry the authority of your system prompt. Make tools idempotent so a retried call doesn't double-charge, sandbox any code execution, and log every model call, tool call, and result so you can reconstruct what happened.
Wrapping This Up
An agent is a model, a set of tools, and a loop with state — nothing more mysterious than that. Build the loop yourself once, write careful tool descriptions, add step limits and human approval for risky actions, and evaluate on realistic tasks. Then pick a framework based on your ecosystem and needs, whether that's LangGraph for stateful control, a provider SDK for tight integration, or Microsoft Agent Framework for .NET and Azure.
Do you need a multi-agent swarm to build something useful? Almost never. Start with the smallest design that could work, measure it, and add complexity only when the measurements say you must.

