Contents
Figure 1: Three specialists, one deliverable — the crew hands work down the line
You ask one LLM to research, write, and edit a report, and it does all three badly. It skips sources, pads the middle, and edits its own mistakes into the final draft. A team of specialized agents fixes that, and CrewAI is a Python framework built for it.
This tutorial builds a three-agent content team from scratch: a researcher, a writer, and an editor. You'll learn the four core concepts, run the crew sequentially, switch to a managed crew, add a custom tool, and move to the YAML project layout.
What CrewAI Does
CrewAI lets you define agents with roles, give them tasks, and group them into a crew that runs the tasks in order. Each agent is an LLM wrapped in a persona and an optional set of tools. The framework handles passing one task's output to the next.
Why bother with multiple agents? Narrow instructions beat broad ones. A "Senior Research Analyst" prompt produces tighter research than "do everything," and an editor who never saw the writer's reasoning catches problems the writer can't. If you're still mapping the ecosystem, the framework comparison in this series shows where CrewAI sits against LangGraph, Pydantic AI, and the rest.
The Four Building Blocks
You need four concepts, and everything else builds on them:
- Agent: A role, a goal, and a backstory, plus optional tools and an LLM.
- Task: A description of the work, the expected output, and the agent responsible.
- Crew: The agents and tasks bundled together with a process that runs them.
- Process: How the crew executes, either sequentially or under a manager.
The role, goal, and backstory fields shape every prompt the agent sees. Write them like you'd brief a new hire.
Setup
Install the package in a clean virtual environment:
pip install crewai
CrewAI needs an LLM. By default it reads an OpenAI key from your environment, so set OPENAI_API_KEY before you run anything. You can point agents at other providers instead, so check the current docs for the model-string format your provider needs.
Step 1: Define Your Agents
Here's the researcher and the writer. The {topic} placeholder gets filled in when you start the crew.
from crewai import Agent
researcher = Agent(
role="Senior Research Analyst",
goal="Find the key facts about {topic}",
backstory="You dig up accurate, well-organized facts and flag anything uncertain.",
verbose=True,
)
writer = Agent(
role="Technical Writer",
goal="Turn research notes about {topic} into a clear 300-word brief",
backstory="You write short, plain-English briefs for busy engineers.",
verbose=True,
)
Set verbose=True while you build. You'll watch each agent's reasoning scroll by, which beats guessing why the output looks odd.
Step 2: Define Tasks and Chain Them
A task needs a description, an expected output, and an owner. The context argument is the key to collaboration: it feeds earlier tasks' outputs into later ones.
from crewai import Task
research_task = Task(
description="Research {topic}. List 5 key facts with a one-line confidence note for each.",
expected_output="A bullet list of 5 facts, each with a confidence note.",
agent=researcher,
)
writing_task = Task(
description="Write a 300-word brief on {topic} using only the research notes.",
expected_output="A brief of about 300 words in Markdown.",
agent=writer,
context=[research_task],
)
Write expected_output carefully. It acts as the acceptance test for the task, and vague outputs produce vague work. Ever wondered why one agent ignores your formatting? Check this field first.
Step 3: Assemble and Run the Crew
Add an editor with its own task, then bundle everything:
from crewai import Crew, Process
editor = Agent(
role="Reviewing Editor",
goal="Tighten the brief about {topic} so it is accurate and about 300 words",
backstory="You cut padding, fix structure, and flag claims that need sources.",
verbose=True,
)
editing_task = Task(
description="Edit the brief about {topic}. Cut padding, fix structure, and flag any claim that needs a source.",
expected_output="A final Markdown brief of about 300 words, plus a short list of flagged claims.",
agent=editor,
context=[research_task, writing_task],
output_file="brief.md",
)
crew = Crew(
agents=[researcher, writer, editor],
tasks=[research_task, writing_task, editing_task],
process=Process.sequential,
verbose=True,
)
result = crew.kickoff(inputs={"topic": "retrieval-augmented generation"})
print(result.raw)
The inputs dictionary fills every {topic} placeholder. Process.sequential means each task runs in list order, with its context passed forward. The editing task sets output_file="brief.md", so the final brief lands on disk.
Step 4: Give an Agent a Tool
Agents get more useful with tools. This custom tool counts words, so the writer can check its own length instead of guessing:
from crewai.tools import tool
@tool("word_counter")
def word_counter(text: str) -> str:
"""Count the words in a piece of text and return the count."""
return f"{len(text.split())} words"
writer = Agent(
role="Technical Writer",
goal="Turn research notes about {topic} into a clear 300-word brief",
backstory="You write short, plain-English briefs for busy engineers.",
tools=[word_counter],
)
Keep the docstring clear and specific. The agent reads it to decide when to call the tool, so a vague docstring means a tool nobody uses. CrewAI also ships a separate crewai_tools package with ready-made tools, such as web search, though most need their own API keys.
Step 5: Try a Managed Crew
Sequential crews follow a fixed script. A hierarchical crew adds a manager that assigns work and reviews results:
managed_crew = Crew(
agents=[researcher, writer, editor],
tasks=[research_task, writing_task, editing_task],
process=Process.hierarchical,
manager_llm="gpt-4o",
verbose=True,
)
A hierarchical crew needs a manager_llm or a custom manager agent — check current model names before you run it. Start sequential and switch only when your task order truly varies. Managers cost extra LLM calls, and a bad manager reroutes work in confusing ways. And if what you actually need is durable state, human approval interrupts, or time travel rather than a manager, that's a LangGraph job, not a process flag.
Step 6: Move to the YAML Layout
Hard-coded personas get messier as crews grow. CrewAI's project template moves them into YAML files and uses decorators to wire everything together. The scaffolding command creates the layout for you:
crewai create crew brief_crew
crewai run
Here's what an agents file looks like:
researcher:
role: Senior Research Analyst
goal: Find the key facts about {topic}
backstory: You dig up accurate, well-organized facts and flag anything uncertain.
And the crew class loads it:
from crewai.project import CrewBase, agent, crew, task
@CrewBase
class BriefCrew:
agents_config = "config/agents.yaml"
tasks_config = "config/tasks.yaml"
@agent
def researcher(self) -> Agent:
return Agent(config=self.agents_config["researcher"], verbose=True)
This split pays off fast. Non-programmers can edit prompts in YAML without touching Python, and you can version prompts separately from logic.
Design Tips That Matter
The same mistakes sink crews again and again:
- Keep roles narrow. "Editor" works better than "Content Specialist."
- Write testable outputs. "Five facts with confidence notes" beats "good research."
- Pass context deliberately. Give each task only the upstream results it needs.
- Limit tool access. Agents with ten tools call the wrong ones.
- Start with two agents. Add the third when you can name the problem it solves.
More agents doesn't mean better results. Each handoff adds latency, cost, and a chance for errors to compound.
Debug a Misbehaving Crew
Something will go wrong, so plan for it. Work through this checklist:
- Read the verbose logs. Find the first task where output drifts.
- Tighten that task's
expected_output. Vague targets cause most drift. - Check the context chain. An agent can't use notes it never received.
- Test agents alone. Run each task on its own before you chain them.
- Cap the cost. Watch token use, since agents loop when instructions conflict.
Then fix one thing at a time. Changing three prompts at once teaches you nothing.
Common Pitfalls
- Treating backstory as decoration. It steers tone and judgment.
- Skipping evaluation. Read the outputs, and compare against a single-prompt baseline.
- Trusting the final text. Agents state invented facts confidently, so verify claims.
- Forgetting secrets. Keep API keys in environment variables, never in your repo.
- Ignoring versions. CrewAI changes quickly, so pin your version in
requirements.txt.
Recommended Books
| Cover | Book | Description | Get it |
|---|---|---|---|
![]() |
AI Agents in Action | - hands-on coverage of agent loops, tools, and multi-agent teams. | View on Amazon |
![]() |
Designing Agentic AI Systems | - architecture decisions for roles, handoffs, and manager patterns. | View on Amazon |
![]() |
Build a Large Language Model (From Scratch) | - understand the models your agents are wrapped around. | View on Amazon |
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
When should you use CrewAI instead of a single prompt or a framework like LangGraph?
Use CrewAI when the work decomposes into distinct roles with clear handoffs — a researcher feeding a writer feeding an editor — and the task order is mostly fixed. Use a single prompt when one pass is enough, and reach for a graph framework such as LangGraph when you need durable state, human-in-the-loop interrupts, or branching that a process enum can't express.
What does the context argument actually do in a CrewAI task?
context lists upstream tasks whose outputs get injected into the task the agent receives. Without it, an agent only sees its own task description; with context=[research_task], the writer's prompt includes the researcher's output. Forgetting the context chain is the most common reason a later agent ignores earlier work.
When do you switch from Process.sequential to Process.hierarchical?
Only when the task order truly varies run to run and a manager should decide routing or review. Hierarchical crews add a manager LLM call per decision, which costs tokens and latency, and a weak manager can reroute work confusingly. Start sequential; switch when you can name a concrete ordering problem the fixed script can't solve.
How do you keep CrewAI costs under control?
Keep agents and tools few, cap task counts, set verbose=False in production, and pin your model choices. Watch token use during development — conflicting instructions make agents loop. Always compare the crew's output against a single-prompt baseline; if the crew doesn't clearly win, simplify it.
Wrapping This Up
CrewAI turns "one LLM doing everything" into specialists with clear jobs: define agents, give them tasks, chain context, and pick a process. Start with a sequential crew of two or three agents, add a tool when an agent needs one, and move to YAML when the prompts multiply.
Build the researcher-writer-editor crew this week on a topic you know well. Then compare its brief against a single-prompt version — the same discipline the from-scratch agent guide drills into every loop you write. If the crew doesn't win, simplify it until it does. New to agents altogether? The beginners guide covers the anatomy these roles sit on top of, and the courses guide maps where to learn next.


