Appearance
CrewAI
If you can describe a business process to a colleague in one sentence — "one researcher gathers material, one analyst organizes it, one editor polishes the draft" — CrewAI is the framework that translates that description straight into code. It pushes the entry cost of multi-agent systems down to the lowest tier among mainstream frameworks, at the price of a thick abstraction layer: when things go wrong, you end up debugging inside the framework. Based on the official docs for CrewAI v1.15.x (the current release line as of August 2026) plus real user feedback, this page covers what it is, how to use it, and when not to.
1. Positioning: Role-Play Multi-Agent Orchestration
CrewAI was created by João Moura, originally as a side project; in October 2024 it closed an $18M Series A led by Insight Partners (with participation from Andrew Ng, HubSpot CTO Dharmesh Shah, Replit CEO Amjad Masad, and others). v1.0 went GA in October 2025, with the officially disclosed figures at the time: 1.4 billion cumulative agentic executions, 60%+ of the Fortune 500 using it (vendor-reported, including trials), 40k GitHub stars, and 1.8M monthly downloads; by mid-2026 the star count had passed 50k. As of August 2026 the latest release line is v1.15.x, MIT-licensed.
Its core metaphor is role-play: an agent = role + goal + backstory. These three fields aren't decoration — they get assembled into the system prompt and directly shape the model's behavior. The mental overhead of this design is minimal: you don't need graph theory, state machines, or event sourcing; you just need to know how to "write a job description."
Two easily confused facts, clarified up front:
- CrewAI is not a wrapper around LangChain. Early versions (before ~0.30) were indeed built on LangChain, but the dependency was removed long ago and the kernel rewritten. Today it's a fully independent Python framework (requires Python >= 3.10, < 3.14) that uses LiteLLM only as the model-routing layer. Articles claiming it's "based on LangChain" are outdated.
- It is not a single-agent framework. Using CrewAI for a single agent is overkill — calling the LLM API directly or using the OpenAI Agents SDK is more appropriate. Its value lies entirely in "multiple agents with division of labor working together," which is also the fundamental starting point when comparing it with LangGraph and others (see Multi-Agent Architecture).
Internally, CrewAI uses a two-layer architecture, and this is the key to understanding it:
┌──────────────────────────────────────────────────────────┐
│ Flows (event-driven workflows) — the control layer │
│ @start / @listen / @router — state-machine orchestration│
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ Crew A │ → │ deterministic│ → │ Crew B │ │
│ │ (autonomous) │ │(plain Python)│ │ (autonomous) │ │
│ └──────────────┘ └──────────────┘ └──────────────┘ │
│ Inside each Crew: │
│ Agent(role+goal+backstory) + Task │
│ → Process.sequential / hierarchical │
└──────────────────────────────────────────────────────────┘In one sentence: Crews handle autonomy, Flows handle control. After v1.0, the official recommendation is to use a Flow as the skeleton of production applications, with Crews demoted to execution units inside the Flow.
2. Core Concepts: Agent / Task / Crew / Process
Agent: three strings decide behavior
python
from crewai import Agent
from crewai_tools import SerperDevTool
researcher = Agent(
role="Senior AI researcher",
goal="Dig up the latest developments on {topic}; only trust verifiable sources",
backstory="You've spent a decade doing technical research and are known for finding primary sources and refusing secondhand claims.",
tools=[SerperDevTool()],
llm="openai/gpt-4o", # routed via LiteLLM; supports hundreds of providers
verbose=True, # print every reasoning step; a must while debugging
max_iter=20, # max agent-loop iterations per task
max_rpm=10, # rate limit so you don't blow through API quotas
allow_delegation=False, # whether the agent may hand tasks to other agents
)role / goal / backstory are required, and CrewAI assembles them into the system prompt. This means the backstory is not superstition: writing it concretely ("you are known for rejecting secondhand sources") produces noticeably steadier output than writing it vaguely ("you are a helpful AI") — essentially Prompt Engineering encapsulated at the framework layer.
Other high-frequency parameters: respect_context_window=True (on by default; summarizes history instead of erroring when the context window overflows), memory=True (enables short-term/long-term/entity memory, which depends on an embedding model underneath — see Memory Systems), reasoning=True (reflect and plan before executing), inject_date=True (inject the current date into the prompt, useful for time-sensitive tasks).
One change worth noting from the 1.x era: allow_code_execution and the built-in CodeInterpreterTool have been removed, with the official recommendation to use standalone sandbox services like E2B or Modal for code execution (a security decision — see Agent Security).
Task: a description plus an expected output
python
from crewai import Task
from pydantic import BaseModel
class ResearchNotes(BaseModel):
key_findings: list[str]
sources: list[str]
research_task = Task(
description="Research the state of industry adoption for {topic} in 2026, separating facts from vendor marketing.",
expected_output="A set of structured research notes with key findings and a source list.",
agent=researcher,
output_pydantic=ResearchNotes, # force structured output
# context=[other_task], # explicitly depend on an upstream task's output
# output_file="output/notes.md", # write to disk
)expected_output is the most underrated field in CrewAI. It isn't just for humans to read — the framework uses it to validate and steer the output format. The more it reads like an acceptance criterion ("3-5 items, each with a URL, facts separated from opinions"), the lower your rework rate.
Crew and Process
A Crew assembles agents and tasks, and process decides the scheduling strategy:
Process.sequential(default): tasks run in list order, with each task's output flowing automatically into the next task's context. Covers 80% of scenarios.Process.hierarchical: introduces a manager agent responsible for dispatching tasks and accepting results — think "a manager leading a team." It requires an additionalmanager_llmormanager_agent. It sounds great, but in practice stability is mediocre — the manager's dispatch quality depends entirely on the model, and slightly complex tasks tend to devolve into repeated re-assignments and loops, with token consumption multiplying. Most teams eventually fall back to sequential + Flow routing.
python
from crewai import Crew, Process
crew = Crew(
agents=[researcher, analyst, writer],
tasks=[research_task, analysis_task, writing_task],
process=Process.sequential,
memory=True,
verbose=True,
max_rpm=20, # crew-level rate limit; overrides per-agent settings
checkpoint=True, # new in 1.x: archive after each task, resumable after interruption
)
result = crew.kickoff(inputs={"topic": "on-device LLMs"})
print(result.raw) # the final task's raw output
print(result.pydantic) # if output_pydantic was defined
print(result.token_usage) # total token consumptionThe kickoff family has several variants; the ones commonly used in production:
kickoff_for_each(inputs_list): run over a list of inputs one by one — e.g., batch-processing 100 articles;akickoff()/akickoff_for_each(): natively async, the first choice for high-concurrency scenarios;kickoff_async(): thread-wrapped async, officially marked as compatibility-only — useakickoffin new code;- The CLI's
crewai replay -t <task_id>: replay from a given task, saving tokens when debugging long pipelines.
3. Flows: Event-Driven Workflows
Crews solve "how do agents collaborate"; Flows solve "how does the whole automation pipeline move." This is the watershed between CrewAI as a "multi-agent toy" and CrewAI as a "business automation tool," and it's where it diverges from pure agent frameworks.
The Flow model is simple: class methods plus three decorators, with state hanging off self.state:
python
from crewai.flow.flow import Flow, listen, start, router, or_
from pydantic import BaseModel
class ContentState(BaseModel): # structured state (an untyped dict-style state also works)
topic: str = ""
draft: str = ""
score: float = 0.0
class ContentPipeline(Flow[ContentState]):
@start() # entry point; there can be several, triggered in parallel
def set_topic(self):
self.state.topic = "on-device LLMs"
@listen(set_topic) # listen for the upstream method's completion event
def research(self):
result = research_crew.kickoff(inputs={"topic": self.state.topic})
return result.raw # return values are passed to downstream listeners
@router(research) # route: pick a branch based on the returned string
def grade(self, notes: str):
# call a scoring agent here, or use plain Python rules
return "publish" if self.state.score > 0.8 else "revise"
@listen("publish")
def publish(self):
pass # publishing logic goes here
@listen(or_("revise", "low_quality")) # or_ / and_ combine trigger conditions
def request_revision(self):
pass
flow = ContentPipeline()
flow.plot("pipeline") # generates an HTML flow diagram — handy for reviews
flow.kickoff()A Flow can mix three kinds of building blocks: full Crews, individual agents, and plain Python functions (querying databases, calling internal APIs, running rule engines — without spending a single LLM call). The production-recommended pattern: write all deterministic logic as plain Python in the Flow, and hand only the steps that genuinely need reasoning to a Crew. This is the key to controlling cost and predictability (see Cost Engineering).
Starting with v1.0, Flows also gained persistence and recovery (the @persist decorator with SQLite/custom backends) and human-in-the-loop pauses that wait for human input (for the corresponding patterns, see Human-in-the-Loop).
4. Hands-On: Getting a Minimal Project Running
The official recommendation is the CLI scaffold (with uv dependency management built in):
bash
uv pip install crewai 'crewai[tools]'
crewai create crew market_research
cd market_research
# edit .env and fill in OPENAI_API_KEY, SERPER_API_KEY
crewai runIn the v1.15 era, the scaffold generates a JSONC-configuration-style project (crew.jsonc + agents/*.jsonc) that runs without a single line of Python:
jsonc
// crew.jsonc
{
"name": "Market Research Crew",
"agents": ["researcher", "analyst"],
"tasks": [
{
"name": "research",
"description": "Research {topic} and collect the most relevant facts.",
"expected_output": "Structured research notes about {topic}.",
"agent": "researcher"
},
{
"name": "analysis",
"description": "Analyze the research and write a concise report.",
"expected_output": "A markdown report with findings and recommendations.",
"agent": "analyst",
"context": ["research"],
"output_file": "output/report.md"
}
],
"process": "sequential",
"inputs": { "topic": "AI Agents" }
}jsonc
// agents/researcher.jsonc
{
"role": "{topic} Senior Researcher",
"goal": "Find accurate and current information about {topic}.",
"backstory": "You are a careful researcher who cites clear evidence.",
"llm": "openai/gpt-4o",
"tools": ["SerperDevTool"]
}Older projects can use crewai create crew <name> --classic for the YAML + @CrewBase/@agent/@task/@crew decorator route, which remains supported. Personal advice: the configuration style fits standardized pipelines (ops can edit them without touching code), while the decorator style fits scenarios needing deep customization.
The security boundary of JSONC projects
The custom:<name> tools and {"python": "module.attribute"} callbacks in a JSONC config execute local Python code at load time. Only run crew projects from trusted sources — this is functionally equivalent to running a script someone handed you.
The debugging trio: verbose=True for the full reasoning process; built-in free tracing since v1.0 (OpenTelemetry-based, no third party needed); and crewai log-tasks-outputs + crewai replay -t <task_id> to replay from a middle task. For a more systematic observability setup, see Observability.
5. CrewAI vs AutoGen vs LangGraph: Design Philosophy Differences
The three frameworks disagree not on feature checklists but on "how should an agent system be thought about":
| Dimension | CrewAI | AutoGen | LangGraph |
|---|---|---|---|
| Core metaphor | Company/team: roles + job descriptions | Group chat: agents as conversation participants | Circuit diagram: nodes + edges + state |
| Orchestration paradigm | Procedural (sequential/hierarchical) + event-driven Flows | Message-driven async actor model | Explicit state graph; you define every transition |
| Abstraction level | Highest — convention over configuration | Mid-high | Lowest — everything is explicit |
| Time to first run | Half an hour | Half a day | One to two days (state/reducer/edge to digest) |
| Fine-grained control | Medium — the Flow layer suffices; Crew internals are a black box | Mid-high | Highest — every step is intervenable |
| Debugging when things break | Painful — thick abstractions | Medium | Relatively good — state transitions are fully explicit |
| Typical user profile | Business automation, content-pipeline teams | Research/experimental multi-agent conversations | Platform teams, complex production systems |
A blunt but accurate summary:
- CrewAI bets that "most multi-agent needs are really pipelines." The role metaphor and Flows are two granularities of the same pipeline idea. The bet paid off — it became the fastest-to-ship framework for content production and business process automation.
- LangGraph bets that "serious systems need explicit state machines." It gives you Turing-complete control, at the price of putting all the complexity in front of you with nothing to hide behind.
- AutoGen bets that "agent collaboration is fundamentally conversation." The message-driven actor model is the most flexible and the most academic — great for exploring novel collaboration patterns, but roundabout in enterprise scenarios where "the process is basically fixed and just needs automating."
In real engineering the three aren't mutually exclusive: a common combination is CrewAI for quick validation and standardized pipelines, with the genuinely critical core path rewritten in LangGraph; or the reverse — a LangGraph graph delegating one node to a CrewAI Crew. For selection methodology, see the framework selection overview.
6. Enterprise Features: AMP and Deployment
Beyond the open-source framework, CrewAI's commercial product is CrewAI AMP (Agent Management Platform, formerly CrewAI Enterprise). From what the official GitHub and docs disclose:
- A unified control plane: managed deployment, version management, environment isolation, safe redeploys, in both cloud-hosted and on-premise forms;
- Observability: real-time traces, metrics, and logs (the open-source version has shipped free tracing since v1.0; AMP adds enterprise retention and alerting on top);
- Governance: RBAC, audit logs, team management;
- Triggers: direct connections to event sources like Gmail, Slack, Salesforce, Outlook, and HubSpot — an incoming event automatically triggers a Crew/Flow;
- Visual Agent Builder: a no-code interface for building agents, aimed at non-developer roles;
- 24/7 enterprise support.
AMP has a free tier (the Crew Control Plane can be tried for free); paid tiers bill by execution, and third-party compilations put Professional at roughly $25/month to start — pricing changes often, so check the official site before signing. The deployment path: crewai run locally → push to GitHub → one-click deploy from the AMP console.
On the partnership side, the official materials name IBM, PwC, NVIDIA, and Arize, Databricks, Galileo, among others; they also claim engineers inside IBM, Microsoft, Walmart, SAP, Adobe, and PayPal use the open-source version. Take "Fortune 500 adoption" numbers with a grain of salt — the counting rules are generous by industry convention (having installed the package counts), so treat them as color, not procurement evidence.
7. Community Reception and Common Criticisms
Positive feedback clusters around a few points: fast to get started ("I had the idea running in an afternoon"), an intuitive role model ("if you know how to manage a human team, you know how to use CrewAI"), and solid docs and official courses (learn.crewai.com; the company claims 100k+ certified developers).
The criticisms cluster just as tightly, ranked by frequency:
- Thick abstractions make failures hard to trace. The most-cited complaint. The happy path is smooth, but once an agent behaves unexpectedly, you're facing three stacked black boxes — the framework's giant assembled system prompt, LiteLLM routing, and the internal agent loop — and debugging often becomes reading framework source code. A common revenge move in the community is to bypass Crews entirely and use only Flows + single LLM calls.
- Token consumption runs high. Assembling role/goal/backstory, passing context between tasks, and the delegation mechanism all inflate token usage; the hierarchical mode's manager repeatedly re-dispatching is an amplifier. Launch without caps on
max_iterandmax_rpmand the bill will teach you a lesson. - Frequent breaking changes in the 0.x era. Upgrade pain was the norm through 2024-2025; v1.0 (October 2025) came with an official API freeze and long-term stability pledge, and the 1.x release line over the following year has indeed focused on bug fixes and features — a promise kept, for now.
- The hierarchical mode doesn't live up to its name. The manager agent's dispatch quality is entirely at the model's mercy; the community consensus is to prefer sequential + explicit Flow routing and treat hierarchical as experimental.
- Anonymous telemetry is on by default. The framework reports version, OS, agent counts, role names, and similar metadata by default (officially, no prompts or business data unless
share_crew=Trueis explicitly enabled). In enterprise environments, remember to setOTEL_SDK_DISABLED=true. - A developer framework, not a business product. There's no business-user console (AMP's Visual Builder is filling that gap), and non-technical teams will hit a wall — this is a deliberate contrast with low-code platforms like Dify and Coze.
A pragmatic take
The best way to use CrewAI is "down-scoped": treat the Flow as the workflow engine and only spin up a Crew at the nodes that genuinely need multi-role reasoning; anything writable as plain Python steps should never go to an agent. Using it as "a workflow framework with agent capabilities" has a much higher success rate than using it as "an autonomous multi-agent system."
8. Where It Fits
Good fit:
- Content production pipelines: research → writing → fact-checking → polishing, with clear roles per stage and explicit acceptance criteria. This is CrewAI's sweetest spot, and all the official examples (job description, trip planner, stock analysis) are exactly this type.
- Business process automation: RevOps reports, KYC pre-screening, meeting-notes distribution, support ticket triage — paired with AMP's Triggers (Gmail/Slack/Salesforce event-driven), you can assemble a usable system quickly.
- Rapidly validating multi-agent ideas: a prototype in half an hour to test whether a task really needs multiple agents, before deciding whether to rewrite the core path in LangGraph.
Poor fit:
- Scenarios where a single agent suffices: adding role/goal/backstory won't make a single agent stronger — only more expensive.
- Hard real-time / low-latency scenarios: framework scheduling overhead plus multi-turn LLM calls; think twice for latency-sensitive products.
- Systems needing precise control over every reasoning step: finance compliance, healthcare, anywhere every step must be explainable and intervenable — go straight to LangGraph or hand-write the agent loop.
- No-code teams: use Dify or Coze instead; don't touch a Python framework.
One final learning suggestion: CrewAI is the fastest path to multi-agent intuition, but don't stop inside the metaphor it hands you. After finishing this page, read AutoGen's "conversation as collaboration" and LangGraph's "graph as control" side by side. Only when all three philosophies make sense to you are you genuinely equipped to choose.
References
- CrewAI official docs — the authoritative source for Concepts (Agents/Crews/Flows), Quickstart, and Enterprise docs; all APIs on this page follow the v1.15.x docs.
- crewAIInc/crewAI · GitHub — source code, MIT license, AMP Suite capability notes, and the original telemetry policy.
- CrewAI Releases — release history; used to verify the v1.15.x changes (latest v1.15.17 as of 2026-08).
- CrewAI OSS 1.0 GA announcement — the source of the official figures: 1.4B executions, 60% of the Fortune 500, and the investor list.
- The Hidden Costs of LangChain, CrewAI, PydanticAI and Others — a representative criticism of CrewAI's rigid abstractions and hidden costs; opinionated, but worth reading as a counterweight.
- AutoGen vs CrewAI: Which AI Agent Framework Should You Use? — a comparison aggregating real Reddit user feedback; one of the main sources for this page's community-reception section.
- CrewAI in Python · Real Python — a high-quality third-party tutorial, good as a hands-on supplement.
- CrewAI pricing compilation (third party) — a roundup of public AMP free/paid tier information; verify against the official site's live prices before signing.