Skip to content

Multi-Agent Systems

At a glance When multi-agent is worth it — and when it's over-engineering. This page unpacks the opposing positions of Anthropic Research and Cognition, five orchestration patterns, the state of the A2A protocol, real-system case studies, and quantified costs, giving you a decision framework for moving from a single agent to many.

Multi-Agent Systems ​

Multi-agent was the single most contested topic in agent engineering in 2025. Anthropic backed Claude's Research feature with an orchestrator-worker multi-agent system and claimed it outperformed a single agent by 90.2%; almost simultaneously, Cognition (the company behind Devin) published "Don't Build Multi-Agents," asserting that most multi-agent architectures are a fragile anti-pattern. Both articles are right — because they are talking about differently shaped tasks.

The goal of this page is to untangle that debate, then hand you an actionable catalog of orchestration patterns and a decision framework: when to split, into what, and what to ask yourself before splitting.

1. Why Multiple Agents at All ​

There are only three engineering-sound reasons to split a task from one agent across many:

1. Context isolation (the strongest reason) ​

An agent's context window is its only workbench. After searching 20 web pages and reading dozens of files, the workbench is crammed with raw noise, and the instructions that actually matter get pushed out of the attention center. Subagents each have their own context window; after rummaging through the noise, they hand back only compressed conclusions — the lead agent's context stays clean throughout.

Anthropic puts it bluntly: "search is fundamentally compression" — subagents act as intelligent filters, exploring vast corpora and distilling only the most important tokens back to the lead agent. This is also Claude Code's subagent design philosophy: the subtask agent's only job is "to answer one clearly defined question," and its investigation never lands in the lead agent's history, which lets the lead agent sustain longer traces (see Claude Code Anatomy and Context Engineering).

2. Specialization ​

Different subtasks need different system prompts, different toolsets, even different models. Forcing a "web browsing agent" and a "code execution agent" to share one prompt and one toolset leaves both under-equipped. Specialization lets each part's context be precisely tailored — the same logic as human organizations, except that an agent's "specialization" is the combination of prompt + tools + model.

3. Parallelism ​

Breadth-first tasks ("find every board member of every company in the S&P 500 information technology sector") decompose naturally into dozens of independent sub-queries. A single agent searching serially is unusably slow; multiple agents in parallel finish in minutes. Anthropic's data: a lead agent launching 3-5 subagents in parallel, each subagent calling 3+ tools in parallel — these two changes cut research time on complex queries by up to 90%.

Why you mostly shouldn't: the two camps ​

This is the most important part of this page, and both articles deserve to be read in the original.

The pro side: Anthropic's "How we built our multi-agent research system" (June 2025)

  • Internal evals: a multi-agent system with Claude Opus 4 as the lead and Sonnet 4 as subagents beat single-agent Opus 4 by 90.2% on internal research evaluations.
  • The key insight: a multi-agent system is essentially "spending more tokens to solve the problem." On the BrowseComp evaluation, three factors explained 95% of the performance variance, with token usage alone explaining 80%. The multi-agent architecture extends token consumption past a single agent's ceiling by using multiple independent context windows.
  • The applicability boundary (drawn by Anthropic itself): high-value tasks that are heavily parallel, whose information exceeds a single context window, and that need to integrate many complex tools. They gave the counter-example too: most coding tasks are far less parallelizable than research tasks, and agents' ability to coordinate and delegate with each other in real time still isn't there.

The con side: Cognition's "Don't Build Multi-Agents" (Walden Yan, June 12, 2025)

Cognition's argument rests on two "context engineering principles":

  1. Share context, and share the full agent trace — not just individual messages.
  2. Actions carry implicit decisions, and conflicting implicit decisions produce bad outcomes.

The "clone Flappy Bird" example: the task is split into "build the background" and "build the bird"; two subagents each produce components that misread the requirements — the background comes out Mario-styled, the bird looks nothing like a game asset — and the merging agent is left helpless with two mutually contradictory half-products. Even when the original task is copied verbatim to every subagent, the two still can't see each other's actions, and their outputs clash stylistically. The root cause: every agent's actions carry implicit decisions, and parallel splitting makes those decisions impossible to align.

Cognition's alternative: default to a single-threaded linear agent; when context is about to overflow, use a dedicated compaction model to compress the action history into key decisions and events (they even fine-tuned a small model for this internally), rather than splitting the context.

How do you reconcile the two camps?

Notice that the two companies' task shapes differ. Anthropic's scenario is breadth-first information retrieval: subtasks are nearly independent, outputs are textual conclusions, merging is cheap, and "reading" aligns far more easily than "writing." Cognition's scenario is software engineering: subtasks share one codebase, and any two agents writing code will stomp on each other. Rule of thumb: subtasks read-only and outputs separable → multi-agent is viable; subtasks write the same state and decisions are coupled → multi-agent is a disaster. Both camps are right within their own boundaries.

An honest status check

As of 2026, multi-agent systems are production-proven for "research/retrieval" products (Anthropic Research and the various Deep Research products), while coding has largely retreated to the conservative shape of "one lead agent + read-only subtasks." Claude Code's subagents never write code in parallel with the lead agent — not for lack of capability, but by deliberate reliability design. Treating multi-agent as the default architecture was still over-engineering in 2026.

2. The Orchestration Pattern Catalog ​

The five patterns below cover 95% of multi-agent systems in industry. When selecting, first get clear about where information flows and where decision authority sits — then draw the arrows.

1. Supervisor / Orchestrator ​

              ┌──────────────────┐
   user request → │   Supervisor     │ → final answer
              │ (plan/delegate/  │
              │   synthesize)    │
              └───┬───┬───┬──────┘
                  │   │   │   delegate subtasks (with clear boundaries)
          ┌───────┘   │   └────────┐
          ▼           ▼            ▼
     ┌───────────┐ ┌───────────┐ ┌───────────┐
     │ Worker A  │ │ Worker B  │ │ Worker C  │  ← each with its own context
     │ (search)  │ │ (code)    │ │ (analysis)│
     └─────┬─────┘ └─────┬─────┘ └─────┬─────┘
          └─────────────┴─────────────┘
              only compressed conclusions flow back
  • Decision authority concentrates in the supervisor; workers are passive execution units.
  • Anthropic Research, Magentic-One, and most Deep Research products use this pattern.
  • Fits: tasks that can be decomposed up front (or iteratively), loosely coupled subtasks, and the need for a single exit that synthesizes.
  • Traps: the supervisor becomes a single point of bottleneck and failure; vague delegation briefs cause workers to duplicate work or leave coverage gaps (early versions of Anthropic's system hit exactly this).

2. Handoff ​

  ┌────────┐   out of scope?      ┌────────┐   ┌────────┐
  │support │ ───────────────────→ │refund  │ → │tech    │
  │ agent  │   hands over the     │ agent  │   │ agent  │
  └────────┘   full conversation  └────────┘   └────────┘
     ↑      at any moment exactly one agent holds "control of the conversation"
  • Only one agent is active at a time; a handoff tool transfers control (and context) wholesale to the next agent. OpenAI's Swarm experiment and the Agents SDK productized this pattern (see OpenAI Agents SDK).
  • Fits: support routing and staged processes (pre-sales → order → after-sales) — essentially "a state machine with LLM routing."
  • The difference from a supervisor: after a handoff the original agent exits, and nobody aggregates above it; a supervisor stays online throughout.

3. Pipeline ​

  input → [retrieval agent] → [analysis agent] → [writing agent] → [proofing agent] → output
             each stage's output is the next stage's input; the direction is fixed
  • Strictly speaking, this isn't necessarily "multi-agent" — if each step is a deterministic call, it's a workflow rather than an agent system. Anthropic's "Building Effective Agents" classifies this as prompt chaining / workflow, explicitly distinct from agents.
  • Fits: tasks with predictable steps and clear per-step quality criteria (report generation, translation review, data processing).
  • This is the cheapest and most reliable pattern. If a pipeline solves it, don't reach for an agent — let alone several.

4. Group Chat / Debate ​

        ┌─────────────────────┐
        │   Shared message board   │
        └─────────────────────┘
           ↑        ↑         ↑         ↑
     ┌─────────┐ ┌───────┐ ┌─────────┐ ┌─────────┐
     │proposer │ │critic │ │retriever│ │moderator│
     └─────────┘ └───────┘ └─────────┘ └─────────┘
        rounds of speaking until convergence or the round cap
  • AutoGen's GroupChat and academic multi-agent debate belong here. All agents read and write the same message board, with a moderator (or a rotation policy) deciding who speaks.
  • Fits: open-ended brainstorming and scenarios that need adversarial review (red-teaming, proposal reviews).
  • Warning: this is the pattern Cognition called out by name. Without a clear convergence mechanism, debates tend to spin their wheels, exchange flattery, or wander off topic, with uncontrollable token burn. Successful production deployments are rare.

5. Hierarchical Teams ​

                 ┌────────────────┐
                 │ Top supervisor │
                 └───────┬────────┘
              ┌──────────┴──────────┐
              ▼                     ▼
     ┌───────────────┐     ┌──────────────────┐
     │ Research lead │     │ Engineering lead │
     └─┬────┬────┬──┘     └──┬────┬────┬────┘
       ▼    ▼    ▼           ▼    ▼    ▼
    [search][read][analyze] [frontend][backend][tests]
  • The recursive version of the supervisor pattern: when a single supervisor manages too many workers, or the task has natural domain layering, introduce middle managers.
  • Fits: large codebase refactors, multi-domain research projects. LangGraph's official docs have a dedicated hierarchical teams tutorial.
  • The cost: every added layer distorts the information one more time, and debugging difficulty and token cost rise in step.

Patterns can be combined: Anthropic Research is "supervisor pattern + each worker running its own internal pipeline," and Magentic-One is "supervisor + a fixed crew of specialists."

3. The Context-Engineering View: What Subagents Are Actually For ​

Strip away all the organizational metaphors and a subagent does exactly one thing technically: run a disposable exploration inside an independent context window, then write compressed conclusions back into the main context.

Three corollaries follow directly:

  1. A subagent's output must be "compressed." If a subagent dumps 50k tokens of raw search results straight back into the lead agent, isolation has failed. Anthropic defines the subagent as an "intelligent filter" — the filtering is the subagent's core competence.
  2. The delegation brief must be "self-contained." Subagents can't see the lead agent's conversation history, so the task description must contain the goal, the output format, the available tools, and the boundaries. Anthropic's lesson: a one-line brief like "research the semiconductor shortage" leads multiple subagents to run identical searches or each wander off — they observed three subagents where two were duplicating an investigation of the 2025 supply chain while the third went off researching the 2021 automotive chip crisis.
  3. Isolation has a price, and the price is exactly what Cognition criticized. Subagents can't see each other, so any task that requires "aligning implicit decisions" should not be split. Context isolation and context sharing are a fundamental tension, and your task's shape decides which side to pick.

The lead agent has a context limit of its own. Anthropic's approach: the lead agent writes its plan into external Memory for persistence, because context beyond 200k tokens gets truncated and the plan must be protected first. For the full "externalized plan + context compaction" combination, see Context Engineering and Memory Design.

An operational test

Draw the task's "information flow": if a subtask's input can be written as a self-contained brief and its output compressed into a few hundred tokens of conclusions, it's a good subagent candidate. If you find yourself copying large chunks of conversation history into the brief, or the subagent's output needs the lead agent to review it line by line before it's usable — the split is wrong; go back to a single agent.

4. Communication and Protocols ​

Shared state vs message passing ​

Multi-agent systems have two basic styles of internal communication:

  • Shared state: all agents read and write the same graph state or message history. LangGraph's StateGraph and AutoGen's GroupChat both work this way. The upside is no information loss (satisfying Cognition's first principle); the downside is that state bloats with the agent count and agents become tightly coupled.
  • Message passing: agents exchange only well-defined message/task objects; internal states are invisible to each other. This is the microservices-style approach, and it's the A2A protocol's approach too. The upside is decoupling and cross-organizational reach; the downside is exactly what Cognition warned about — "insufficient context transfer."

Rule of thumb: multi-agent within one process, one team, one codebase → shared state; agent collaboration across processes, teams, or organizations → message passing (protocolized).

The A2A protocol: cross-organizational agent interoperability ​

A2A (Agent2Agent) is an open protocol released by Google on April 9, 2025, with 50+ technology partners backing it at launch; on June 23, 2025, Google donated it to the Linux Foundation as the Agent2Agent project. By April 2026 (the first anniversary), the protocol's progress included:

  • A stable v1.0 specification: multi-protocol support, enterprise-grade multi-tenancy, modernized security flows, and migration paths for early adopters.
  • Supporting organizations grew from 50+ to 150+, including AWS, Cisco, Google, IBM, Microsoft, Salesforce, SAP, and ServiceNow.
  • Built into cloud platforms: Microsoft integrated A2A into Azure AI Foundry and Copilot Studio; AWS supports it through Amazon Bedrock AgentCore Runtime.
  • SDKs expanded from Python-only to five languages (Python, JavaScript, Java, Go, .NET), and the core repo passed 22,000 GitHub stars.
  • Signed Agent Cards: cryptographic signatures verify an agent's identity, answering "how do I know the agent on the other side is who it claims to be?"
  • The ecosystem's edge extended to payments: the Agent Payments Protocol (AP2) now has 60+ payment and financial institutions behind it.

A2A's core abstraction isn't complicated:

  Client Agent                               A2A Server (Remote Agent)
      │                                          │
      │  1. GET /.well-known/agent-card          │   ← discovery: capabilities, skills, auth
      │ ←────────────── Agent Card ──────────────│
      │                                          │
      │  2. authenticate per the Agent Card      │
      │                                          │
      │  3. POST sendMessage / sendMessageStream │   ← task: the Task lifecycle
      │ ───────────── Task (submitted) ─────────→│
      │ ←── TaskStatusUpdateEvent (working) ─────│   (status pushed over SSE)
      │ ←── TaskArtifactUpdateEvent (artifact) ──│   ← artifacts flow back
      │ ←── TaskStatusUpdateEvent (completed) ───│

The key design trade-offs are clearest when contrasted with MCP:

DimensionMCPA2A
What it solveshow an agent connects to tools/datahow an agent collaborates with other agents
What's on the other endstateless tools with fixed functionsautonomous agents capable of multi-turn negotiation
Interaction granularitysingle function callslong-running operations (LRO), streaming, multi-turn conversations
Internal visibilitytool implementation transparent to the calleropaque execution: no internal logic, memory, or tools exposed
StewardLinux FoundationLinux Foundation (since June 2025)

The two are complementary stacks: an agent uses MCP internally to reach its own tools, and A2A externally to negotiate with other people's agents. The A2A documentation stresses that "agents are not tools" — wrapping an agent as a stateless tool call throws away the multi-turn capabilities of negotiation, clarification, and delegation. For more MCP detail, see Tools & MCP.

Protocols are not silver bullets

A2A standardizes the "envelope" — discovery, authentication, the task state machine — not the "letter inside." Whether two agents can actually collaborate still depends on how clearly the task is described and whether output formats were agreed on in advance. Protocols solve integration cost; they do not solve the context-alignment problem Cognition raised. For cross-organizational work, that means investing more design effort in the task schema than in the prompt.

5. Real-System Case Studies ​

Anthropic Research: the production template for orchestrator-worker ​

Referenced several times above; here are the engineering details in one place:

  • Architecture: the LeadResearcher analyzes the query, sets strategy, writes the plan to Memory so it survives truncation, then creates multiple subagents in parallel to search; results are aggregated, with another round of subagents appended as needed; a separate CitationAgent handles citation attribution at the end.
  • Effort-scaling rules are written into the prompt: simple factual queries get 1 agent + 3-10 tool calls; direct comparisons get 2-4 subagents at 10-15 calls each; complex research gets 10+ subagents with clear divisions of labor. The most common early failure was spawning 50 subagents for a simple question.
  • Tool-description maintainability: they built a "tool-testing agent" that repeatedly exercised flawed MCP tools and rewrote their descriptions, cutting downstream agents' task completion time by 40%.
  • Production operations: errors compound multiplicatively along long traces, so they built checkpoint recovery (no restarting from scratch), rainbow deployments (new and old versions run in parallel with traffic gradually shifted, so running agents aren't interrupted), and high-level decision-pattern tracing that never reads conversation content (a privacy/observability compromise). For more observability practice, see Observability & Tracing.
  • The remaining bottleneck: subagents currently run synchronously — the lead must wait for the whole batch to finish before continuing, with no way to course-correct mid-flight. Making them async is on the roadmap, but it brings new problems of result coordination, state consistency, and error propagation.

Manus: plan-execute-verify multi-agent + a virtual machine ​

Manus (released March 2025 by the Monica team; announced as joining Meta in late 2025 while keeping independent operations) is the most talked-about product in the general agent space. Public materials — including Manus's own context-engineering blog posts and third-party reverse engineering — sketch this architecture:

  • Tasks run in an isolated cloud VM/sandbox (reportedly using E2B as the sandbox infrastructure), with subagents operating the browser, terminal, and filesystem inside that complete computing environment.
  • It uses a Planner / Executor / Verifier division of labor: the Planner decomposes the task into a todo.md-style checklist, the Executor works through it step by step, and the Verifier checks the results.
  • The engineering lessons Manus has shared publicly center on context management: designing context around the KV cache, managing tool visibility by "masking rather than deleting," and treating the filesystem as external memory.

One caveat: Manus has never published a full architecture paper, so these details come from official blog fragments and community reverse engineering, with limited resolution. See the Manus case study.

Magentic-One: Microsoft's generalist multi-agent baseline ​

Magentic-One is Microsoft Research's generalist multi-agent system from November 2024 (arXiv:2411.04468), built on AutoGen and officially integrated with the AutoGen 0.4 release in January 2025:

  • A fixed crew of specialists: one Orchestrator coordinating four specialist agents — WebSurfer (browser operation), FileSurfer (local files), Coder (writing code), ComputerTerminal (executing code).
  • The dual-ledger mechanism: the Orchestrator maintains a Task Ledger (task decomposition and planning) and a Progress Ledger (progress tracking and error recovery), driven by a two-loop structure of an outer loop (updating the task ledger) and an inner loop (tracking progress).
  • Positioned as a generalist baseline on benchmarks like GAIA; it later spawned Magentic-UI with a human-collaboration interface (open-sourced May 2025).

Magentic-One's value isn't the benchmark numbers but the "generalist" division-of-labor idea it demonstrates: specialization not by industry (financial analysts, legal advisors) but by capability modality (browse the web, read files, write code, run code) — which lets the crew migrate across task domains without modification.

6. The Quantified Price in Cost and Reliability ​

Multi-agent isn't free performance; the price is written in the ledger:

  • Token cost: Anthropic's production data — agent interactions consume roughly 4× the tokens of chat, and multi-agent systems roughly 15×. That means multi-agent is only economically justified when "the task's value covers 15× the cost." For a more systematic cost model, see Cost & Latency Optimization.
  • Reliability: errors multiply rather than add in multi-agent systems. In Anthropic's words, "problems that are minor in traditional software can be catastrophic for agents"; one failed step can carry the whole agent swarm onto a completely different trajectory.
  • Debugging cost: non-determinism + multiple executors = you can't localize problems by replaying. Without full trace infrastructure (see Observability & Tracing), a multi-agent system is basically unmaintainable.
  • Evaluation difficulty: the same input can take completely different legitimate paths, so you can't assert "correct steps" — you can only evaluate outcomes. Anthropic does scaled evaluation with a single LLM judge + 0.0-1.0 scoring + pass/fail grading, and stresses starting from a small sample of about 20 real queries. See Evaluation Systems for the methodology.
   Single agent                Multi-agent
   ─────────────────────────────────────────────
   Cost:       4× chat         15× chat            (~4x)
   Failure:    one point       N points + coordination failures (new category)
   Debugging:  replay a single trace  cross-agent trace correlation required
   Evaluation: path + result   results only
   Ceiling:    capped by context  scales with parallelism  ← this is what you buy

7. The Decision Framework: When to Move Beyond a Single Agent ​

Ask yourself five questions in order; any "yes" means stay where you are:

  1. Can a pipeline solve it? Predictable steps → use prompt chaining; you don't even need an agent.
  2. Can a single agent with better context management solve it? Try compressing history, external memory, and better tool descriptions first (remember that 40% figure). Cognition's default answer — the single-threaded linear agent — goes further than most people assume.
  3. Is the task genuinely parallelizable? Do subtasks share writable state? If yes (the same codebase, the same document), merge conflicts will eat the gains from splitting.
  4. Does the task's value support 15× the tokens? Multi-agent on low-value, high-frequency tasks is money-burning.
  5. Can each subtask be written as a self-contained brief? If you can't write it, the split boundary is drawn wrong.

Only after all five pass should you go multi-agent, and start with the smallest form: one supervisor plus a few read-only workers; consider more complex topologies only after that works. For implementation, LangGraph's supervisor/hierarchical support and AutoGen are the two mainstream starting points; to write a supervisor skeleton from scratch yourself, start from Build Your Own Agent and add dispatch logic. For the security dimension, see Security & Alignment — multi-agent expands the prompt-injection attack surface from "one agent" to "every message between agents."

python
# A minimal usable orchestrator-worker skeleton (stdlib only, framework-agnostic)
import asyncio

async def worker(brief: str) -> str:
    """Subagent: independent context, executes a self-contained brief,
    returns only a compressed conclusion.

    In a real implementation this is a full agent loop (see /components/agent-loop),
    with its own system prompt, toolset, and context window.
    """
    return await run_agent_loop(
        system="You are a researcher. Return a conclusion of at most 300 words, with sources.",
        task=brief,  # the brief must be self-contained: goal, boundaries, output format
    )

async def supervisor(query: str) -> str:
    # 1. Plan: split the query into non-overlapping briefs (the emphasis is "non-overlapping")
    briefs = await plan(query)  # e.g. ["key points of company A's 2025 annual report", "key points of company B's report for the same period", ...]

    # 2. Dispatch in parallel: workers don't talk to each other; read-only, separable outputs
    findings = await asyncio.gather(*[worker(b) for b in briefs])

    # 3. Synthesize: the lead agent only touches compressed conclusions, keeping its context clean
    return await synthesize(query, findings)

A multi-agent system isn't a smarter agent; it's an organization. An organization's output ceiling is higher, but its management cost is real — the whole point of this page is to make sure you actually need that "company" before you pay for it.

References ​