Skip to content

Subagents & Multi-Agent Orchestration

At a glance What subagents really are — context isolation and parallelization, not "multiple AIs holding a meeting." Claude Code's delegate-and-summarize mechanism, the orchestrator-worker architecture and hard numbers behind Anthropic's multi-agent research system, Cognition's counter-argument in "Don't Build Multi-Agents," and a decision framework for choosing between a strong single-threaded agent and a multi-agent setup.

Subagents & Multi-Agent Orchestration ​

Subagents answer this question: when a task exceeds what a single context window can carry, or contains several independent directions that could advance simultaneously, how does the harness hand work off — without drowning in what it handed off?

The core thesis of this page up front: a subagent is, at bottom, context isolation plus parallelization — a space-for-time move in context engineering — not "multiple AIs holding a meeting." Plenty of multi-agent frameworks market themselves as "simulating a team": a PM agent, a programmer agent, and a QA agent sitting around a table. That metaphor of human organization obscures the actual engineering variables. Only three things decide whether a multi-agent system succeeds: what's in each agent's context, what gets exchanged between them, and who answers for conflicting conclusions.

Why subagents exist: three real motivations ​

Strip away the "AI team" narrative and the rationale for subagents reduces to three — all pointing at the same scarce resource: the main context window.

Motivation 1: Protect the main context ​

The simplest and most common reason. A research-style task — "figure out how this repo's auth module works" — may require reading dozens of files and running a dozen searches, generating tens of thousands of tokens of raw material. Stuff it all into the main conversation, and by the time the main agent is ready to do real work, the context is already full of intermediate debris — triggering compaction and the loss of early details (see the dilemma in context engineering).

A subagent quarantines that pile of intermediate material inside its own context window: it reads everything, then returns only a few hundred tokens of conclusions to the main agent. What the main context sees is not the expedition but the expedition report. This is the same idea as compression in memory systems — trading a summary for space — except the compression happens at the agent boundary rather than inside the session.

Motivation 2: Divide the task, separate the concerns ​

A subagent can be given its own system prompt, its own subset of tools, even its own (cheaper or more specialized) model. That lets roles with fixed identities and clear boundaries — code review, test writing, log analysis — be packaged into reusable expert units. Note: the value here isn't "role-playing," it's that each unit carries a different context configuration and permission surface — a review agent can be given read-only tools only, which is at the same time a permissions decision.

Motivation 3: Parallel exploration ​

A single-threaded agent explores serially: it can't investigate direction B before finishing direction A. For breadth-first tasks ("find every tech company board member in the S&P 500"), serial execution means total time equals the sum of all directions, and the early choice of direction path-dependently pollutes later judgment. Multiple subagents fan out in parallel, each exploring with its own isolated context, and total time approaches that of the slowest direction. Anthropic's measurements (detailed below) show the payoff for this class of task is overwhelming.

text
┌───────────────────────── Main agent context ──────────────────────────┐
│  User task                                                            │
│  Description of subtask A ──────────────┐                             │
│  Description of subtask B ──────────┐   │  ← inbound: task            │
│                                     ▼   ▼    descriptions only        │
│                              ┌────────────┐ ┌────────────┐             │
│                              │ Subagent A │ │ Subagent B │  ← each     │
│                              │ own window │ │ own window │    with an  │
│                              │ reads 20   │ │ runs 15    │    isolated │
│                              │ files…     │ │ searches…  │    context; │
│                              └─────┬──────┘ └─────┬──────┘    parallel  │
│  Summary A ◄───────────────────────┘              │                    │
│  Summary B ◄──────────────────────────────────────┘  ← outbound:       │
│                                                         summaries only │
│                                                                       │
│  Tens of thousands of tokens of intermediate output stay in the       │
│  subagents' contexts and are discarded when they finish               │
└───────────────────────────────────────────────────────────────────────┘

Read this diagram and the entire economics of subagents becomes clear: the main context is the scarce asset; subagents are consumables. What you delegate is intent; what you get back is conclusions; the process is burned on the spot.

Claude Code's subagent mechanism: delegate, get a summary back ​

Claude Code's subagents are the cleanest engineering implementation of this idea, worth dissecting layer by layer (mechanism per its official docs; for the full product analysis see Case Study: Claude Code).

An isolated context window ​

Each subagent runs in a context window completely isolated from the main conversation. It inherits nothing from the main thread — none of the user's chatter, none of the earlier tool output, no todo list. All it sees is: its own system prompt + the task description the main agent wrote for it. That's how thorough the isolation is — and also where its cost lies (Cognition's critique, below, hits exactly this point).

Configuration as Markdown ​

A subagent isn't code; it's a Markdown file with YAML frontmatter, placed in the project's .claude/agents/ directory:

yaml
---
name: code-reviewer
description: Senior code review expert. Use proactively after completing a code change.
tools: Read, Grep, Glob, Bash
---

You are a senior code reviewer. When given a code change:
1. Check correctness, security vulnerabilities, and adherence to project conventions;
2. Organize feedback by severity;
3. Do not modify code — output review comments only.

Three design details worth noting:

  • description is the routing key. Whether the main agent delegates a task to a given subagent depends on matching the task against this field. A vague description means messy delegation — the same law as "the tool description determines tool usage" from the tool system.
  • The tools field is the permission boundary. A subagent can be restricted to a read-only toolset, which blocks file modification at the harness level. An isolated context plus a narrowed tool surface makes the subagent a natural permission sandbox.
  • The body is a complete system prompt. The subagent's behavior is shaped entirely by this prompt; it doesn't know — and doesn't need to know — what happened in the main conversation.

The delegate-and-summarize protocol ​

The main agent initiates delegation through a Task tool: it passes in the subagent type and a self-contained task description ("self-contained" is the key — the subagent can't see the main conversation, so anything not written in the description doesn't exist for it). After the subagent runs its own complete agent loop, only its final message returns to the main context; the dozens of intermediate tool calls all stay in the subagent's window.

This protocol cuts both ways:

  • Benefit: for an exploration that might burn tens of thousands of tokens, the main context pays only the cost of "task description + summary of conclusions."
  • Cost: the main agent has zero visibility into the subagent's execution. If the summary is wrong, incomplete, or overly optimistic, the main agent has no raw material to double-check against — it must trust a second-hand report. This is the shared soft spot of every delegation-based architecture.

Anthropic's multi-agent research system: the affirmative evidence for orchestrator-worker ​

Claude Code's subagents solve "context protection." Anthropic's June 2025 engineering post, How we built our multi-agent research system, shows what subagents look like scaled up for "parallel exploration" — the architecture behind Claude's Research feature.

The architecture: lead agent + parallel subagents ​

It's a textbook orchestrator-worker structure (corresponding to the orchestrator-workers pattern Anthropic defined in Building effective agents):

text
      User query: "Find all the notable companies in field X"
                        │
                        ▼
              ┌───────────────────┐
              │   LeadResearcher   │  ← sets research strategy, writes
              │   (orchestrator)   │    the plan to Memory so context
              └─────────┬─────────┘    truncation can't lose it
        ┌───────────────┼───────────────┐
        ▼               ▼               ▼
  ┌───────────┐   ┌───────────┐   ┌───────────┐
  │ Subagent 1│   │ Subagent 2│   │ Subagent 3│  ← run in parallel, each
  │ direction │   │ direction │   │ direction │    with its own context
  │ A search  │   │ B search  │   │ C search  │    window
  └─────┬─────┘   └─────┬─────┘   └─────┬─────┘
        └───────────────┼───────────────┘
                        ▼
              LeadResearcher synthesizes findings,
              decides whether more research is needed
              (can spawn more subagents)
                        │
                        ▼
              CitationAgent verifies sources → final report

The post explains why this architecture suits research: search is fundamentally compression — distilling insight from a sea of material. Subagents compress different slices of the corpus in parallel and return the most important tokens, condensed, to the lead agent. Independent contexts also bring separation of concerns (different tools, prompts, exploration trajectories), which reduces path dependence.

The numbers: huge gains, huge burn ​

What makes this post valuable is that it publishes hard numbers rarely seen elsewhere:

  • Performance: a multi-agent system with Claude Opus 4 as lead and Claude Sonnet 4 as subagents beat single-agent Opus 4 by 90.2% on internal research evals. Breadth-first queries (requiring several independent directions to be pursued at once) benefited most — e.g., "find the board members of every company in the S&P 500 IT sector": serial single-agent search was slow and leaky, while the multi-agent decomposition found the correct answer.
  • Where the performance comes from: in an analysis of the BrowseComp eval (which measures a browsing agent's ability to locate hard-to-find information), three factors explained 95% of performance variance, and token usage alone explained 80%. The essential function of the multi-agent architecture thus stands revealed: it's a token-scaling machine — spreading inference budget across multiple parallel contexts, breaking past a single agent's window and serial limits.
  • Cost: an ordinary agent uses about 4× the tokens of a chat interaction; a multi-agent system about 15×. The conclusion is blunt: multi-agent only makes economic sense when the task's value justifies paying that premium for performance.

The scars: coordination complexity explodes ​

The post is equally honest about the lessons between prototype and production, each directly useful to harness designers:

  • Delegation must be specific. Early on, the lead agent would issue vague instructions like "research the semiconductor shortage" — one subagent went off studying the 2021 automotive chip crisis while two others redundantly investigated the 2025 supply chain, overlapping in places and leaving gaps elsewhere. The fix: require every subtask to carry an objective, an output format, tool guidance, and clear boundaries.
  • Calibrate effort. Early agents would spin up 50 subagents for a simple query. The fix: write explicit effort-scaling rules into the prompt — a simple factual query gets 1 agent and 3–10 tool calls; a direct comparison gets 2–4 subagents; only complex research gets 10+.
  • Parallelism is free performance. Introducing two levels of parallelism (the lead running 3–5 subagents simultaneously, and each subagent calling 3+ tools in parallel) cut research time on complex queries by up to 90%.
  • Synchronous execution is the bottleneck. In the implementation at the time, the lead agent synchronously waited for each batch of subagents; it couldn't course-correct mid-flight and subagents couldn't coordinate with each other, so the whole system was blocked by its slowest member. Async execution is the obvious direction — but it introduces new hard problems in result reconciliation, state consistency, and error propagation.
  • Old production-engineering problems, amplified. Long-running stateful agents need checkpoint-resume rather than restart-from-scratch; nondeterministic behavior is only debuggable with end-to-end tracing (echoing observability); deploying updates can't interrupt running agents — use rainbow deployments to shift traffic gradually.

The counter-argument: Cognition's "Don't Build Multi-Agents" ​

The very same week Anthropic's post came out (June 2025), Cognition — the company behind Devin — published a pointed rebuttal: Don't Build Multi-Agents by Walden Yan. Given that Cognition builds a coding agent — precisely the domain where multi-agent is hardest — the piece deserves to be taken seriously.

Two principles ​

Starting from a reliability theory of long-running agents, the post proposes two principles of context engineering and argues for defaulting away from any architecture that violates them:

Principle 1: Share context — and share the full agent trace, not just messages. The classic counter-example: "build a Flappy Bird clone" gets split into two subtasks — subagent 1 builds "a scrolling background," subagent 2 builds "a bird that moves up and down." Subagent 1's background comes out looking like Super Mario; subagent 2's bird looks nothing like a game asset. The agent responsible for merging them faces two mutually misunderstood artifacts and has nowhere to start. You might think handing the original task verbatim to each subagent solves this — but in real systems, the meaning of a task gets clarified progressively through rounds of conversation and tool calls. A static task description handed to a subagent can't carry the full semantics accumulated in the main conversation.

Principle 2: Actions carry implicit decisions, and conflicting decisions produce bad outcomes. Even when two subagents understand the task correctly, each makes a host of unspecified implicit choices while working — this time, maybe the bird's and the background's visual styles simply don't match. Two locally correct agents produce one globally wrong system.

The conclusion: single-threaded first ​

Cognition's default architecture is the single-threaded linear agent: one continuous context, where every decision is visible to later steps. Context overflow is a real constraint, but the answer is compression and memory techniques (maintaining continuity across a long task), not shredding the work and distributing it to executors who can't see each other.

The two principles don't inherently conflict with subagent mechanisms

Look closely and you'll see Cognition's target is the architecture of "parallel workers producing artifacts that get merged afterwards" — not delegation in all forms. A subagent that only reads, explores, and returns a summary (like Claude Code's "go look at how this module works") produces no artifacts that need merging, so conflicting implicit decisions never arise. What's genuinely dangerous is multiple subagents writing to the same system in parallel — code, design, configuration. To check whether a multi-agent design steps on the mine, verify exactly one thing: do the subagents' outputs need to be consistent with each other?

Where each applies: the debate isn't actually a contradiction ​

Read the two posts side by side and you'll notice each quietly limits its scope — and the scopes barely overlap:

DimensionAnthropic (multi-agent research system)Cognition (Don't Build Multi-Agents)
Task typeOpen-ended research, breadth-first information gatheringProgramming — building one coherent artifact
Subagent outputIndependent information summaries; aggregation completes the taskComponents that must assemble into one system
ParallelismNaturally high: directions barely depend on each otherNaturally low: components are coupled at their interfaces
Context-sharing needsLow: directions don't need each other's findingsHigh: every decision shapes how other parts are interpreted
Verdict+90.2% performance, worth the 15× tokensViolates the two principles; default to not doing it

Anthropic itself acknowledges this boundary, in plain words: "domains requiring all agents to share the same context — or with many dependencies between agents — don't fit multi-agent systems today. In most coding tasks the truly parallelizable share is far smaller than in research, and LLM agents still aren't good at coordinating and delegating with other agents in real time."

So the right reading isn't "Anthropic says multi-agent good, Cognition says multi-agent bad." It's one unified principle:

Value of subagents = gains from parallel exploration + gains from context isolation − costs of coordination and context severing.

Research tasks score high on the first two and low on the third; coding tasks are exactly the reverse. This also explains why Anthropic uses multi-agent heavily in its Research product while deliberately shaping Claude Code's subagents into the conservative form of "isolated exploration + summary returned" — the same company, making different calls for different task structures.

A decision framework: strong single-threaded agent or multi-agent? ​

If you're designing a harness, walk a concrete task through these four questions in order:

1. Can the task decompose into loosely coupled subtasks? Do the subtasks' outputs need to be consistent with each other? Yes (the same codebase, the same design) → lean single-threaded. No (independent fact-finding, independent areas of the codebase) → keep going.

2. How much context sharing do the subtasks need? Does finishing subtask B require knowing subtask A's execution details (not just its conclusions)? Yes → single-threaded; the cost of the severed context will eat the parallel gains.

3. Does it fit in a single context window? If yes and the task isn't breadth-first → single-threaded; don't import needless complexity (consistent with Anthropic's "start with the simplest thing that could work" principle in Building effective agents). If no → consider compression and memory first; consider delegation only after that fails.

4. Can the task's value pay the token premium? Multi-agent's ~15×-of-chat token cost is Anthropic's measured number, not a theoretical estimate. Using multi-agent orchestration for everyday chores is, most likely, paying the electricity bill for an org chart.

text
              ┌─────────────────────────────┐
              │    A complex task arrives    │
              └─────────────┬───────────────┘
                            ▼
              Do the subtasks' outputs need to be
              mutually consistent? (writing one codebase?)
                   │                    │
                 yes│                    │no
                   ▼                    ▼
            ┌────────────┐   Does it fit in one window?
            │ Single-     │      │           │
            │ threaded    │    yes│           │no
            │ strong agent│       ▼           ▼
            │ + todo /    │  Single-    Read-only exploration,
            │ memory /    │  threaded   summaries back?
            │ compression │   │              │         │
            └────────────┘    │            yes│         │no (writing one system)
                              │               ▼         ▼
                              │        Delegated      High-risk zone:
                              │        subagents      the parallel-write
                              │        (Claude Code   architecture
                              │        style, low     Cognition warns
                              │        parallelism)   against
                              │               │
                              │   Breadth-first + high value?
                              │               ▼
                              │   orchestrator-worker
                              │   (Research style, massive parallelism)

A rule of thumb

Read-only delegation is safe; parallel writes are dangerous. Subagents that "look," "search," "try," "evaluate," then come back with a summary are almost always net positive. Subagents each "writing" part of the same artifact must first answer Cognition's two principles — and if you can't, fall back to single-threaded. This shares roots with the rule of thumb from planning: only when decomposition outgrows what a todo list can carry do subagents get their turn.

The seams with the rest of the harness ​

Subagents aren't an isolated component; they connect to other harness parts at several key seams that need designing together:

  • With the Agent Loop: each subagent runs a complete but usually more constrained loop (fewer tools, clearer stop conditions). The orchestrating layer's loop gains one new kind of action: delegate and wait.
  • With context engineering: subagents are context management's "spend total tokens to buy quality" move — burn more tokens overall to protect the main context's attention budget. How the task description is written determines whether isolation protects or severs.
  • With permissions: narrowing a subagent's tool surface is the cheapest permission control there is. An exploration subagent with read-only tools only has a naturally lower ceiling on the damage it can do than the main agent.
  • With memory systems: in Anthropic's Research system, the lead agent writes the research plan into Memory so context truncation can't lose it — on long tasks, continuity across subagent rounds still depends on externalized memory; subagents don't solve that.
  • With observability: debugging difficulty explodes with agent count — "the user reports the agent missed obvious information, and we can't see why" was Anthropic's own lived experience. Without end-to-end tracing, a multi-agent system is unmaintainable.

Further reading ​

References ​