Appearance
Claude Code
If you could study only one coding agent to understand "how a modern agent should be built," the answer is Claude Code. It is not the most feature-rich, nor the most deeply IDE-integrated (that's Cursor's territory), but it is the product that has pushed the co-design of "model capability" and the "agent harness" furthest—Anthropic builds both the model and the shell, so each side's iterations can feed the other.
This page draws on three kinds of public information: Anthropic's official releases and documentation, community reverse engineering of the npm package's source maps (Claude Code at one point shipped source maps capable of reconstructing its source code, which produced several high-quality reverse-engineering reports), and a large body of hands-on engineering write-ups. Claims about internals are explicitly labeled "reverse-engineering findings" and kept separate from official information.
1. Product Evolution: From Terminal Toy to Full-Platform Harness
Claude Code's evolution is itself a product lesson—every step expanded the coverage of "the same harness" rather than rewriting the product.
| Date | Milestone | Significance |
|---|---|---|
| 2025-02 | Released as a research preview alongside Claude 3.7 Sonnet, CLI only | Validated "agentic coding in the terminal" as a form factor |
| 2025-05 | GA alongside the Claude 4 family; GitHub Actions integration launched | Went from toy to tool, entering CI/CD pipelines |
| 2025-07 onwards | Subagents (Task tool + .claude/agents/) and custom slash commands rolled out | Moved from a single agent to orchestratable multi-agent |
| 2025-09-29 | Claude Code 2.0 + Sonnet 4.5 released the same day: native VS Code extension, checkpoints, refreshed terminal UI; Claude Code SDK renamed Claude Agent SDK | The "shell" became a reusable, productized asset |
| 2025-10-20 | Claude Code on the web: Code tab on claude.ai + iOS app, agents running in cloud sandboxes | The CLI was no longer the only entry point; the harness went to the cloud |
| 2025-11 | Opus 4.5 released, 80.9% on SWE-bench Verified | The model side kept widening the gap |
| 2026-04 | Desktop app rebuilt: sidebar for managing multiple parallel sessions | From "one agent" to "managing a fleet of agents" |
| 2026-07 | Desktop app gained a built-in browser for preview-and-debug while coding | Moving closer to a "complete development environment" |
A few notable decisions:
The terminal is home, but not the boundary. Product manager Cat Wu put it plainly in a TechCrunch interview: the CLI will always be the home base, but "Claude Code everywhere" is the direction. The web version is not a separate product—it is the same agent logic running in cloud sandboxes, so you can launch and monitor tasks from your phone.
Turning it into an SDK externalizes internal capability. The September 2025 renaming was more than a name change: the Claude Agent SDK exposes Claude Code's own runtime—the same tool catalog, the same permission system, the same loop. This means Anthropic believes the harness itself has standalone product value; see Claude Agent SDK.
Commercially, it is no longer a side project. According to TechCrunch's October 2025 report, Claude Code's annualized revenue exceeded $500 million and its user base grew 10x since GA in May. There are also claims that roughly 90% of Claude Code's code is written by Claude itself—the most aggressive dogfooding imaginable.
2. Architecture: Main Loop + Subagents + Five Extension Surfaces
Claude Code is not open source, but the source map in the npm package let the community reconstruct its TypeScript source. The most substantial reverse-engineering effort covered all ~512,664 lines of code (the awesome-cc-harness project); combined with earlier reports from Reid Barber and ShareAI Lab (v1.0.33), the architecture's outline is now fairly clear:
┌───────────────────────── User Interface Layer ──────────────────────────┐
│ CLI REPL / VS Code extension / Web & Desktop / Agent SDK / GitHub App │
└──────────────────────────────┬──────────────────────────────────────────┘
│
┌──────────────────────────────▼──────────────────────────────────────────┐
│ Main Agent Loop (Query Loop) │
│ while(true): stream model output → parse tool_use → dispatch tools │
│ → feed results back → continue until the model stops calling tools │
│ Reverse finding: the core loop is ~30 lines, wrapped in ~1,800 lines │
│ of error-recovery logic │
└───────┬──────────────┬──────────────┬──────────────┬──────────────────┘
│ │ │ │
┌───────▼──────┐ ┌─────▼────────┐ ┌───▼──────────┐ ┌─▼────────────────┐
│ Tool System │ │ Task │ │ Permission │ │ Context │
│ 40+ atomic │ │ Subagents │ │ Model │ │ Engineering │
│ parallel R / │ │ isolated ctx │ │ defense in │ │ auto-compaction │
│ serial W │ │ returns only │ │ depth + │ │ pipeline │
│ │ │ conclusions │ │ allowlist │ │ 180K→~45K │
└───────┬──────┘ └──────────────┘ └──────────────┘ └──────────────────┘
│
┌───────▼────────────── Extension Surface ────────────────────────────────┐
│ CLAUDE.md memory │ Hooks (events) │ MCP servers │ Skills / commands │
└─────────────────────────────────────────────────────────────────────────┘The Main Agent Loop: Simple Core, Heavy-Duty Recovery
The most counterintuitive finding from the reverse engineering: the core loop is tiny—a while(true) that executes tool calls as the model emits them, feeds results back, and continues until the model stops calling tools. This matches the textbook structure we describe in Agent Loop exactly.
The real engineering mass lives outside the loop: roughly 1,800 lines of error recovery and retry logic (rate limiting, network interruptions, context overflow, resuming interrupted streams), layers of cascading AbortControllers (so pressing Esc cleanly interrupts any stage), and read/write concurrency control under streaming execution. One analysis drew a pointed conclusion from this: only about 1.6% of the codebase is "AI decision logic"—the rest is operational harness: permissions, recovery, compaction, routing. The exact figure may not be rigorous (unverified—double-check before citing), but the direction is right: the moat of a production agent is not the prompt; it is the dirty work.
Task Subagents: Context Isolation Comes First
The main agent spawns subagents through the Task tool. A subagent gets a focused task description and a clean context window, and when it finishes it returns only its conclusions to the main agent. Since July 2025, users can also define custom subagents with Markdown in the .claude/agents/ directory—specifying a dedicated system prompt, restricting available tools, even pinning a specific model.
The essence of this design is not the romantic narrative of "multi-agent collaboration" but context engineering: tasks like exploring a codebase—"reading 50 files just to answer one question"—burn tokens inside the subagent's isolated context, and the main agent's context keeps only the answer. This is the same orchestrator-worker pattern discussed in Multi-Agent Architecture, but the motivation is context economy first, division of labor second.
3. Tool Design Philosophy: General-Purpose Atomic Tools, No Specialized Macro Tools
Claude Code's toolset is the Unix philosophy applied to agents: no high-semantics macro tools like "refactor function" or "fix lints"—just a set of general-purpose atomic tools:
- Files:
Read,Write,Edit(exact old_string/new_string replacement),Glob,Grep - Execution:
Bash(including background tasks),NotebookEdit - Information:
WebFetch,WebSearch - Meta-tools:
Task(spawns subagents),TodoWrite(task list)
Why this is the right call—at least four reasons:
- Atomic tools have an unlimited composition space. Models already know how to write code and shell commands, so
Grep + Read + Editcan compose any code operation. A macro tool, by contrast, needs a new one for every new scenario and can never keep up with the long tail of demand. - Tools double as interface training. Because the model maker builds its own shell, the model can be aligned to these tools' calling conventions during training (how to write Edit's old_string so it doesn't mismatch, how to phrase Grep regexes)—a synergy no third-party shell can replicate.
- Bash is the ultimate escape hatch. Any operation a specialized tool cannot cover—running tests, git operations, builds, starting services—converges on Bash. The harness only needs to police Bash's permissions rather than enumerate every capability.
- Reverse engineering shows tool dispatch itself is deliberate: read-only tools (Read/Grep/Glob) run in parallel, while write operations (Edit/Write/Bash) run serially under a lock to avoid concurrent-write conflicts. This "parallel reads, serial writes" partitioning is an elegant compromise between performance and safety.
A Direct Corollary for Builders
If you are designing agent interfaces for internal tools, first ask: "Can this operation be composed from five or fewer atomic tools?" If yes, don't build a new tool. Tool-count bloat is the most common form of misplaced cleverness in agent systems—every extra tool makes the system prompt longer, raises the odds the model picks the wrong tool, and grows your maintenance surface. See Tools & MCP.
4. Context Engineering in Practice: Compaction, Isolation, External Memory
Claude Code's context strategy boils down to: keep out of context whatever can stay out; throw out whatever got in; and save to disk before throwing anything away. It is the most complete industrial-grade specimen of context engineering.
Auto-Compaction
As a session approaches the context window limit, Claude Code triggers compaction automatically: it first sends a request carrying a "summarize" instruction (that history most likely hits the prompt cache, keeping the cost manageable), then replaces the message history with the generated summary. Before compacting it clears out the oldest tool outputs first, and only summarizes everything if that is not enough. Reverse engineering describes it as a multi-stage pipeline that can compress roughly 180K tokens down to about 45K. Users can also trigger it manually with /compact, attaching a "focus instruction" about what to preserve.
A key detail: after compaction, CLAUDE.md is re-injected from disk. This explains why veterans all say "put rules in CLAUDE.md, don't just say them in conversation"—spoken rules get compacted away; files on disk do not.
Subagent Isolation
As described above, Task subagents are the primary isolation mechanism: reading code, running exploratory tasks, doing research—all "high token burn, low information yield" work happens inside an isolated context, and the main agent receives only distilled conclusions.
File References and External Memory
@path/to/filereferences a file explicitly, putting control in the user's hands instead of letting the model grep blindly.- CLAUDE.md loads in layers: user-level (
~/.claude/CLAUDE.md), project-level (repo root), and subdirectory-level, each taking effect in turn, with@importsupport for referencing other files. It is the vehicle for "cross-session memory"; see Memory Systems. For writing your own agent config file, see Writing AGENTS.md. - Checkpoints, introduced in 2.0, snapshot the workspace before every modification and pair with
/rewindfor rollback—essentially taking "undoability" out of the model's hands and handing it to deterministic filesystem machinery.
An Underrated Design Decision
Claude Code has no built-in vector-retrieval/RAG-style codebase index. It bets on "agentic search": letting the model find code with Grep/Glob the way a human would, backed by a strong enough model and a large enough context window. This is the opposite road from Cursor (which leans heavily on codebase embedding indexes). Practice in 2025–2026 shows that with strong models, agentic search often beats crude embedding retrieval on precision—at the cost of higher token consumption. Related discussion in RAG.
5. Permissions and Trust: Turning Trust into a Configurable State Machine
Letting an agent that can execute arbitrary shell commands loose on your machine raises one core question: trust. Claude Code's answer is a layered permission model (for a fuller discussion see Agent Security).
Permission Modes: Pick a Gear by Task Risk
The classic four modes (cycle through them in-session with Shift+Tab):
| Mode | Behavior | Use Case |
|---|---|---|
default | Read tools allowed; edits/Bash require approval each time | Everyday development |
acceptEdits | File edits auto-approved; Bash still asks | Fewer interruptions during heavy refactoring |
plan | Read-only; the agent can only investigate and propose | Aligning on a plan before touching code |
bypassPermissions | Everything allowed (prominent warning at startup) | Sandboxed/containerized/CI environments |
Versions in 2026 added finer-grained modes such as auto, manual, and dontAsk (the --permission-mode CLI flag lists what is currently supported), but the idea is unchanged: autonomy is granted per mode, not claimed by the model.
Allowlists: Granular Down to Command Patterns
permissions.allow / deny in .claude/settings.json supports command-prefix granularity:
json
{
"permissions": {
"allow": [
"Bash(npm run test:*)",
"Bash(git status)",
"Read(**/*)"
],
"deny": [
"Bash(rm -rf *)",
"Read(./.env)"
]
}
}deny takes precedence over allow. Reverse engineering indicates permission evaluation is layered defense in depth (mode → allowlist/denylist → interactive confirmation → sandbox), each layer intercepting independently so safety never rides on a single point. Settings are layered too: managed (pushed by enterprise IT) → user (~/.claude/) → project (.claude/settings.json, committed with the repo) → local (settings.local.json, kept out of git).
6. Hooks and Extensibility: A Deterministic Layer Outside the Model
CLAUDE.md persuades the model—it will probably comply, but there is no guarantee. Hooks manage the model—they run your shell commands at fixed points in the lifecycle, deterministically, regardless of what the model wants. This is the piece of Claude Code's extensibility design that the industry has copied most.
The Event Model
Hooks attach to agent lifecycle events. The five most used in production:
SessionStart ──→ UserPromptSubmit ──→ [ Agent Loop ]
│
┌── PreToolUse ──→ tool execution ──→ PostToolUse ──┐
│ (can block/rewrite args) (validate/format) │
└──────────── wraps every tool call ────────────────┘
│
Stop (main agent done) / SubagentStop / Notification / PreCompact ...By 2026 the event catalog has grown to 20+ events (subagent start/stop, pre-compaction, notifications, and so on), but mastering a handful covers 80% of use cases.
Configuration and the Verdict Protocol
json
{
"hooks": {
"PreToolUse": [
{
"matcher": "Bash",
"hooks": [
{
"type": "command",
"command": "python3 .claude/hooks/guard.py"
}
]
}
]
}
}A hook script receives JSON on stdin (tool name, arguments, etc.) and expresses its verdict via exit code: exit 0 allows, exit 2 blocks and feeds stderr back to the model (note: exit 1 is merely an error, not a block). Typical uses:
PreToolUseblocksrm -rf, blocks reads of.env, and forces database migrations through approvalPostToolUseruns prettier / eslint / type checks after every Edit and feeds the errors back to the modelStopenforces "no finishing while tests fail," forcing the agent to keep fixing
The Full Extension Surface
Beyond hooks, Claude Code has four more extension points; together they amount to nearly a platform:
- MCP: the standard protocol for external tools and data sources; see Tools & MCP
- Custom slash commands: drop a Markdown file into
.claude/commands/to register a new command—essentially a parameterized prompt template - Skills: capability packs loaded on demand, more structured than commands
- Claude Agent SDK: embed the entire harness in your own program (Python
claude-agent-sdk/ TS@anthropic-ai/claude-agent-sdk); CI, chat bots, and internal platforms can all grow a Claude Code core
python
# Minimal Claude Agent SDK (Python) example
import anyio
from claude_agent_sdk import query, ClaudeAgentOptions
async def main():
options = ClaudeAgentOptions(
allowed_tools=["Read", "Grep", "Bash"],
permission_mode="acceptEdits",
cwd="/path/to/repo",
)
async for message in query(
prompt="Find all unresolved TODOs in this repo and group them by module",
options=options,
):
print(message)
anyio.run(main)The Trust Boundary of Hooks
Hooks in a project-level .claude/settings.json are "executable commands distributed with the repo"—clone a malicious repo and trust it, and you have executed the author's arbitrary script. This is the same class of RCE surface as git hooks and direnv. Reviewing an unfamiliar repo's .claude/ directory should become muscle memory, just like checking postinstall in package.json.
7. Why It Became the Benchmark: Success Factors
The coding-agent space is crowded (Cursor, Copilot, Windsurf, Aider, Codex CLI...). That Claude Code ended up as the benchmark, I would argue, is four factors multiplying together—remove any one and it does not happen:
1. Model and harness co-designed under one roof. This is the deepest moat. When Anthropic trains Claude, it can directly optimize "performance inside the Claude Code shell"—tool-calling habits, endurance on long tasks, recovery after interruption—while shell-side iteration (system prompt tweaks, tool definition changes, new compaction strategies) immediately feeds back into benchmark runs. Third-party shells can only wait for the model maker's next release. Sonnet 4.5 hit 77.2% on SWE-bench Verified at launch, and Opus 4.5 reached 80.9%—all numbers measured with model and harness together.
2. Disciplined form-factor choices. Starting from the terminal was a mocked decision (in early 2025 everyone was racing to build IDE plugins), but the terminal means: proximity to the real development environment (tests, git, and builds all live there), natural scriptability, and zero-cost entry into CI. Only after the capabilities proved out did it grow an IDE extension, web, and desktop—each step a new shell over the same harness, no rewrite.
3. Layering deterministic mechanisms beneath model capability. CLAUDE.md handles "soft rules," hooks handle "hard rules," checkpoints handle "undoability," permission modes handle "autonomy"—everything that "shouldn't be trusted to the model" was given a deterministic mechanism. This lets the agent be both empowered and controllable, and it is the key to enterprise adoption.
4. Extreme dogfooding. Anthropic claims about 90% of Claude Code's code is written by Claude. That flywheel—build your product with your product, so problems surface and get fixed the same day—is an iteration speed no outside competitor can replicate.
Honesty requires stating the limits too: its strengths are bound to Anthropic models; the subscription tiers (Pro $20/month, Max $100–200/month) are not cheap for heavy users; agentic search's token overhead is considerable on very large monorepos; and the maintainability debate around "90% of code written by the agent" has never gone away. Being the benchmark does not mean having no weak spots.
8. A Reusable Checklist for Building Your Own Agent
If you are building your own agent from scratch, these are the Claude Code designs worth stealing one by one, ordered by return on effort:
- Keep the loop minimal; invest the engineering in recovery. The core loop fits in thirty lines; spend your effort on rate-limit retries, interruption recovery, and timeout fallbacks. What kills agents in production is never "not smart enough"—it is "crashed and never came back."
- Atomic tools plus one escape hatch. Three classes of atomic tools—read, write, execute—are enough to start; Bash (or an equivalent) is the escape hatch, and policing its permissions covers 80% of your risk.
- Treat context as a budget. Auto-compaction as the safety net + subagent isolation to save tokens + on-disk files as durable memory—none of the three is optional. Rules relayed verbally get compacted away; put them in files.
- Soft rules in prompts, hard rules in code. "Try to write tests" belongs in the system prompt; "you may not stop while tests fail" belongs in a Stop hook. Any constraint where "the model occasionally ignores it" is unacceptable must be enforced deterministically.
- Make permissions modes, not pop-ups. The endgame of per-action confirmations is users mashing Enter. Tier by risk (read-only / editable / fully automatic) and let users pick the right gear once, up front.
- Undoability matters more than trustworthiness. Rather than praying the model never deletes the wrong file, snapshot before every change (checkpoints). Rollback capability is what turns
bypassPermissionsfrom a gamble into an engineering practice. - Separate read and write scheduling. Read-only tools in parallel, write tools in series—one rule, a meaningful latency win.
- Test the harness and the model together. When you evaluate your agent you are testing a "model × shell" combination; any change on either side needs a full eval regression run. See Agent Evaluation and Observability.
References
- Enabling Claude Code to work more autonomously — Anthropic — Official release of 2025-09-29: Claude Code 2.0, the VS Code extension, checkpoints, and the Agent SDK renaming.
- Anthropic brings Claude Code to the web — TechCrunch — Web launch, pricing tiers, revenue and user-growth numbers, and the Cat Wu interview.
- anthropics/claude-code Releases — GitHub — The official changelog for tracking features and fixes per release.
- awesome-cc-harness — GitHub — A systematic reverse engineering of all 512K lines of Claude Code's TypeScript source (agent loop, tool system, permission model, compaction pipeline); the source of many of this page's architecture findings.
- Reverse engineering Claude Code — Reid Barber — The earliest source-map reverse engineering, from early 2025, reconstructing the REPL structure and tool definitions.
- Inside Claude Code: A Deep-Dive Reverse Engineering Report — ShareAI Lab — A 50K+ line deobfuscation analysis targeting v1.0.33.
- Claude Code Hooks official documentation — The authoritative reference for hook events, matchers, and the JSON output protocol.
- claude_code_docs_map.md — Simon Willison — Observations on Claude Code's system prompt and docs-index design, drawn from its own self-querying behavior.