Skip to content

Case Study: Claude Code

At a glance A deep dive into Claude Code's harness design — the minimal Unix-style toolset, the planning mechanisms behind TodoWrite and plan mode, CLAUDE.md memory files, subagents, hooks, MCP, and permission modes, plus why it chose the terminal over the IDE. The site's component framework, assembled in full inside a real product.

The Claude Code Case Study: A Harness of Minimal Thickness ​

Earlier, Anatomy of the Harness took an agent system apart into its components: the loop, context, tools, planning, memory, subagents, permissions. Now let's look at a real product that ships all of those components, each in its simplest workable form — Claude Code.

There are three reasons it serves as this site's flagship case study. First, it's a rare example of the model and the harness vertically integrated under one roof: Anthropic both trains the model and builds the system that wraps it, and openly admits that the product's edge lives in the latter — the Claude Agent SDK is officially positioned as "the same tools, agent loop, and context management that power Claude Code" (see What Is an Agent Harness?). Second, its design philosophy is radically restrained. Creator Boris Cherny put it bluntly in his Latent Space interview: "Claude Code is less a product than a Unix utility," and the team's product principle is "do the simple thing first." Third, its internals are unusually well documented in public: official docs, engineering blog posts, and a reverse-engineering analysis of its system prompt and tool definitions let us discuss the harness on evidence rather than guesswork.

The argument of this article: Claude Code's strength lies not in any single clever mechanism, but in cutting every component down to "just enough" and handing all the complexity it saved to the model. This is the most important lesson for understanding modern agent harness design.

Overall Architecture: Claude Code's Harness in One Diagram ​

text
┌────────────────────── Claude Code (Harness) ──────────────────────┐
│                                                                   │
│  user input ──> system prompt + CLAUDE.md memory +                │
│                 todo list + history                               │
│                          │                                        │
│                          ▼                                        │
│                 ┌─────────────────┐  Task tool   ┌─────────────┐  │
│                 │    Agent Loop   │ ────────────>│  Subagents  │  │
│                 │think→act→observe│ <── summary ─│(own context)│  │
│                 └────────┬────────┘              └─────────────┘  │
│                          │                                        │
│        ┌─────────────────┼─────────────────┐                      │
│        ▼                 ▼                 ▼                      │
│   Bash/Read/Edit     Glob/Grep        TodoWrite                   │
│   Write/WebFetch     (retrieval tools)  ExitPlanMode/Task         │
│  ──────┴──── permission modes + hooks cross-cut every tool call ──│
│                          │                                        │
│    the real environment in the terminal (filesystem/shell/git)    │
└───────────────────────────────────────────────────────────────────┘

Cross-check it item by item against the anatomy diagram: loop, context, tools, planning, memory, subagents, permissions — nothing missing, and each one has a concrete implementation you can take apart. Let's go through them one at a time.

The Toolset: Unix Tools Make the Best Agent Tools ​

Claude Code ships only a dozen or so built-in tools, and nearly all of them restate the Unix philosophy (the tool list and definitions are covered in the reverse-engineering analysis; the permission rules section of the official settings docs also lists the tool names):

ToolUnix mental modelResponsibility
Bashthe shell itselfRun arbitrary commands: builds, tests, git, package management
Read / Write / Editcat / redirection / sedRead and write files, make precise replacements
Glob / Grepfind / grepSearch the codebase by filename and content
TodoWritenone (a pure context tool)Maintain the task list — see next section
Taskxargs-style fan-outSpawn subagents
WebFetch / WebSearchcurl / search enginesFetch external information

Note what this list doesn't contain: no read_function_definition, no find_symbol_by_name, no semantic index, no LSP-style structured code tools. A 2023-style "AI IDE" would build a pile of bespoke tools for code understanding; Claude Code hands the model grep. Why is this actually the right call?

First, the model's pretraining corpus is full of Unix. The debugging and editing workflows that programmers worldwide have demonstrated on Stack Overflow, GitHub, and in documentation for two decades overwhelmingly exist as grep/bash/git. With these tools, the model is working in its native tongue; every novel tool you invent asks it to learn a dialect nobody but your product speaks.

Second, general-purpose tools fail more gracefully than specialized ones. A specialized tool encodes "how to use it" into the interface, and the model dies the moment it gets one field wrong. grep has no fields to get wrong — if a search comes up empty, try another pattern. The retry cost is near zero, which exactly matches the trial-and-error rhythm of the agent loop.

Third, composability. The value of Unix tools isn't any single tool but their pipe-style composition. Bash alone turns an entire ecosystem (gh, aws, npm, test frameworks, linters) into the agent's capability set, without the harness author writing a single line of adapter code. The official best practices go so far as to explicitly recommend that users "have Claude interact with external services through CLI tools like gh."

Anthropic's own tool-design methodology (Writing tools for agents) distills this into a few principles: don't wrap APIs one by one; build a small number of well-thought-out tools for high-value workflows; keep tool return values high in signal and low in noise; polish tool descriptions as carefully as you would a prompt. Claude Code's built-in toolset is what you get when those principles are applied to the coding domain — a deliberately small number of tools, clean boundaries, plain-text returns. More in Tools & MCP.

A commonly misunderstood point

A "minimal toolset" does not mean "a lazy harness." Quite the opposite: precisely because the tools are raw, correctness of usage, output truncation, and context management for oversized results all fall on the harness and the prompt to handle. The reverse-engineering analysis shows that Claude Code's system prompt carries extensive discipline clauses — "read before you edit," "Edit requires an exact match," "prefer dedicated tools over Bash." The more general the tools, the heavier the prompt engineering. Complexity doesn't disappear; it migrates from the interface layer to the context layer.

Planning: The Two-Layer Structure of TodoWrite and Plan Mode ​

Claude Code's planning mechanism is the main subject of Planning & Task Decomposition; here I'll just sketch what it looks like inside the product.

TodoWrite is the reference implementation of interleaved planning: a three-state task list (pending / in_progress / completed) lives in the context, and the model is asked to maintain it constantly — mark things done the moment they finish, keep exactly one item in_progress at a time, add a new task when blocked rather than force-marking completion. It has no runtime logic whatsoever; the mechanism is pure cognitive offloading: the list stays visible, pushing back against goal drift over long tasks.

Plan mode is a human-in-the-loop gate layered on top of the loop: Shift+Tab or claude --permission-mode plan enters a read-only mode where the model can investigate but not modify anything, and finally submits its plan for user approval through the ExitPlanMode tool; execution begins only once approved. The official best practices prescribe a four-stage workflow: explore → plan → implement → commit.

Why you want both mechanisms

TodoWrite solves "the model forgetting things on its own" (externalizing state within a session); plan mode solves "show a human before you act" (an approval point before irreversible operations). The first targets the model's cognitive limits, the second calibrates human trust. Systems that keep only one — say, Cursor's early planning feature, which was plan mode only, or open-source clones with only a todo list — each cover only half the problem.

Memory: The Minimalism of CLAUDE.md ​

Memory systems usually get designed to look complicated: vector stores, retrieval pipelines, summarization, forgetting curves. Claude Code's answer is one Markdown file, read into context at the start of every session. That's it. In Boris Cherny's words: the team studied all kinds of memory architectures and external products and ended up deciding to "ship the simplest thing — a file with some stuff in it that gets automatically read into context" (Latent Space interview).

But this "simplest form" contains several deliberate layers (see the official memory docs):

  • Layered scopes: organization level (managed policy) → user level (~/.claude/CLAUDE.md) → project level (./CLAUDE.md, checked into version control and shared with the team) → local level (CLAUDE.local.md, gitignored). The lookup walks upward from the working directory, level by level, and the results are concatenated into context; instructions closer to where you are appear later.
  • Lazy loading: a CLAUDE.md in a subdirectory isn't loaded at startup; it enters the context only when the model actually reads a file in that directory — a plain form of context engineering: don't pay tokens for instructions you aren't using.
  • The @path import syntax: you can link in a README, package.json, or other documents, up to four levels of recursion.
  • Explicit self-awareness: the official docs repeatedly stress that CLAUDE.md is context, not enforced configuration — it's injected as a user message and the model "tries to follow it," with no guarantee. If you want enforcement, use hooks (next section).

The engineering lesson here is that it redefines "memory" as "convention": CLAUDE.md lives in no storage other than git, needs no index, can be edited and reviewed directly by humans, and evolves alongside the codebase. It also spawned a cross-product convention (AGENTS.md and its kin); the official docs even teach you how to use an @AGENTS.md import so multiple agent tools can share one set of instructions. For a comparison with heavier memory architectures, see Memory Systems.

Subagents: Context Isolation Is the Only Goal ​

Claude Code's Task tool can delegate a subtask to a subagent: the subagent runs in its own context window and, when done, reports only its conclusions back to the main session. You can also define custom subagents as Markdown files with frontmatter under .claude/agents/ — specifying the system prompt, the subset of available tools, even which model to run on (official subagents docs).

The most important design judgment here: a subagent's first value is not role-playing but context budget management. The official best practices say it plainly: researching a codebase means reading lots of files, and the contents eat the main session's context; let the subagent do the reading, and the main session receives a single page of conclusions. Compared with the "expert persona" story ("dispatch a security auditor"), this is a plainer and more honest motive. For the full discussion of multi-agent collaboration, see Subagents & Multi-Agent Orchestration.

Hooks: Turning "Suggestions" into "Guarantees" ​

CLAUDE.md persuades; hooks enforce. This is a pair of mechanisms the official docs deliberately set in opposition.

Hooks are shell commands you configure in settings, attached to agent lifecycle events: PreToolUse (before a tool call, can block it), PostToolUse (after a tool call), UserPromptSubmit, Stop, SessionStart, and others (the full event model is in the official hooks docs). The key semantics: a PreToolUse hook that exits with code 2 blocks the tool call and feeds stderr back to the model.

The division of labor in the official best practices is stated with precision: "Instructions in CLAUDE.md are advisory; hooks are deterministic — for the things that must happen every time, zero exceptions." Must run eslint after every edit? Write a hook. Writing to the migrations directory is forbidden? Write a hook. Don't count on the model to remember.

From a harness perspective, hooks matter because they open up a programmable interface for the user that doesn't depend on the model's cooperation: the model is stochastic, the hook is deterministic, and only the combination wraps agent behavior inside an enforceable boundary. It also exposes an honest architectural fact — constraining model behavior by prompts alone is not enough in serious engineering settings.

Permission Modes: A Spectrum of Autonomy ​

By default, every tool call with side effects (writing files, running bash, calling MCP) pops up for user approval — safe, but by the tenth click approvals become a rubber stamp. Claude Code turns this problem into an explicit spectrum (details in the official settings docs and the best-practices post):

ModeBehaviorWhen to use
defaultFirst use of each action type requires approval; can whitelist item by itemEveryday interactive use
acceptEditsFile edits are auto-approved; bash and the like still require approvalTrust edits, stay wary of commands
planRead-only; all modifications forbiddenResearch and plan approval
autoA separate classifier model reviews every command, blocking only the high-risk onesYou trust the direction and don't want to click through everything
bypassPermissionsEverything allowed (i.e. --dangerously-skip-permissions)Unattended batch runs inside a sandbox/container

Two design choices worth noting. First, allowlists can go down to the command prefix (allow npm run lint but not arbitrary bash) — this granularity directly determines whether the system is usable in production. Second, auto mode uses "another model as the reviewer" rather than a rules engine, an admission that judging dangerousness itself takes semantic understanding. There's also OS-level sandboxing to physically restrict file and network access. There is no "correct point" on this spectrum, only the point that matches the task's risk — for the full framework, see Permissions, Safety & Human-in-the-Loop.

MCP: The Tool System's Open Boundary ​

However general the built-in tools are, they can't cover the whole world. Claude Code's open interface is MCP (Model Context Protocol, the protocol Anthropic open-sourced in November 2024; spec at modelcontextprotocol.io): one claude mcp add line connects an MCP server, and tools from Notion, Figma, databases, or internal services join the model's tool list.

Worth noting is the warning Anthropic gives in its tool-design post: MCP makes "hundreds of tools" possible, but more tools is not better — wrapping API endpoints one by one is the most common anti-pattern; tools should be merged by workflow (schedule_event rather than list_users + list_events + create_event). In other words, MCP solves the "can you connect" problem; whether the connection is any good remains a harness engineering problem. The related trade-offs are explored in Tools & MCP.

Why a Terminal, Not an IDE ​

This is Claude Code's most counterintuitive choice — and the one hindsight has most vindicated. In 2024–2025 the prevailing view was that AI coding products should live inside the IDE (the triumph of the Cursor route seemed to prove it). Claude Code went and built a CLI instead.

In the Latent Space interview, Boris Cherny gave two layers of reasons. The historical layer: it began as a small terminal experiment for him alone to get familiar with the API, and the "do the simple thing first" culture kept it in its simplest form. The strategic layer is more interesting — they predicted models would get better fast, so rather than hand-crafting a thick UI/scaffold for today's models, they stayed bare bones and gave users raw access to the model, betting on model progress. He also described a three-layer mental model: whatever capability can be trained into the model goes into the model; what can't goes into the scaffolding (Claude Code itself); what still can't is left for users to compose (tmux with multiple windows running parallel sessions, or claude -p inside a GitHub Action) — the harness occupies only the middle layer, deliberately kept thin.

The terminal form brings properties an IDE can't offer:

  • Composability: cat error.log | claude, claude -p in CI, shell loops that migrate files in bulk — it honors the Unix text-I/O contract and slots naturally into any existing workflow.
  • Environment universality: SSH, containers, remote servers, headless CI runners — anywhere a shell runs, it runs.
  • Not mutually exclusive with the IDE: it isn't a Cursor replacement but an orthogonal tool — many users run both.

The cost of honesty

The costs of the terminal form are real: a poor diff-review experience, a high barrier for users who aren't CLI natives, poor readability for long output. Anthropic later shipped an IDE extension and desktop/web forms as complements. The lesson of this case is not "the terminal is always right" but that the interface choice should serve the harness's core assumption — Claude Code assumes "the agent works autonomously; humans approve and supervise," which the terminal matches exactly; Cursor assumes "humans write code, AI assists," so it lives in the IDE.

Design Philosophy, Summed Up: Simplicity as a Bet on Model Progress ​

Read Claude Code's component choices together and a single thread runs through them:

  • memory = one Markdown file, not a vector database;
  • context compaction (compact) = "have Claude summarize the messages so far," not a clever selective-truncation algorithm (Boris's exact words, after trying several approaches: "we just did the simplest thing");
  • tools = grep and bash, not a semantic code index;
  • planning = one three-state list, not a plan-and-execute two-model architecture;
  • enforcement = shell-script hooks, not a rules engine.

This isn't crudeness; it's design for disposability: every mechanism is simple enough to be discarded or simplified when the model gets stronger. The opposite is the thick scaffolding built to compensate for weak models — one model upgrade, and it turns from an asset into a liability (this dynamic is discussed in detail in Model vs. Harness). The principle Anthropic gives in Building effective agents is the same one: find the simplest thing that works, and add complexity only when it's clearly necessary.

Of course, "simple" rests on a hidden premise: your model has to be strong enough. grep can stand in for semantic retrieval because frontier models genuinely use grep skillfully; CLAUDE.md gets followed because the model's instruction-following is up to par. You can copy this design philosophy, but answer one question first: can your model carry it? That's exactly why harness engineering has to co-evolve with model capability — the matching problem that Design Principles keeps coming back to.

Further Reading ​

References ​