Skip to content

Anatomy of the Harness

At a glance One diagram to understand the overall architecture of an Agent Harness: how the user, the environment, the agent loop, the eight core components, and the LLM fit together — with a full data-flow walkthrough and role-based reading routes.

Anatomy of the Harness ​

This is the hub page of the site. The three preceding guide posts answered "what is a harness," "why it matters," and "where it came from." This page does something more concrete: take a production-grade agent harness and break it into an architecture diagram you can point at — what each component is responsible for, its hardest design problem, and which interface it uses to talk to its neighbors.

After this page, you should be able to take any agent product (Claude Code, Cursor, Devin…) and place it on this diagram, naming which box its center of gravity falls into.

How to read this

If you haven't read What Is an Agent Harness? yet, read that first to build the intuition. If you already accept the claim "the harness sets the ceiling of agent capability" (argued in Model vs. Harness), read straight on.

The one big diagram ​

An agent harness divides into four layers: user, harness, model, environment. The model does exactly one thing — generate output given input. Everything else happens in the harness layer.

text
┌─────────────────────────────────────────────────────────────────┐
│ User layer    Task handoff · mid-course feedback · approvals ·  │
│               interrupting and correcting                       │
└──────────────────────────────┬──────────────────────────────────┘
                               ▼
┌─────────────────────────────────────────────────────────────────┐
│ Harness layer (everything this site dissects)                   │
│                                                                 │
│   ┌────── Agent loop (the beating heart of the system) ──────┐  │
│   │                                                          │  │
│   │  assemble context ─▶ call model ─▶ parse actions ─▶      │  │
│   │  execute ─▶ observe & feed back                          │  │
│   │       ▲                                            │     │  │
│   │       └────────────────────────────────────────────┘     │  │
│   │      until: task done / asks for help / circuit breaker  │  │
│   └──────────────────────────────────────────────────────────┘  │
│                                                                 │
│   In-loop components   context engineering · tools · planning   │
│                        · memory                                 │
│   Extension components subagent orchestration · skill injection │
│   Cross-cutting        permissions & safety · evals &           │
│                        observability (every loop turn passes    │
│                        through them)                            │
└──────────────────────────────┬──────────────────────────────────┘
                               ▼
┌─────────────────────────────────────────────────────────────────┐
│ Model layer    The LLM: the system's only "thinker"             │
│                What it sees at each step and where its output   │
│                goes — all decided by the harness                │
└──────────────────────────────┬──────────────────────────────────┘
                               ▼
┌─────────────────────────────────────────────────────────────────┐
│ Environment layer   file system · terminal · browser · external │
│                     APIs · code repositories                    │
│                     Where tool execution actually lands; where  │
│                     observations come from                      │
└─────────────────────────────────────────────────────────────────┘

Three sentences to keep in mind while reading the diagram:

  1. The loop is the trunk; components are organs hung on it. Context engineering decides what the "assemble context" step produces; the tool system decides what the "execute" step can do; the permission system cuts in sideways before every action fires.
  2. The model layer's inputs and outputs are fully mediated by the harness. The model never "sees" the environment directly — it sees an environment representation assembled by context engineering; its "actions" never take effect directly either — they must pass parsing, permission checks, and a tool runtime. That is the literal meaning of the word harness.
  3. The user doesn't only appear at the beginning. Approvals, interruptions, and mid-course feedback are normal in-loop events, not exceptions. Anthropic's Building effective agents (December 2024) defines an agent as a system where "the LLM dynamically directs its own processes and tool usage, maintaining control over how it accomplishes the task" — yet even their own product lists human checkpoints as standard practice.

Dissecting the components ​

The nine subsections below each correspond to a component page on this site. The format is uniform: responsibility → key design questions → interfaces with neighbors.

The Agent Loop ​

Responsibility: the system's main control flow. Decides when to call the model, how to parse output, the order in which actions execute, and when to stop. It's the only part of the harness that is "always running"; every other component is invoked by it at specific moments.

Key design questions:

  • How many actions may the model emit per turn? (Serial single action, or a parallel batch — directly affects latency and consistency)
  • How are termination conditions defined? Beyond the model declaring itself done, is external verification needed (run the tests, check the diff)?
  • Exception paths: how many retries on a malformed output? How does the circuit breaker trip on infinite loops? What happens when the cost ceiling is exceeded?

Interfaces: upstream, takes tasks and interrupt signals from the user layer; gets each turn's model input from "context engineering"; hands parsed actions to the "tool system"; seeks approval for high-risk operations from "permissions & safety"; exposes structured events for every step to "evals & observability."

→ See The Agent Loop

Context Engineering ​

Responsibility: at the start of every loop turn, decide what the model "sees" this time — the system prompt, conversation history, tool results, retrieved code, memory fragments, skill instructions: with what structure, in what order, at what length, packed into a finite context window. Since 2024 this has been widely regarded as the highest-leverage component in the industry: the model is fixed, but the input is entirely under your control.

Key design questions:

  • What's the context budget? When exceeded — truncate, compress into summaries, or offload to external storage for on-demand retrieval?
  • Do tool results get denoised? (One grep can return thousands of lines)
  • What goes in the system prompt (stable, cacheable) versus the conversation flow (dynamic)? Cache hit rate directly determines cost and time-to-first-token.

Interfaces: inputs come from "memory" (what to fetch), "skill injection" (what to mount), the "tool system" (fresh observations), and "planning" (current plan state); there is exactly one output — the message list handed to the loop and passed to the LLM.

→ See Context Engineering

Tools & MCP ​

Responsibility: define what the model can do to the environment: the tool catalog, each tool's schema and description, the execution runtime (sandbox, timeouts, resource limits), and the serialization format of results. A tool description is itself documentation written for the model to read — how well it's written directly changes model behavior.

Key design questions:

  • General-purpose tools (bash, file read/write) or domain-specific ones (apply_patch, search_symbol)? The SWE-agent paper (NeurIPS 2024) named this problem agent-computer interface (ACI) design and showed that the same model jumps from 3.8% to 12.5% pass@1 on SWE-bench purely through interface design.
  • Tool granularity: one do-everything edit_file, or a set of small tools? Granularity affects the model's choice difficulty and error-recovery cost.
  • Adopt standardized protocols like MCP (Model Context Protocol) in exchange for a tool ecosystem?

Interfaces: schemas go to "context engineering" to be packed into the prompt; execution requests come from the loop; every execution passes "permissions & safety" first; results flow back into the context after denoising; all call logs land in "evals & observability."

→ See Tools & MCP

Planning & Task Decomposition ​

Responsibility: turn "a task too big to finish in one turn" into a trackable sequence of steps. The design space spans a huge range: free-form planning in the model's own thinking (ReAct-style, Yao et al., ICLR 2023), explicit TODO-list tools, all the way up to separate planner-executor layered architectures.

Key design questions:

  • Explicit or implicit? Explicit plans can be inspected, corrected, and shown to the user, but add overhead — and plans go stale against environment feedback. How often to replan?
  • Where does the plan live? In the context (visible to the model every turn but costs budget) or in external state (needs active re-injection to be noticed)?
  • Who verifies plan progress — the model's self-assessment, or deterministic checks by the loop (tests passing, subtasks checked off)?

Interfaces: plan state is a major input to "context engineering"; decomposed subtasks may be dispatched to "subagents"; plan persistence is handed to "memory."

→ See Planning & Task Decomposition

Memory Systems ​

Responsibility: fight the finitude of the context window. Two time scales: within a session — how to compress, summarize, and filter when history grows too long; across sessions — user preferences, project conventions, lessons learned: in what form to store them and how to recall them next time.

Key design questions:

  • When to write: record everything (noise explosion) or write on explicit trigger (the model or user decides "this is worth remembering")?
  • How to recall: inject everything, keyword matching, vector retrieval, or let the model actively query via a tool? Wrong recall is worse than no recall — stale memories will reliably mislead the model.
  • Who may write memory? Model-written memory files are an attack surface that requires auditing.

Interfaces: the read side hangs off the "context engineering" assembly pipeline; the write side is usually exposed as a tool in the "tool system"; persisted content falls under the audit scope of "permissions & safety."

→ See Memory Systems

Subagents & Multi-Agent Orchestration ​

Responsibility: delegate part of a task together with a clean context. The subagent's most important value isn't parallelism — it's context isolation: the main agent's context stays unpolluted by dozens of exploratory searches and receives only distilled conclusions.

Key design questions:

  • Delegation boundaries: which tasks are worth splitting out? The communication cost (the task description must be self-contained) often exceeds the benefit.
  • What do the subagent and the main agent share? Tools, file system, memory — the more shared, the easier; the more isolated, the safer.
  • Hierarchy depth: may a subagent spawn sub-subagents? Runaway recursive delegation is a black hole for cost and latency.

Interfaces: subtasks are the output of "planning"; each subagent runs a complete "agent loop" internally, with its own "context engineering" and restricted "tool system" and "permissions" configuration; its returned result is one observation in the main agent's context.

→ See Subagents & Multi-Agent Orchestration

Permissions, Safety & Human-in-the-Loop ​

Responsibility: answer "is this allowed?" before any action takes effect. Includes: tool tiering (read-only / reversible write / irreversible write / external side effects), approval interactions (ask every time / remember the choice / pre-authorization rules), sandbox isolation, and making "when to stop and ask a human" a first-class protocol.

Key design questions:

  • Expressiveness of the permission model: blunt per-tool rules, or fine-grained matching on "tool × argument pattern × target path"? bash rm -rf / and bash ls shouldn't be in the same tier.
  • Approval fatigue: ask too often and users will blindly click "allow" — reducing interruptions through rules, allowlists, and risk tiers is one of the few pure-UX problems in the safety component.
  • The exchange rate between autonomy and reversibility: the more irreversible the operation (sending email, pushing code, deleting data), the heavier the human confirmation required.

Interfaces: sits between the loop's "execute" step and the "tool system"; approval requests go up to the user layer; every allow/deny decision is recorded into "evals & observability"; enforces stricter defaults on "subagents" than on the main agent.

→ See Permissions, Safety & Human-in-the-Loop

Skills & Knowledge Injection ​

Responsibility: give a general agent on-demand domain expertise — package the how-to manuals for a class of work (procedures, templates, scripts, checklists) into discoverable, mountable units that get injected into the context when a relevant task appears, instead of permanently stuffing all knowledge into the system prompt.

Key design questions:

  • Discovery: how does the model know "this skill should load now"? Self-retrieval by name/description, keyword triggers, or a routing model?
  • Progressive disclosure: first a one-line summary; the full text when needed; the bundled script when needed — each level of depth costs only that level of context.
  • Where's the boundary between skills, ordinary documents, the system prompt, and tools? (Rule of thumb: skills teach the model "procedural knowledge"; tools give the model "capability"; memory remembers "facts.")

Interfaces: injected content enters the prompt via "context engineering"; a skill can declare which "tools" it needs; skill packs themselves are a supply chain in the "permissions" sense (third-party skills need auditing).

→ See Skills & Knowledge Injection

Evaluation & Observability ​

Responsibility: answer questions at two levels. Runtime observability: can every loop step's inputs and outputs, tool calls, token spend, and latency be replayed in full? Offline evaluation: when any harness component changes, did things get better or worse — and by what benchmark, what metric?

Key design questions:

  • The event model: structuring every loop step as an event stream is the common choice of modern agent frameworks — OpenHands (formerly OpenDevin, platform paper) made the "action + observation" event stream the central abstraction of its whole architecture, with observability, replay, and multi-agent coordination all built on it.
  • Eval attribution: did the score change come from the model, the prompt, the tools, or environment noise? Without per-component controlled experiments, harness iteration is gambling.
  • Cost observability: the cost variance of agent tasks is enormous (the same task may "complete" in 5 turns or 50); an evaluation that doesn't track cost is incomplete.

Interfaces: a pure consumer — it subscribes to every event produced by the loop, tools, and permissions; in return, its findings set the iteration direction for all components.

→ See Evaluation & Observability

Walk the full data flow ​

Paper understanding only goes so far. Walk the whole chain with one concrete task: "add rate limiting to this repo's login endpoint."

Step 0 · The task enters. The user types the task. The harness does three things: load project-level conventions (a memory file like CLAUDE.md), match potentially relevant skills, initialize the permission session.

Step 1 · Assemble the first context. Context engineering produces:

text
system:    Role and behavior rules + tool usage norms (stable prefix, cache hit)
context:   Project conventions (memory) + rate-limiting best practices
           (skill, mounted because of a keyword hit)
user:      "Add rate limiting to this repo's login endpoint"
tools:     schemas for bash / read_file / edit_file / grep / todo_write

Step 2 · The first model call. The model outputs a tool call, not an answer:

json
{ "tool": "grep", "args": { "pattern": "login", "path": "src/", "type": "py" } }

Step 3 · Parse, authorize, execute. The loop parses the action → the permission system sees that grep is read-only and waves it through under pre-authorization → the tool runtime executes and returns 47 matches. The tool system denoises: truncate to the 10 most relevant, keeping the list of file paths.

Step 4 · The observation flows back. The denoised result is appended to the conversation history as a tool_result message. The context is now: initial context + step 2's tool call + step 3's result. Note: the model never "saw" the repository — it saw the repository as translated by the harness.

Step 5 · The loop continues. The model reads key files → calls todo_write to lay down a three-step plan (the planning component activates) → reads the existing middleware code → drafts an edit_file.

Step 6 · One permission interception. edit_file is a reversible write, allowed by the preset rules; but when the model then wants to run bash pytest, it hits a pattern the user hasn't pre-authorized. The loop pauses and the approval request goes up to the user layer. The user clicks "allow test commands for this session" — the permission system remembers this decision, and similar commands don't prompt again.

Step 7 · Verification-driven convergence. The tests fail for two rounds (the rate limiter's counter has a race condition under concurrency); the model fixes the implementation based on the failure output fed back. The third round passes.

Step 8 · Termination and settlement. The model declares completion + the loop's external verification confirms (tests all green, diff non-empty) → the loop exits. At wrap-up: the memory system writes "this project uses pytest-xdist; test commands need -n auto" into project memory; the observability component records this session: 23 loop turns, 41k tokens, 2 user approvals.

text
User task
   │
   ▼
┌─ loop turn n ───────────────────────────────────────────┐
│  Context engineering assembles the prompt               │
│  (memory + skills + history + tool results)             │
│        │                                                │
│        ▼                                                │
│  LLM outputs: text or tool call                         │
│        │                                                │
│        ▼                                                │
│  Permission check ──denied/needs approval──▶ user ──▶   │
│        │ allowed                   result fed back      │
│        ▼                                                │
│  Tool runtime executes (sandbox · timeout · denoise)    │
│        │                                                │
│        ▼                                                │
│  Observation returns as tool_result ──▶ turn n+1        │
└──────────────────────────────────────────────────────────┘
   │
   ▼  model declares done ∧ external verification passes
Memory settled · events archived · session ends

The key engineering insight

Across this entire chain, the only uncontrollable step is the model's sampling in step 2; every other step is deterministic code. The essence of harness engineering is to make deterministic everything that can be made deterministic (authorization, denoising, verification, circuit-breaking), so the model only exercises judgment where judgment is truly needed.

Trade-offs ​

Nothing on this architecture diagram is free. The four most central tensions:

TensionPull leftPull rightTypical stance
Autonomy vs. controlFewer approvals, longer unattended loopsHuman confirmation for every writeTier by operational reversibility, not a global blunt rule
Context richness vs. attention dilutionMore content = more complete information for the modelMore content = key signals drown and costs riseAggressive denoising + on-demand retrieval (toolized search)
General tools vs. bespoke interfacesOne bash for everything; portableCustom ACI: fewer errors, shorter trajectories (the SWE-agent lesson)Customize for the target task set; keep an escape hatch
Deep single agent vs. parallel multi-agentCoherent context, zero communication lossIsolated contexts, parallelizable explorationDefault to a single agent; delegate only when context pollution is evident

Two through-lines, argued fully in Harness Design Principles:

  • Simplicity first. The first recommendation in Anthropic's much-cited essay: find the simplest solution first and add complexity only when demonstrably necessary — many scenarios just need optimizing a single LLM call (add retrieval, add examples); they don't need an agent, let alone a multi-agent system.
  • Every component must be observable and individually replaceable. Design component interfaces as structured data (message lists, event streams, tool schemas), not implicit conventions — that's the precondition for attribution evals and piece-by-piece iteration later.

Reading routes by role ​

Researchers ​

You care about "which designs actually move the capability ceiling." Suggested order:

  1. Model vs. Harness: Why the Harness Sets the Ceiling — the core claim and the evidence
  2. A Brief History — how the questions evolved from ReAct to today
  3. The Paper Map → Classic Papers, Annotated → Frontier Developments
  4. SWE-agent and OpenHands — the two most research-flavored case studies: ACI and event streams
  5. Evaluation & Observability — the experimental methodology you can't do agent research without

Engineers ​

You care about "I want to build one — where do I start, how do I avoid the holes." Suggested order:

  1. This page — build the component map
  2. The Agent Loop → Context Engineering → Tools & MCP — the minimum viable trio
  3. Permissions, Safety & Human-in-the-Loop — required reading before launch
  4. Build a Minimal Harness from Scratch — get the trio running with your own hands
  5. Common Pitfalls & Anti-Patterns + Case Study: Claude Code — calibrate your trade-offs against a mature product

Product managers ​

You care about "where are the capability boundaries, where does the experience difference come from, how should I write requirements." Suggested order:

  1. What Is an Agent Harness? → Model vs. Harness
  2. This page — focus on the "Trade-offs" section
  3. Claude Code, Cursor, Devin — the harness trade-offs behind three product shapes
  4. Planning & Task Decomposition and Subagents — the technical answer to "how big a task can an agent actually handle"
  5. Harness Design Principles — a shared language for talking with the engineering team

Further reading ​

References ​