Appearance
Anatomy of the Harness
This is the hub page of the site. The three preceding guide posts answered "what is a harness," "why it matters," and "where it came from." This page does something more concrete: take a production-grade agent harness and break it into an architecture diagram you can point at — what each component is responsible for, its hardest design problem, and which interface it uses to talk to its neighbors.
After this page, you should be able to take any agent product (Claude Code, Cursor, Devin…) and place it on this diagram, naming which box its center of gravity falls into.
How to read this
If you haven't read What Is an Agent Harness? yet, read that first to build the intuition. If you already accept the claim "the harness sets the ceiling of agent capability" (argued in Model vs. Harness), read straight on.
The one big diagram
An agent harness divides into four layers: user, harness, model, environment. The model does exactly one thing — generate output given input. Everything else happens in the harness layer.
text
┌─────────────────────────────────────────────────────────────────┐
│ User layer Task handoff · mid-course feedback · approvals · │
│ interrupting and correcting │
└──────────────────────────────┬──────────────────────────────────┘
▼
┌─────────────────────────────────────────────────────────────────┐
│ Harness layer (everything this site dissects) │
│ │
│ ┌────── Agent loop (the beating heart of the system) ──────┐ │
│ │ │ │
│ │ assemble context ─▶ call model ─▶ parse actions ─▶ │ │
│ │ execute ─▶ observe & feed back │ │
│ │ ▲ │ │ │
│ │ └────────────────────────────────────────────┘ │ │
│ │ until: task done / asks for help / circuit breaker │ │
│ └──────────────────────────────────────────────────────────┘ │
│ │
│ In-loop components context engineering · tools · planning │
│ · memory │
│ Extension components subagent orchestration · skill injection │
│ Cross-cutting permissions & safety · evals & │
│ observability (every loop turn passes │
│ through them) │
└──────────────────────────────┬──────────────────────────────────┘
▼
┌─────────────────────────────────────────────────────────────────┐
│ Model layer The LLM: the system's only "thinker" │
│ What it sees at each step and where its output │
│ goes — all decided by the harness │
└──────────────────────────────┬──────────────────────────────────┘
▼
┌─────────────────────────────────────────────────────────────────┐
│ Environment layer file system · terminal · browser · external │
│ APIs · code repositories │
│ Where tool execution actually lands; where │
│ observations come from │
└─────────────────────────────────────────────────────────────────┘Three sentences to keep in mind while reading the diagram:
- The loop is the trunk; components are organs hung on it. Context engineering decides what the "assemble context" step produces; the tool system decides what the "execute" step can do; the permission system cuts in sideways before every action fires.
- The model layer's inputs and outputs are fully mediated by the harness. The model never "sees" the environment directly — it sees an environment representation assembled by context engineering; its "actions" never take effect directly either — they must pass parsing, permission checks, and a tool runtime. That is the literal meaning of the word harness.
- The user doesn't only appear at the beginning. Approvals, interruptions, and mid-course feedback are normal in-loop events, not exceptions. Anthropic's Building effective agents (December 2024) defines an agent as a system where "the LLM dynamically directs its own processes and tool usage, maintaining control over how it accomplishes the task" — yet even their own product lists human checkpoints as standard practice.
Dissecting the components
The nine subsections below each correspond to a component page on this site. The format is uniform: responsibility → key design questions → interfaces with neighbors.
The Agent Loop
Responsibility: the system's main control flow. Decides when to call the model, how to parse output, the order in which actions execute, and when to stop. It's the only part of the harness that is "always running"; every other component is invoked by it at specific moments.
Key design questions:
- How many actions may the model emit per turn? (Serial single action, or a parallel batch — directly affects latency and consistency)
- How are termination conditions defined? Beyond the model declaring itself done, is external verification needed (run the tests, check the diff)?
- Exception paths: how many retries on a malformed output? How does the circuit breaker trip on infinite loops? What happens when the cost ceiling is exceeded?
Interfaces: upstream, takes tasks and interrupt signals from the user layer; gets each turn's model input from "context engineering"; hands parsed actions to the "tool system"; seeks approval for high-risk operations from "permissions & safety"; exposes structured events for every step to "evals & observability."
→ See The Agent Loop
Context Engineering
Responsibility: at the start of every loop turn, decide what the model "sees" this time — the system prompt, conversation history, tool results, retrieved code, memory fragments, skill instructions: with what structure, in what order, at what length, packed into a finite context window. Since 2024 this has been widely regarded as the highest-leverage component in the industry: the model is fixed, but the input is entirely under your control.
Key design questions:
- What's the context budget? When exceeded — truncate, compress into summaries, or offload to external storage for on-demand retrieval?
- Do tool results get denoised? (One
grepcan return thousands of lines) - What goes in the system prompt (stable, cacheable) versus the conversation flow (dynamic)? Cache hit rate directly determines cost and time-to-first-token.
Interfaces: inputs come from "memory" (what to fetch), "skill injection" (what to mount), the "tool system" (fresh observations), and "planning" (current plan state); there is exactly one output — the message list handed to the loop and passed to the LLM.
→ See Context Engineering
Tools & MCP
Responsibility: define what the model can do to the environment: the tool catalog, each tool's schema and description, the execution runtime (sandbox, timeouts, resource limits), and the serialization format of results. A tool description is itself documentation written for the model to read — how well it's written directly changes model behavior.
Key design questions:
- General-purpose tools (
bash, file read/write) or domain-specific ones (apply_patch,search_symbol)? The SWE-agent paper (NeurIPS 2024) named this problem agent-computer interface (ACI) design and showed that the same model jumps from 3.8% to 12.5% pass@1 on SWE-bench purely through interface design. - Tool granularity: one do-everything
edit_file, or a set of small tools? Granularity affects the model's choice difficulty and error-recovery cost. - Adopt standardized protocols like MCP (Model Context Protocol) in exchange for a tool ecosystem?
Interfaces: schemas go to "context engineering" to be packed into the prompt; execution requests come from the loop; every execution passes "permissions & safety" first; results flow back into the context after denoising; all call logs land in "evals & observability."
→ See Tools & MCP
Planning & Task Decomposition
Responsibility: turn "a task too big to finish in one turn" into a trackable sequence of steps. The design space spans a huge range: free-form planning in the model's own thinking (ReAct-style, Yao et al., ICLR 2023), explicit TODO-list tools, all the way up to separate planner-executor layered architectures.
Key design questions:
- Explicit or implicit? Explicit plans can be inspected, corrected, and shown to the user, but add overhead — and plans go stale against environment feedback. How often to replan?
- Where does the plan live? In the context (visible to the model every turn but costs budget) or in external state (needs active re-injection to be noticed)?
- Who verifies plan progress — the model's self-assessment, or deterministic checks by the loop (tests passing, subtasks checked off)?
Interfaces: plan state is a major input to "context engineering"; decomposed subtasks may be dispatched to "subagents"; plan persistence is handed to "memory."
→ See Planning & Task Decomposition
Memory Systems
Responsibility: fight the finitude of the context window. Two time scales: within a session — how to compress, summarize, and filter when history grows too long; across sessions — user preferences, project conventions, lessons learned: in what form to store them and how to recall them next time.
Key design questions:
- When to write: record everything (noise explosion) or write on explicit trigger (the model or user decides "this is worth remembering")?
- How to recall: inject everything, keyword matching, vector retrieval, or let the model actively query via a tool? Wrong recall is worse than no recall — stale memories will reliably mislead the model.
- Who may write memory? Model-written memory files are an attack surface that requires auditing.
Interfaces: the read side hangs off the "context engineering" assembly pipeline; the write side is usually exposed as a tool in the "tool system"; persisted content falls under the audit scope of "permissions & safety."
→ See Memory Systems
Subagents & Multi-Agent Orchestration
Responsibility: delegate part of a task together with a clean context. The subagent's most important value isn't parallelism — it's context isolation: the main agent's context stays unpolluted by dozens of exploratory searches and receives only distilled conclusions.
Key design questions:
- Delegation boundaries: which tasks are worth splitting out? The communication cost (the task description must be self-contained) often exceeds the benefit.
- What do the subagent and the main agent share? Tools, file system, memory — the more shared, the easier; the more isolated, the safer.
- Hierarchy depth: may a subagent spawn sub-subagents? Runaway recursive delegation is a black hole for cost and latency.
Interfaces: subtasks are the output of "planning"; each subagent runs a complete "agent loop" internally, with its own "context engineering" and restricted "tool system" and "permissions" configuration; its returned result is one observation in the main agent's context.
→ See Subagents & Multi-Agent Orchestration
Permissions, Safety & Human-in-the-Loop
Responsibility: answer "is this allowed?" before any action takes effect. Includes: tool tiering (read-only / reversible write / irreversible write / external side effects), approval interactions (ask every time / remember the choice / pre-authorization rules), sandbox isolation, and making "when to stop and ask a human" a first-class protocol.
Key design questions:
- Expressiveness of the permission model: blunt per-tool rules, or fine-grained matching on "tool × argument pattern × target path"?
bash rm -rf /andbash lsshouldn't be in the same tier. - Approval fatigue: ask too often and users will blindly click "allow" — reducing interruptions through rules, allowlists, and risk tiers is one of the few pure-UX problems in the safety component.
- The exchange rate between autonomy and reversibility: the more irreversible the operation (sending email, pushing code, deleting data), the heavier the human confirmation required.
Interfaces: sits between the loop's "execute" step and the "tool system"; approval requests go up to the user layer; every allow/deny decision is recorded into "evals & observability"; enforces stricter defaults on "subagents" than on the main agent.
→ See Permissions, Safety & Human-in-the-Loop
Skills & Knowledge Injection
Responsibility: give a general agent on-demand domain expertise — package the how-to manuals for a class of work (procedures, templates, scripts, checklists) into discoverable, mountable units that get injected into the context when a relevant task appears, instead of permanently stuffing all knowledge into the system prompt.
Key design questions:
- Discovery: how does the model know "this skill should load now"? Self-retrieval by name/description, keyword triggers, or a routing model?
- Progressive disclosure: first a one-line summary; the full text when needed; the bundled script when needed — each level of depth costs only that level of context.
- Where's the boundary between skills, ordinary documents, the system prompt, and tools? (Rule of thumb: skills teach the model "procedural knowledge"; tools give the model "capability"; memory remembers "facts.")
Interfaces: injected content enters the prompt via "context engineering"; a skill can declare which "tools" it needs; skill packs themselves are a supply chain in the "permissions" sense (third-party skills need auditing).
→ See Skills & Knowledge Injection
Evaluation & Observability
Responsibility: answer questions at two levels. Runtime observability: can every loop step's inputs and outputs, tool calls, token spend, and latency be replayed in full? Offline evaluation: when any harness component changes, did things get better or worse — and by what benchmark, what metric?
Key design questions:
- The event model: structuring every loop step as an event stream is the common choice of modern agent frameworks — OpenHands (formerly OpenDevin, platform paper) made the "action + observation" event stream the central abstraction of its whole architecture, with observability, replay, and multi-agent coordination all built on it.
- Eval attribution: did the score change come from the model, the prompt, the tools, or environment noise? Without per-component controlled experiments, harness iteration is gambling.
- Cost observability: the cost variance of agent tasks is enormous (the same task may "complete" in 5 turns or 50); an evaluation that doesn't track cost is incomplete.
Interfaces: a pure consumer — it subscribes to every event produced by the loop, tools, and permissions; in return, its findings set the iteration direction for all components.
→ See Evaluation & Observability
Walk the full data flow
Paper understanding only goes so far. Walk the whole chain with one concrete task: "add rate limiting to this repo's login endpoint."
Step 0 · The task enters. The user types the task. The harness does three things: load project-level conventions (a memory file like CLAUDE.md), match potentially relevant skills, initialize the permission session.
Step 1 · Assemble the first context. Context engineering produces:
text
system: Role and behavior rules + tool usage norms (stable prefix, cache hit)
context: Project conventions (memory) + rate-limiting best practices
(skill, mounted because of a keyword hit)
user: "Add rate limiting to this repo's login endpoint"
tools: schemas for bash / read_file / edit_file / grep / todo_writeStep 2 · The first model call. The model outputs a tool call, not an answer:
json
{ "tool": "grep", "args": { "pattern": "login", "path": "src/", "type": "py" } }Step 3 · Parse, authorize, execute. The loop parses the action → the permission system sees that grep is read-only and waves it through under pre-authorization → the tool runtime executes and returns 47 matches. The tool system denoises: truncate to the 10 most relevant, keeping the list of file paths.
Step 4 · The observation flows back. The denoised result is appended to the conversation history as a tool_result message. The context is now: initial context + step 2's tool call + step 3's result. Note: the model never "saw" the repository — it saw the repository as translated by the harness.
Step 5 · The loop continues. The model reads key files → calls todo_write to lay down a three-step plan (the planning component activates) → reads the existing middleware code → drafts an edit_file.
Step 6 · One permission interception. edit_file is a reversible write, allowed by the preset rules; but when the model then wants to run bash pytest, it hits a pattern the user hasn't pre-authorized. The loop pauses and the approval request goes up to the user layer. The user clicks "allow test commands for this session" — the permission system remembers this decision, and similar commands don't prompt again.
Step 7 · Verification-driven convergence. The tests fail for two rounds (the rate limiter's counter has a race condition under concurrency); the model fixes the implementation based on the failure output fed back. The third round passes.
Step 8 · Termination and settlement. The model declares completion + the loop's external verification confirms (tests all green, diff non-empty) → the loop exits. At wrap-up: the memory system writes "this project uses pytest-xdist; test commands need -n auto" into project memory; the observability component records this session: 23 loop turns, 41k tokens, 2 user approvals.
text
User task
│
▼
┌─ loop turn n ───────────────────────────────────────────┐
│ Context engineering assembles the prompt │
│ (memory + skills + history + tool results) │
│ │ │
│ ▼ │
│ LLM outputs: text or tool call │
│ │ │
│ ▼ │
│ Permission check ──denied/needs approval──▶ user ──▶ │
│ │ allowed result fed back │
│ ▼ │
│ Tool runtime executes (sandbox · timeout · denoise) │
│ │ │
│ ▼ │
│ Observation returns as tool_result ──▶ turn n+1 │
└──────────────────────────────────────────────────────────┘
│
▼ model declares done ∧ external verification passes
Memory settled · events archived · session endsThe key engineering insight
Across this entire chain, the only uncontrollable step is the model's sampling in step 2; every other step is deterministic code. The essence of harness engineering is to make deterministic everything that can be made deterministic (authorization, denoising, verification, circuit-breaking), so the model only exercises judgment where judgment is truly needed.
Trade-offs
Nothing on this architecture diagram is free. The four most central tensions:
| Tension | Pull left | Pull right | Typical stance |
|---|---|---|---|
| Autonomy vs. control | Fewer approvals, longer unattended loops | Human confirmation for every write | Tier by operational reversibility, not a global blunt rule |
| Context richness vs. attention dilution | More content = more complete information for the model | More content = key signals drown and costs rise | Aggressive denoising + on-demand retrieval (toolized search) |
| General tools vs. bespoke interfaces | One bash for everything; portable | Custom ACI: fewer errors, shorter trajectories (the SWE-agent lesson) | Customize for the target task set; keep an escape hatch |
| Deep single agent vs. parallel multi-agent | Coherent context, zero communication loss | Isolated contexts, parallelizable exploration | Default to a single agent; delegate only when context pollution is evident |
Two through-lines, argued fully in Harness Design Principles:
- Simplicity first. The first recommendation in Anthropic's much-cited essay: find the simplest solution first and add complexity only when demonstrably necessary — many scenarios just need optimizing a single LLM call (add retrieval, add examples); they don't need an agent, let alone a multi-agent system.
- Every component must be observable and individually replaceable. Design component interfaces as structured data (message lists, event streams, tool schemas), not implicit conventions — that's the precondition for attribution evals and piece-by-piece iteration later.
Reading routes by role
Researchers
You care about "which designs actually move the capability ceiling." Suggested order:
- Model vs. Harness: Why the Harness Sets the Ceiling — the core claim and the evidence
- A Brief History — how the questions evolved from ReAct to today
- The Paper Map → Classic Papers, Annotated → Frontier Developments
- SWE-agent and OpenHands — the two most research-flavored case studies: ACI and event streams
- Evaluation & Observability — the experimental methodology you can't do agent research without
Engineers
You care about "I want to build one — where do I start, how do I avoid the holes." Suggested order:
- This page — build the component map
- The Agent Loop → Context Engineering → Tools & MCP — the minimum viable trio
- Permissions, Safety & Human-in-the-Loop — required reading before launch
- Build a Minimal Harness from Scratch — get the trio running with your own hands
- Common Pitfalls & Anti-Patterns + Case Study: Claude Code — calibrate your trade-offs against a mature product
Product managers
You care about "where are the capability boundaries, where does the experience difference come from, how should I write requirements." Suggested order:
- What Is an Agent Harness? → Model vs. Harness
- This page — focus on the "Trade-offs" section
- Claude Code, Cursor, Devin — the harness trade-offs behind three product shapes
- Planning & Task Decomposition and Subagents — the technical answer to "how big a task can an agent actually handle"
- Harness Design Principles — a shared language for talking with the engineering team
Further reading
- The component panorama: The Agent Loop · Context Engineering · Tools & MCP · Planning & Task Decomposition · Memory Systems · Subagents & Multi-Agent Orchestration · Permissions, Safety & Human-in-the-Loop · Skills & Knowledge Injection · Evaluation & Observability
- The empirical sources behind the architecture diagram: SWE-agent, OpenHands, Aider, LangGraph
- Quick reference: Glossary, Curated Resources
References
- Anthropic, Building effective agents, 2024-12 — the classic workflow/agent distinction and the "simplicity first" engineering advice.
- Yang et al., SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering, NeurIPS 2024 — the origin of the ACI concept and the 12.5% pass@1 (GPT-4) SWE-bench result.
- Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models, ICLR 2023 — the foundational work on implicit planning (alternating reasoning and acting).
- Wang et al., OpenHands: An Open Platform for AI Software Developers as Generalist Agents, 2024 — the platform design behind the event-stream architecture and sandboxed execution.