Skip to content

Harness Design Principles

At a glance Actionable harness design principles distilled from the practice of Anthropic, Cognition, 12-Factor Agents, and mainstream coding agents — each with a rationale and a positive/negative example.

Harness Design Principles ​

In Model vs. Harness we made the point that the same model, wrapped in different harnesses, can deliver night-and-day results. Which raises the question: when you set out to design a harness, what should you actually follow?

The bad news: this field doesn't yet have mature design laws on the order of database normalization or REST constraints. The good news: starting in late 2024, a handful of engineering write-ups rapidly converged on a set of principles that have been validated again and again. This page distills them into eight actionable rules, each with a statement, a rationale, a positive example, and a counterexample.

What this page is

This isn't an academic-style taxonomy — it's an engineering checklist. For every principle, ask yourself: "Is my harness violating it? And if it is, do I have good reason?"

Where the Principles Come From: Three Main Sources ​

Before laying out the principles, a word on where they come from. Understanding the context behind each source matters more than memorizing the entries.

Source 1: Anthropic, "Building Effective Agents" (December 2024) ​

On December 19, 2024, Anthropic's Erik Schluntz and Barry Zhang published "Building Effective Agents." The article made three core contributions:

  1. It drew a sharp line between workflow and agent: workflows are systems where "LLMs and tools are orchestrated through predefined code paths"; agents are systems where "the LLM dynamically directs its own processes and tool usage." The whole industry adopted this distinction.
  2. It laid out five foundational workflow patterns: prompt chaining, routing, parallelization, orchestrator-workers, and evaluator-optimizer.
  3. It kept hammering on simplicity first: in essence, "find the simplest solution that meets the need, and add complexity only when truly necessary," along with the candid admission that "agents trade latency and cost for better task performance, and whether that trade is worth it is yours to weigh."

Notably, the article comes from the team that built Claude and Claude Code — people who ship agent products at scale, yet wrote a guide saying, in effect, "don't rush into agents." That restraint is itself a principle.

Source 2: Cognition, "Don't Build Multi-Agents" (June 2025) ​

In June 2025, Walden Yan of Cognition (the company behind Devin) published this highly influential piece, putting forward two principles of context engineering:

  1. Share context: pass along the full agent trace, not just individual messages.
  2. Actions carry implicit decisions: every action an agent takes encodes its understanding of and decisions about the task; if different parts of the system act on different assumptions, their results will inevitably conflict.

The article's conclusion is radical: these two principles matter enough that "any architecture that violates them should be ruled out by default" — including the multi-agent parallel architectures that were fashionable at the time. The Flappy Bird example in the piece has become a classic: split "clone Flappy Bird" between two subagents, and one produces a Mario-style background while the other produces a bird that looks nothing like a game asset — leaving the agent in charge of merging the two staring at incompatible outputs with no way in.

Source 3: 12-Factor Agents (HumanLayer, 2025) ​

Borrowing from the famous Twelve-Factor App, HumanLayer's Dex Horthy proposed 12 factors for building reliable LLM applications; the list has circulated widely in engineering circles. All twelve:

  1. Natural Language to Tool Calls
  2. Own your prompts
  3. Own your context window
  4. Tools are just structured outputs
  5. Unify execution state and business state
  6. Launch/Pause/Resume with simple APIs
  7. Contact humans with tool calls
  8. Own your control flow
  9. Compact Errors into Context Window
  10. Small, Focused Agents
  11. Trigger from anywhere
  12. Make your agent a stateless reducer

The tone of the list is anti-framework: own your prompts, your context, and your control flow yourself, rather than handing them to a black-box framework.

What the three sources agree on

The three sources differ in stance and in depth, yet they share one striking consensus: the core of harness engineering is context engineering. For all its intelligence, the model is no longer the system's only bottleneck — what you feed it, what you hide from it, and when you interrupt it are.

Principle 1: Simplicity First — Prefer Workflows over Agents ​

Statement: Start from the simplest thing that could work — a single LLM call → a fixed workflow → a single-agent loop → a multi-agent system. Every step up in complexity must be paid for with a specific problem the tier below cannot solve.

Rationale: Complexity isn't free. Anthropic points out that agents trade latency and cost for task performance; Cognition demonstrated the structural flaws of multi-agent architectures around context sharing. Each layer of complexity brings new failure modes: workflows fail with predictable code bugs, agents fail with unpredictable model behavior, and multi-agent systems fail with systemic decision conflicts. Debugging cost rises exponentially along that gradient.

A common complexity ladder:

Complexity ↑      Predictability ↓      Debugging difficulty ↑

Single LLM call         translation, classification, extraction
    │
Fixed workflow          RAG, prompt chaining, routing
    │  (steps are predictable; control flow lives in code)
Single-agent loop       coding assistants, deep research
    │  (steps are unpredictable; control flow lives in the model)
Multi-agent system      parallel research, long-running tasks
    (decisions are scattered, context is fragmented — use with caution)

Good example: SWE-agent runs on a fairly plain single-agent loop — the model operates step by step inside a command-line interface designed specifically for it. The paper's focus isn't how clever the architecture is; it's how the design of the agent-computer interface (ACI) lets a simple loop perform strongly.

Bad example: An internal ticketing system chains four agents — "classifier agent + summarizer agent + reply agent + review agent" — when in reality every step's input and output is fully deterministic and a prompt-chaining workflow would do. The three extra agent loops contribute nothing but latency, token cost, and unpredictable cascading failures.

The litmus test

Ask yourself: "If I replaced this agent with a piece of deterministic code plus one LLM call, what would I lose?" If the answer is "nothing," what you need is a workflow, not an agent.

Principle 2: The Model Decides, the Harness Constrains ​

Statement: The model owns whatever requires judgment — what to do next, whether an answer is good enough. The harness owns whatever requires determinism — which tools may be called, which files may be touched, how much money may be spent, and when it must stop and ask a human.

Rationale: This is separation of concerns, agent-era edition. The model's strength is fuzzy judgment in an open world; its weakness is deterministic execution — and code is exactly the opposite. Putting constraints in the prompt ("please don't delete any files") delegates deterministic work to the component least suited for it — the model will violate the instruction with some probability, and that probability climbs as context grows longer and tasks get more complex. Putting constraints in the harness (the tool layer simply has no delete capability) is what counts as a guarantee in engineering terms.

┌─────────────────────────────────────────────┐
│                 User intent                  │
└──────────────────┬──────────────────────────┘
                   ▼
┌─────────────────────────────────────────────┐
│  Model layer: decisions                      │
│  "What's the next step?"                     │
│  "Is this result good enough?"               │
│  "Do I need to ask the user?"                │
└──────────────────┬──────────────────────────┘
                   ▼ proposes an action (tool call)
┌─────────────────────────────────────────────┐
│  Harness layer: constraints                  │
│  · Permissions: is this tool call allowed?   │
│  · Budget: steps/tokens/time exceeded?       │
│  · Guardrails: does this action need human   │
│    approval?                                 │
│  · State: packages results and errors to     │
│    feed back to the model                    │
└──────────────────┬──────────────────────────┘
                   ▼ execute or reject
               External environment

Good example: Claude Code's permission system puts the verdict on "can this Bash command run" in the harness's permission layer, backed by allowlists, directory boundaries, and human-confirmation hooks — the model is free to decide, but out-of-bounds actions are deterministically blocked before execution. See Permissions, Safety & Human-in-the-Loop.

Bad example: You write "you may only modify files under src/" in the system prompt, then hand the model a bare write_file(path, content) tool. The first time the model helpfully "fixes" a type declaration inside node_modules, you'll understand the difference between prompt constraints and enforced constraints.

Principle 3: Few Tools, Done Well ​

Statement: The tool set should be small and orthogonal, with every tool having a clear purpose, good error messages, and deterministic side effects. Better to give the model five well-designed tools than thirty APIs wrapped in a hurry.

Rationale: Tools are the model's hands, but every extra hand consumes attention budget inside the context window. Tool descriptions themselves cost tokens; the more tools there are, the likelier the model is to pick the wrong one or pass wrong arguments. One of the SWE-agent paper's core findings is that interface design significantly affects agent performance — they built the model a custom file viewer with line numbers that shows one window at a time, instead of exposing raw cat/sed, and that ACI design got it a then-best 12.5% pass@1 on SWE-bench. Tools are an interface designed for the model as a "new category of end user," not API documentation for human eyes.

Good example: A good read_file tool: line-numbered output, truncation with a note when the file is too long ("N lines total, showing 1–200"), and a "file not found — do you need to list_dir first?" response for missing files — the error message itself teaches the model how to recover.

Bad example: Registering all 80 endpoints of an internal OpenAPI spec as 80 tools, each description an auto-generated English blurb from Swagger. The model dithers between get_user_by_id and fetch_user_details_v2, guesses at parameter formats, and retries blindly after getting a 422.

A micro-example of tool design

A poor error return:

json
{ "error": "exit code 1" }

A good error return:

json
{
  "error": "grep: pattern 'handleRequest(' not found in any file.",
  "hint": "Searched 1,247 files under /repo. Did you mean 'handle_request'? Python code typically uses snake_case."
}

The latter turns "failure" into "information the model can act on for its next step." That's what Cognition means by: actions carry information, and error messages are context too.

Principle 4: Context Is Everything (Garbage In, Garbage Out) ​

Statement: The quality of the model's output at each step is almost entirely determined by the context it sees at that moment. The harness's first job is deciding what goes into the context at each step — and what comes out.

Rationale: Cognition calls context engineering "the number one job of engineers building AI agents," and that's not rhetoric. The same model, given the full error stack and the relevant files, can fix the bug; given just "the build failed," it can only guess. Context engineering covers: the system prompt, how tool results are organized, compression of message history, what gets injected from retrieval, and what gets passed between subagents — see Context Engineering.

This principle has two corollaries that map exactly onto Cognition's two context-engineering principles:

  • Pass subagents the full trace, not just the conclusion. Process information like "the main agent already searched X and ruled out hypothesis Y" determines whether a subagent will duplicate work or make contradictory assumptions.
  • Watch out for contradictory instructions in context. The model's attention is sensitive to conflict: the system prompt says "be concise," the user message pastes in "explain every step in detail," and the conversation history carries a third style of its own — output quality degrades under that tug-of-war.

Good example: Claude Code's subagent design matches Cognition's observation: subagents are typically dispatched to "answer a well-defined question" (say, "where is this function called?"), not to write code in parallel — because subagents writing code in parallel can't share each other's decision context, and their outputs are guaranteed to conflict. A subagent's exploration doesn't pollute the main context; what comes back is the distilled answer. See Subagents & Multi-Agent Orchestration.

Bad example: A support agent stuffs the verbatim text of every ticket from the past three months into each turn's context — stale policies, complaints already resolved, wrong promises made by other support reps. The model gets pulled off course by the noise and starts citing a refund policy from two years ago. Garbage in, garbage out — no matter how strong the model is.

Principle 5: Failures Must Be Visible and Recoverable ​

Statement: The harness must turn every failure — a tool error, a parse failure, invalid model output — into context the model can read and act on, and it must let the whole system resume from the point of failure instead of starting over.

Rationale: Agent loops fail by nature: the model hallucinates parameters that don't exist, the outside world times out, tests break. The problem isn't failure itself; it's what state the system is in after a failure. Factor 9 of 12-Factor Agents puts it bluntly: compact errors into the context window — errors aren't exceptions to be caught and swallowed; they're information to feed back to the model. Layer on factors 5 and 6 (unify execution state and business state; launch/pause/resume with simple APIs) and the conclusion follows: an agent's execution state should be serializable, so an interrupted loop can pick up from its checkpoint.

A minimal, workable failure-handling skeleton:

python
result = run_tool(call)
if result.ok:
    ctx.append(tool_message(result.output))
else:
    # Don't raise it, don't swallow it — compress it into context the model can use
    ctx.append(tool_message(
        f"ERROR: {result.summary}\n"      # one line saying what happened
        f"stderr (last 20 lines): {result.tail}"  # keep only the tail so the window doesn't blow up
    ))

# Execution state can be persisted at any time: interrupt → resume = reload ctx and keep looping
checkpoint.save(ctx, step=i)

Good example: When Aider can't apply the model's suggested edits (an edit-format match failure), it feeds the specifics of the failure back to the model so it can self-correct; and the whole edit-test-feedback loop can be interrupted by the user at any time, with the option to roll back to any git commit. Failures are both visible and survivable. See Case Study: Aider.

Bad example: A long-running agent's tool times out at step 40; the exception propagates to the top level, the process exits, and all forty steps of intermediate work — half-written files, already-verified findings — are lost, forcing a rerun from step 1. The worse version: the timeout is swallowed by except: pass, the model believes the command succeeded, and it pushes ahead another twenty steps on a false premise.

Principle 6: Design for Uncertainty ​

Statement: Design every model output as "a sample from a probability distribution," not "the return value of a program" — critical paths need validation, retries, and fallback paths, rather than an assumption that the model gets it right the first time.

Rationale: This is the most fundamental difference between harnesses and traditional software engineering. In a traditional system, a function call either returns the right result or throws; in an agent system, the model may return something that looks right but is actually wrong, and the same input may produce different output tomorrow. Hence: validate structured output against a schema instead of hard-parsing with regex; require confirmation for high-risk actions; and patterns like evaluator-optimizer are, at bottom, an admission that "the first generation probably isn't good enough," with validation built explicitly into the loop.

Good example: Generate code → run tests → feed failures back → fix. This loop exists in different forms in SWE-agent, OpenHands, and Aider. It doesn't assume the model writes correct code on the first try; it turns "correct" into a convergence process and puts the most deterministic judges available — the compiler and the test suite — inside the loop.

Bad example: Letting the model generate SQL and execute it directly against the production database, with no dry-run, no read-replica validation, no row-count cap. The first time the model writes DELETE FROM users WHERE id > 5 instead of DELETE FROM users WHERE id = 5, a design that got lucky becomes an incident.

Principle 7: Evals-Driven Iteration ​

Statement: Every change to the harness — swapping a prompt, adding a tool, reorganizing context — should be proven an improvement by evals, not by gut feel.

Rationale: A harness is the classic "pull one lever and the whole machine moves" system: change a single sentence in the system prompt and pass rates on task type A might rise 5 points while type B drops 10 — and you'll only notice A. Without evals, iteration is Brownian motion. Public cases bear this out again and again: SWE-bench for SWE-agent and OpenHands; the public coding benchmark Aider maintains for the evolution of its edit formats — the trade-offs among Aider's many edit formats (whole, diff, udiff, edit-fenced, and so on) were settled by running the benchmark, not by the designer's taste. This also explains why Evaluation & Observability isn't a nice-to-have: without traces, you can't even locate why an eval failed.

Good example: The improvement loop: collect real failure cases → turn them into an eval set → modify the harness → verify on the eval set → after shipping, fold new failures back into the eval set. The eval set grows along with the system and is one of the most valuable assets a harness team owns.

Bad example: "The new prompt feels more professional" — a quick look at the output on three hand-tested samples, then a full rollout. Two weeks later user-reported issues climb, so you roll back, and nobody can ever say what behavior that "more professional" prompt actually changed.

Start small

Evals don't need to be a full-blown platform on day one. Twenty real failure cases you labeled yourself, plus one script that runs the harness automatically and tallies pass rates, already puts you ahead of where most teams start.

Principle 8: Autonomy Needs Proportionate Guardrails ​

Statement: However much autonomy you give an agent, you need guardrails to match; guardrail strength should scale with an action's irreversibility and blast radius, rather than being spread evenly.

Rationale: Autonomy and risk are not linearly related. Let an agent roam a read-only codebase and the risk is near zero; let it run database migrations freely and a single mistake is a disaster. Sound guardrail design is tiered:

Low risk ◄────────────────────────────────► High risk
Read-only ops       Reversible writes       Irreversible / high blast radius
(read/grep)         (edit/test)             (rm / deploy / send messages)
   │                   │                        │
Full autonomy      Autonomy + audit        Human approval before execution

This follows straight from Factor 7 of 12-Factor Agents (contact humans with tool calls): "request human approval" is itself modeled as a tool call that the agent can initiate whenever it judges necessary, while the harness handles the state management of pausing, notifying, waiting, and resuming — the human isn't a passive bottleneck in the loop but a part explicitly designed into it.

Good example: Claude Code's default posture: reads are free, file writes prompt for confirmation, Bash commands are matched against rules to decide whether to allow or ask, and the user can tune it level by level — autonomy as a dial, not a switch.

Bad example: Both extremes are failure modes. One end: "for safety, confirm every step" — by the fiftieth click of "Allow," the user has stopped reading the content, and the guardrail degenerates into ceremony. The other end: "for smoothness, run everything with --dangerously-skip-permissions" — right up until the agent deletes a directory it deemed "surplus" during a cleanup task. The goal of guardrail design is to spend human attention on the few actions that genuinely require judgment.

Tensions Between the Principles ​

The eight principles don't always get along. Mature harness design is largely a matter of balancing between these tensions:

TensionOne endThe other endTypical trade-off point
Simplicity vs. capabilityPrinciple 1 (simplicity first)Long-running, open-ended tasks genuinely need agentsCover 80% with workflows; escalate to agents only for the remaining 20%
Autonomy vs. safetyPrinciple 8 (guardrails)Every added confirmation erodes the agent's ability to keep working continuouslyTier by risk; no one-size-fits-all
Context completeness vs. context cleanlinessPrinciple 4 (share the full trace)The window is finite, and noise is harmfulSubagents return distilled conclusions; compress history rather than keeping it all
Speed vs. reliabilityPrinciple 6 (validate and retry)Every validation step adds latency and costFast path for low-risk steps; evaluator-optimizer only where risk is high

There is no universal optimum

The optimal resolution of these tensions depends on your task mix, your cost of failure, and user expectations. A coding agent can tolerate a long validation loop (as long as the tests can run); a support agent can't. The principles are the map; your evals are the compass.

A Cheat Sheet ​

#PrincipleOne-line self-check
1Simplicity firstWhat specific problem does this layer of complexity actually solve?
2Model decides, harness constrainsIs this constraint written in the prompt, or in code?
3Few tools, done wellIf you deleted this tool, would task success rate really drop?
4Context is everythingIf you were the model seeing only this context, could you get this step right?
5Failures visible and recoverableWhere does this error go right now? Does the model know? Can we redo it?
6Design for uncertaintyIf the model gets this step wrong, who notices first?
7Evals-driven iterationDoes eval data back the last change you made?
8Autonomy with guardrailsIs this action irreversible? Who approved it?

Further Reading ​

References ​