Appearance
Model vs. Harness: Why the Harness Sets the Ceiling
Start with a phenomenon that once baffled the entire industry.
In early 2024, the best score on SWE-bench (a benchmark that measures AI bug-fixing ability on real GitHub issues) was 1.96% — fewer than 2 out of 100 real issues solved. In March 2024, Cognition launched Devin, self-reporting 13.86% — a sevenfold jump. A few months later, the SWE-agent paper hit 12.47% (full set) / 18% (Lite) with GPT-4 Turbo. By October 2024, Anthropic had pushed Claude 3.5 Sonnet to 49% using a deliberately minimal scaffold.
Models did improve during this period — but every one of those jumps was far larger than the model iterations behind them. What actually happened: people gradually figured out how to build a harness around the model.
This page's argument, in one sentence:
The core claim
At equal model capability, the harness is the dominant variable in agent performance. A second-tier model with a first-rate harness usually beats a first-rate model with a second-rate harness. And "the harness matters" is simultaneously true with "models will eventually internalize harness tricks" — understanding the interplay between these two forces is a fundamental skill of agent engineering.
1. Pin down the two terms
Model: a static weights file. Given a token sequence, it outputs a probability distribution over the next token. It has no loop, no tools, no memory, and no awareness that it's executing a task. The model is "potential."
Harness (scaffold): the executable system around the model — deciding what context the model sees at each step, which tools it can call, how its output gets parsed into actions, how it recovers from failure, and when to stop. The harness is the mechanism that "converts potential into capability."
text
┌────────────────────────────── Harness ───────────────────────────────┐
│ │
│ Context Tool layer Control Memory Human- │
│ assembly loop in-loop │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌────────┐ ┌──────────┐ │
│ │ Repo map │ │ bash │ │ plan → │ │ Notes/ │ │ Approval │ │
│ │ Relevant │ │ file edit│ │ act → │ │ retrieve│ │ Interrupt│ │
│ │ files │ │ test run │ │ observe │ │ │ │ /takeover│ │
│ │ Truncate/│ │ │ │ │ │ │ │ │ │
│ │ compress │ │ │ │ │ │ │ │ │ │
│ └────┬─────┘ └────┬─────┘ └────┬─────┘ └───┬────┘ └────┬─────┘ │
│ └──────────────┴──────┬─────┴────────────┴────────────┘ │
│ ▼ │
│ ┌─────────────────┐ │
│ │ LLM │ ← only decides "what to │
│ │ (the model) │ output next" │
│ └─────────────────┘ │
└──────────────────────────────────────────────────────────────────────┘The industry has a precise definition of "scaffold." In its SWE-bench technical post, Anthropic wrote that scaffolding "is responsible for constructing the prompt fed to the model, parsing the model's output to execute actions, and managing the interaction loop — the results of the model's previous action are incorporated into the next prompt" — and stated explicitly: "even with the same underlying model, agent performance on SWE-bench can vary dramatically based on the scaffolding." That sentence is one of the few pieces of consensus in agent engineering written into a vendor's own technical blog.
2. The evidence: how big is the harness gap at fixed model?
Asserting "the harness matters" persuades no one; look at the data. The four pieces of evidence below all satisfy the same control condition: same model (or same era), different harness.
Evidence 1: SWE-agent's ACI ablation (the cleanest controlled experiment)
The SWE-agent paper (NeurIPS 2024) ran one of the classic ablations in agent engineering: hold GPT-4 Turbo fixed and change only the interface between the model and the computer — what the authors call the agent-computer interface (ACI).
| Configuration | Model | SWE-bench Lite resolved |
|---|---|---|
| Retrieval-augmented generation (RAG, non-interactive) | GPT-4 Turbo | ~3.8% |
| Bare shell (raw terminal only) | GPT-4 Turbo | Well below the ACI version |
| SWE-agent (carefully designed ACI) | GPT-4 Turbo | 18.00% (54/300) |
Two numbers matter:
- Versus the RAG approach, the interactive harness delivered a 6.7× improvement in resolution rate (at 8–13× the cost);
- Versus "just give it a bare shell," an ACI designed for the model (dedicated edit commands, formatted feedback, truncation protection for long outputs) delivered a 64% relative improvement.
Not a single weight changed. Only "the interface the model sees" changed, and performance rose 64%. That is the leverage of the harness.
What "designing the interface for the model" means
SWE-agent found that tools designed for humans (like vim, interactive commands) are a disaster for models; a model's strength is emitting complete, strictly formatted content in one shot. So the ACI turned "edit a file" into a dedicated edit command, forced truncation of long outputs with clear markers, and formatted error messages the way a model can easily parse. The interface's user changed from human to model — a paradigm shift for the whole field.
Evidence 2: the SWE-bench leaderboard is itself a giant natural experiment
A systematic analysis of the official SWE-bench leaderboard (Dissecting the SWE-Bench Leaderboards) found that most entries use the same small pool of models — on the Lite and Verified boards, Claude 3.5 Sonnet was used by 21 and 24 entries respectively.
What does that mean? Entries flying different teams' flags with scores spread across a dozen-plus percentage points often run on the same model. The score differences between them are almost purely harness differences: how context is packed, how tools are designed, how trajectories are controlled, how submissions self-check. SWE-bench accidentally became an instrument for measuring the state of the art in "harness engineering."
Evidence 3: Devin's 7× jump
When Cognition launched Devin in March 2024, it self-reported solving 13.86% of SWE-bench issues, at a time when the best unassisted score was 1.96%. Note the timing: Devin launched with a contemporaneous GPT-4-class model — no exclusive model. The jump from 1.96% to 13.86% came almost entirely from the harness: a full sandboxed environment, its own shell/editor/browser, multi-step planning and execution loops. The case was widely debated afterwards (including methodological criticism of its evaluation subset choice), but it established one fact: with the model fixed, an engineered harness can lift a benchmark score by an order of magnitude.
Evidence 4: Anthropic's "minimal harness" at 49%
The most intriguing evidence comes from Anthropic itself. In an October 2024 engineering post, they published Claude 3.5 Sonnet (new) scoring 49% on SWE-bench Verified (SOTA at the time was 45%) — then open-sourced their entire scaffold:
- A prompt under 300 words;
- Two tools: a Bash tool and a file-editing tool (
str_replace_editor); - No planning module, no multi-agent, no elaborate retrieval pipeline — the loop runs until the model itself says "done" or the 200k context is exhausted.
The same post includes a table of four models running the same scaffold: Claude 3 Opus 22% → older 3.5 Sonnet 33% → previous SOTA 45% → new 3.5 Sonnet 49%. The table demonstrates both directions at once: with the harness fixed, model progress helps (22% → 49%), while across the same period the leaderboard's assorted scaffolds scattered same-model scores across a wide range.
Even more worth stealing is the "mistake-proofing" they built into the harness:
- The file-editing tool requires absolute paths — because models frequently mangle relative paths after
cd-ing away from the root; - Edits use string replacement (
old_str→new_str), executed only whenold_stroccurs exactly once in the file; otherwise a clear error message is returned for the model to retry; - Tool descriptions read like product docs: every misuse pattern they observed in testing is pre-empted in the text.
Note the nature of these details
"Require absolute paths" is neither a model capability nor a prompt trick — it's eliminating a known model failure mode with deterministic code. This kind of work — "wherever the model tends to fall, install a railing" — is the daily bread of harness engineering, and even the model vendors do it seriously.
3. Three mechanisms by which the harness compensates for the model
Abstracting from the cases above, harness compensation for model weaknesses comes down to three paths.
1. Context management: solving "can't see" and "can't fit"
The model's first weakness is a finite context window, with attention degrading over very long contexts. A mid-sized codebase easily exceeds any model's window. The harness's answer is to put the model on an information diet:
- Retrieval and navigation: repo maps, symbol indexes, expanding along call relationships — instead of dumping whole files in;
- Truncation and summarization: force-truncate long tool outputs, compress old trajectories into summaries (Claude Code's compaction is the canonical example);
- Subagent isolation: hand "exploration"-type tasks — high token, low information density — to subagents, returning only conclusions to the main context.
This practice now has a name: context engineering. See Context Engineering.
2. Error recovery: solving "makes mistakes" and "doesn't notice"
Models hallucinate, write broken syntax, and circle in dead loops. A bare model is helpless here; a harness can:
- Immediate feedback loops: feed the real execution result of every tool call (compile errors, test failures, exit codes) straight back, letting the model self-correct next round — successful SWE-bench runs typically take dozens to hundreds of rounds;
- Mistake-proof tool design: kill known failure modes at the tool layer, as Anthropic does (unique-match validation, absolute paths, timeout protection);
- External verification: use deterministic programs (linters, type checkers, test suites) as the judge instead of trusting the model's self-assessment. "The model says it's fixed" doesn't count; "the tests are green" does.
3. Tool fallback: solving "can't do"
Some things models are inherently bad at: precise arithmetic, byte-exact manipulation of long strings, remembering state from three days ago, accessing real-time information. The harness principle: whenever deterministic code can do it, never let probabilistic generation do it:
- Arithmetic goes to a calculator/code execution, not the model's mental math;
- File diffs go to a
str_replacetool with exact matching, not the model rewriting the whole file; - Cross-session state goes to the file system and memory modules, not the model's "recollection."
A rough but useful heuristic: whenever you catch yourself begging the model in a prompt to carefully do something mechanical ("please double-check character by character…"), that's usually a sign the job should be taken away from the model and turned into a tool.
4. The reverse arrow: how model progress rewrites the harness
If the story ended at "the harness matters," it would be a static conclusion. Real history is a two-way street: every model generation internalizes a batch of yesterday's harness tricks. This is often described as the digestion of the capability overhang — abilities that already existed in the model but needed external tricks to elicit get trained directly into the next generation's weights.
Several internalizations that have already happened:
| Yesterday's harness trick | Today's model capability |
|---|---|
| Writing "let's think step by step" in the prompt (CoT prompting) | Reasoning models (o1, DeepSeek-R1, etc.) train long chains of thought directly into the model; thinking becomes a native output |
Hand-writing ReAct loops, parsing Thought/Action/Observation with regex | Native tool calling (function calling): structured output comes straight from the model, and the parser disappears |
| Carefully curating few-shot examples | With stronger instruction following, zero-shot + clear instructions is usually enough |
| Multi-agent debate/voting to reduce errors | Reliability gains from a single model + long thinking absorb much of the "self-debate" into one forward pass |
Anthropic's advice in "Building effective agents" (December 2024) reflects exactly this dynamic view: the most successful implementations aren't piled up from complex frameworks but built from "simple, composable patterns"; and "give the model as much control as possible, keep the scaffolding minimal." The stronger the model, the more the harness should subtract — because elaborate pre-set workflows constrain a model that's already capable of planning its own path.
What this means for you
Don't treat the tricks in your harness as permanent assets. For every mechanism, ask: "Is this compensating for a model weakness, or doing something the model already does well?" The former (sandboxes, permissions, test verification, mistake-proof tools) holds value long-term; the latter (elaborate output-format constraints, hard-coded workflows) most likely becomes a liability the day the next model ships — it's not just useless, it actively limits the new model.
5. Rethinking the "just a wrapper" sneer
With the previous two sections in hand, we can revisit the question that's been argued since 2023: "isn't it just a wrapper?"
The sneer assumes all value lives at the model layer and the surrounding system is superficial packaging. History has handed down two opposite verdicts on it:
Verdict one: pure prompt wrappers did die. Products that only did "wrap the user's input in a carefully written prompt" (early writing assistants, role-play chats) were mowed down in waves by model iteration — because their value happened to be exactly "what the model temporarily couldn't do," and model progress specializes in digesting that kind of value.
Verdict two: system-level harnesses survived, and grew more valuable. The distance between Claude Code, Cursor, or Devin and "one API call" is sandbox isolation, permission systems, tool ecosystems, context assembly, observability, team workflows — these are software engineering assets. They don't depreciate with model generations; they appreciate (a stronger model raises the ceiling of the same harness).
So the right framing of the "wrapper" debate isn't "does the shell have value" but:
The litmus test
A harness's value ≈ the engineering assets it has accumulated that won't evaporate with model progress — environments, tools, data flywheels, workflow integrations — minus the temporary tricks it depends on that models will internalize. The thicker the former, the deeper the moat; the thicker the latter, the greater the danger.
The SWE-bench leaderboard is the bluntest footnote to this: dozens of teams, the same model, wildly different scores. If the shell didn't matter, the scores would cluster. They don't — because what's inside the shell was never just a prompt; it's complete, hard-to-replicate systems engineering.
6. For engineers: where to invest — the model layer or the harness layer
Down to decisions. Suppose your resources are limited — which layer do you invest in?
| Your situation | Advice | Why |
|---|---|---|
| Product/business team | Invest firmly in the harness; use the best commercial model APIs | The model layer is a trillion-dollar arms race you can't win; the harness layer is the only place you can accumulate differentiation |
| Model training team | Invest in models, but feed harness data back into training | Agentic capability is itself trained (tool use, long-horizon tasks); trajectories produced by harnesses are training fuel |
| Researcher | Treat the harness as a first-class research subject, not subsidiary engineering | ACI, context management and related directions have already proven top-conference-worthy (SWE-agent is a NeurIPS paper) |
| Individual developer | Master 1–2 mature harnesses before building your own | Understanding Claude Code's / Aider's design decisions teaches you far more than writing a scaffold from scratch |
Three more concrete rules of action:
- Exhaust the harness before switching models. Models iterate quarterly; your harness can iterate weekly. Leaderboard evidence shows harness optimization on a fixed model routinely delivers double-digit-percentage gains — far cheaper than waiting for the next model.
- Build evals before optimizing. Without your own task eval set, you can't tell whether an improvement came from the model or the harness — and you won't notice when "a new model turns an old trick into a negative optimization."
- Architect for model turnover. Design the "compensate for model weakness" parts of the harness (format constraints, step-by-step guidance) as removable modules; when a new model ships, your first instinct should be to try deleting them, not stacking new ones.
Further reading
- What Is an Agent Harness? — back to the definition, the boundaries of the term
- Anatomy of the Harness — a layer-by-layer dissection of the skeleton
- The Agent Loop — the harness's heart: the observe-think-act loop
- Context Engineering — the first battlefield where the harness compensates for the model
- Tools & MCP — how to design interfaces for models, not humans
- Case Study: SWE-agent — the full story of the ACI ablation
- Case Study: Claude Code — what a production-grade harness looks like
- Harness Design Principles — turning this page's conclusions into engineering discipline
References
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering (NeurIPS 2024) — the 12.47% (full) / 18.00% (Lite) results with GPT-4 Turbo, and the ablation showing the ACI's 64% relative gain over a bare shell
- Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet (Anthropic engineering blog, 2024-10) — the 49% score, the full minimal scaffold, the "scaffolding dramatically affects results with the same model" statement, and the mistake-proofing details
- Building effective agents (Anthropic, 2024-12-19) — the workflow/agent distinction, "simple composable patterns," and "give the model more control"
- Dissecting the SWE-Bench Leaderboards (arXiv:2506.17208) — statistics on models used by leaderboard entries (Claude 3.5 Sonnet appearing 21/24 times on Lite/Verified)
- Cognition: Introducing Devin (2024-03) — Devin's self-reported 13.86% on SWE-bench (note: evaluated on a random 25% subset of the full set; the methodology was challenged by Answer.AI and others — mind the caveats when citing)