Skip to content

Model vs. Harness: Why the Harness Sets the Ceiling

At a glance The same model inside different agent harnesses (scaffolds) can perform several times apart on real tasks. Hard numbers from SWE-bench to argue that, at equal model capability, the harness is the dominant variable in agent performance — plus a look at how model progress rewrites harness design in return.

Model vs. Harness: Why the Harness Sets the Ceiling ​

Start with a phenomenon that once baffled the entire industry.

In early 2024, the best score on SWE-bench (a benchmark that measures AI bug-fixing ability on real GitHub issues) was 1.96% — fewer than 2 out of 100 real issues solved. In March 2024, Cognition launched Devin, self-reporting 13.86% — a sevenfold jump. A few months later, the SWE-agent paper hit 12.47% (full set) / 18% (Lite) with GPT-4 Turbo. By October 2024, Anthropic had pushed Claude 3.5 Sonnet to 49% using a deliberately minimal scaffold.

Models did improve during this period — but every one of those jumps was far larger than the model iterations behind them. What actually happened: people gradually figured out how to build a harness around the model.

This page's argument, in one sentence:

The core claim

At equal model capability, the harness is the dominant variable in agent performance. A second-tier model with a first-rate harness usually beats a first-rate model with a second-rate harness. And "the harness matters" is simultaneously true with "models will eventually internalize harness tricks" — understanding the interplay between these two forces is a fundamental skill of agent engineering.

1. Pin down the two terms ​

Model: a static weights file. Given a token sequence, it outputs a probability distribution over the next token. It has no loop, no tools, no memory, and no awareness that it's executing a task. The model is "potential."

Harness (scaffold): the executable system around the model — deciding what context the model sees at each step, which tools it can call, how its output gets parsed into actions, how it recovers from failure, and when to stop. The harness is the mechanism that "converts potential into capability."

text
┌────────────────────────────── Harness ───────────────────────────────┐
│                                                                      │
│   Context        Tool layer       Control       Memory     Human-   │
│   assembly                        loop                     in-loop  │
│  ┌──────────┐  ┌──────────┐  ┌──────────┐  ┌────────┐  ┌──────────┐  │
│  │ Repo map │  │ bash     │  │ plan →   │  │ Notes/ │  │ Approval │  │
│  │ Relevant │  │ file edit│  │ act →    │  │ retrieve│ │ Interrupt│  │
│  │ files    │  │ test run │  │ observe  │  │        │  │ /takeover│  │
│  │ Truncate/│  │          │  │          │  │        │  │          │  │
│  │ compress │  │          │  │          │  │        │  │          │  │
│  └────┬─────┘  └────┬─────┘  └────┬─────┘  └───┬────┘  └────┬─────┘  │
│       └──────────────┴──────┬─────┴────────────┴────────────┘        │
│                             ▼                                        │
│                    ┌─────────────────┐                               │
│                    │     LLM         │  ← only decides "what to      │
│                    │   (the model)   │     output next"              │
│                    └─────────────────┘                               │
└──────────────────────────────────────────────────────────────────────┘

The industry has a precise definition of "scaffold." In its SWE-bench technical post, Anthropic wrote that scaffolding "is responsible for constructing the prompt fed to the model, parsing the model's output to execute actions, and managing the interaction loop — the results of the model's previous action are incorporated into the next prompt" — and stated explicitly: "even with the same underlying model, agent performance on SWE-bench can vary dramatically based on the scaffolding." That sentence is one of the few pieces of consensus in agent engineering written into a vendor's own technical blog.

2. The evidence: how big is the harness gap at fixed model? ​

Asserting "the harness matters" persuades no one; look at the data. The four pieces of evidence below all satisfy the same control condition: same model (or same era), different harness.

Evidence 1: SWE-agent's ACI ablation (the cleanest controlled experiment) ​

The SWE-agent paper (NeurIPS 2024) ran one of the classic ablations in agent engineering: hold GPT-4 Turbo fixed and change only the interface between the model and the computer — what the authors call the agent-computer interface (ACI).

ConfigurationModelSWE-bench Lite resolved
Retrieval-augmented generation (RAG, non-interactive)GPT-4 Turbo~3.8%
Bare shell (raw terminal only)GPT-4 TurboWell below the ACI version
SWE-agent (carefully designed ACI)GPT-4 Turbo18.00% (54/300)

Two numbers matter:

  • Versus the RAG approach, the interactive harness delivered a 6.7× improvement in resolution rate (at 8–13× the cost);
  • Versus "just give it a bare shell," an ACI designed for the model (dedicated edit commands, formatted feedback, truncation protection for long outputs) delivered a 64% relative improvement.

Not a single weight changed. Only "the interface the model sees" changed, and performance rose 64%. That is the leverage of the harness.

What "designing the interface for the model" means

SWE-agent found that tools designed for humans (like vim, interactive commands) are a disaster for models; a model's strength is emitting complete, strictly formatted content in one shot. So the ACI turned "edit a file" into a dedicated edit command, forced truncation of long outputs with clear markers, and formatted error messages the way a model can easily parse. The interface's user changed from human to model — a paradigm shift for the whole field.

Evidence 2: the SWE-bench leaderboard is itself a giant natural experiment ​

A systematic analysis of the official SWE-bench leaderboard (Dissecting the SWE-Bench Leaderboards) found that most entries use the same small pool of models — on the Lite and Verified boards, Claude 3.5 Sonnet was used by 21 and 24 entries respectively.

What does that mean? Entries flying different teams' flags with scores spread across a dozen-plus percentage points often run on the same model. The score differences between them are almost purely harness differences: how context is packed, how tools are designed, how trajectories are controlled, how submissions self-check. SWE-bench accidentally became an instrument for measuring the state of the art in "harness engineering."

Evidence 3: Devin's 7× jump ​

When Cognition launched Devin in March 2024, it self-reported solving 13.86% of SWE-bench issues, at a time when the best unassisted score was 1.96%. Note the timing: Devin launched with a contemporaneous GPT-4-class model — no exclusive model. The jump from 1.96% to 13.86% came almost entirely from the harness: a full sandboxed environment, its own shell/editor/browser, multi-step planning and execution loops. The case was widely debated afterwards (including methodological criticism of its evaluation subset choice), but it established one fact: with the model fixed, an engineered harness can lift a benchmark score by an order of magnitude.

Evidence 4: Anthropic's "minimal harness" at 49% ​

The most intriguing evidence comes from Anthropic itself. In an October 2024 engineering post, they published Claude 3.5 Sonnet (new) scoring 49% on SWE-bench Verified (SOTA at the time was 45%) — then open-sourced their entire scaffold:

  • A prompt under 300 words;
  • Two tools: a Bash tool and a file-editing tool (str_replace_editor);
  • No planning module, no multi-agent, no elaborate retrieval pipeline — the loop runs until the model itself says "done" or the 200k context is exhausted.

The same post includes a table of four models running the same scaffold: Claude 3 Opus 22% → older 3.5 Sonnet 33% → previous SOTA 45% → new 3.5 Sonnet 49%. The table demonstrates both directions at once: with the harness fixed, model progress helps (22% → 49%), while across the same period the leaderboard's assorted scaffolds scattered same-model scores across a wide range.

Even more worth stealing is the "mistake-proofing" they built into the harness:

  • The file-editing tool requires absolute paths — because models frequently mangle relative paths after cd-ing away from the root;
  • Edits use string replacement (old_str → new_str), executed only when old_str occurs exactly once in the file; otherwise a clear error message is returned for the model to retry;
  • Tool descriptions read like product docs: every misuse pattern they observed in testing is pre-empted in the text.

Note the nature of these details

"Require absolute paths" is neither a model capability nor a prompt trick — it's eliminating a known model failure mode with deterministic code. This kind of work — "wherever the model tends to fall, install a railing" — is the daily bread of harness engineering, and even the model vendors do it seriously.

3. Three mechanisms by which the harness compensates for the model ​

Abstracting from the cases above, harness compensation for model weaknesses comes down to three paths.

1. Context management: solving "can't see" and "can't fit" ​

The model's first weakness is a finite context window, with attention degrading over very long contexts. A mid-sized codebase easily exceeds any model's window. The harness's answer is to put the model on an information diet:

  • Retrieval and navigation: repo maps, symbol indexes, expanding along call relationships — instead of dumping whole files in;
  • Truncation and summarization: force-truncate long tool outputs, compress old trajectories into summaries (Claude Code's compaction is the canonical example);
  • Subagent isolation: hand "exploration"-type tasks — high token, low information density — to subagents, returning only conclusions to the main context.

This practice now has a name: context engineering. See Context Engineering.

2. Error recovery: solving "makes mistakes" and "doesn't notice" ​

Models hallucinate, write broken syntax, and circle in dead loops. A bare model is helpless here; a harness can:

  • Immediate feedback loops: feed the real execution result of every tool call (compile errors, test failures, exit codes) straight back, letting the model self-correct next round — successful SWE-bench runs typically take dozens to hundreds of rounds;
  • Mistake-proof tool design: kill known failure modes at the tool layer, as Anthropic does (unique-match validation, absolute paths, timeout protection);
  • External verification: use deterministic programs (linters, type checkers, test suites) as the judge instead of trusting the model's self-assessment. "The model says it's fixed" doesn't count; "the tests are green" does.

3. Tool fallback: solving "can't do" ​

Some things models are inherently bad at: precise arithmetic, byte-exact manipulation of long strings, remembering state from three days ago, accessing real-time information. The harness principle: whenever deterministic code can do it, never let probabilistic generation do it:

  • Arithmetic goes to a calculator/code execution, not the model's mental math;
  • File diffs go to a str_replace tool with exact matching, not the model rewriting the whole file;
  • Cross-session state goes to the file system and memory modules, not the model's "recollection."

A rough but useful heuristic: whenever you catch yourself begging the model in a prompt to carefully do something mechanical ("please double-check character by character…"), that's usually a sign the job should be taken away from the model and turned into a tool.

4. The reverse arrow: how model progress rewrites the harness ​

If the story ended at "the harness matters," it would be a static conclusion. Real history is a two-way street: every model generation internalizes a batch of yesterday's harness tricks. This is often described as the digestion of the capability overhang — abilities that already existed in the model but needed external tricks to elicit get trained directly into the next generation's weights.

Several internalizations that have already happened:

Yesterday's harness trickToday's model capability
Writing "let's think step by step" in the prompt (CoT prompting)Reasoning models (o1, DeepSeek-R1, etc.) train long chains of thought directly into the model; thinking becomes a native output
Hand-writing ReAct loops, parsing Thought/Action/Observation with regexNative tool calling (function calling): structured output comes straight from the model, and the parser disappears
Carefully curating few-shot examplesWith stronger instruction following, zero-shot + clear instructions is usually enough
Multi-agent debate/voting to reduce errorsReliability gains from a single model + long thinking absorb much of the "self-debate" into one forward pass

Anthropic's advice in "Building effective agents" (December 2024) reflects exactly this dynamic view: the most successful implementations aren't piled up from complex frameworks but built from "simple, composable patterns"; and "give the model as much control as possible, keep the scaffolding minimal." The stronger the model, the more the harness should subtract — because elaborate pre-set workflows constrain a model that's already capable of planning its own path.

What this means for you

Don't treat the tricks in your harness as permanent assets. For every mechanism, ask: "Is this compensating for a model weakness, or doing something the model already does well?" The former (sandboxes, permissions, test verification, mistake-proof tools) holds value long-term; the latter (elaborate output-format constraints, hard-coded workflows) most likely becomes a liability the day the next model ships — it's not just useless, it actively limits the new model.

5. Rethinking the "just a wrapper" sneer ​

With the previous two sections in hand, we can revisit the question that's been argued since 2023: "isn't it just a wrapper?"

The sneer assumes all value lives at the model layer and the surrounding system is superficial packaging. History has handed down two opposite verdicts on it:

Verdict one: pure prompt wrappers did die. Products that only did "wrap the user's input in a carefully written prompt" (early writing assistants, role-play chats) were mowed down in waves by model iteration — because their value happened to be exactly "what the model temporarily couldn't do," and model progress specializes in digesting that kind of value.

Verdict two: system-level harnesses survived, and grew more valuable. The distance between Claude Code, Cursor, or Devin and "one API call" is sandbox isolation, permission systems, tool ecosystems, context assembly, observability, team workflows — these are software engineering assets. They don't depreciate with model generations; they appreciate (a stronger model raises the ceiling of the same harness).

So the right framing of the "wrapper" debate isn't "does the shell have value" but:

The litmus test

A harness's value ≈ the engineering assets it has accumulated that won't evaporate with model progress — environments, tools, data flywheels, workflow integrations — minus the temporary tricks it depends on that models will internalize. The thicker the former, the deeper the moat; the thicker the latter, the greater the danger.

The SWE-bench leaderboard is the bluntest footnote to this: dozens of teams, the same model, wildly different scores. If the shell didn't matter, the scores would cluster. They don't — because what's inside the shell was never just a prompt; it's complete, hard-to-replicate systems engineering.

6. For engineers: where to invest — the model layer or the harness layer ​

Down to decisions. Suppose your resources are limited — which layer do you invest in?

Your situationAdviceWhy
Product/business teamInvest firmly in the harness; use the best commercial model APIsThe model layer is a trillion-dollar arms race you can't win; the harness layer is the only place you can accumulate differentiation
Model training teamInvest in models, but feed harness data back into trainingAgentic capability is itself trained (tool use, long-horizon tasks); trajectories produced by harnesses are training fuel
ResearcherTreat the harness as a first-class research subject, not subsidiary engineeringACI, context management and related directions have already proven top-conference-worthy (SWE-agent is a NeurIPS paper)
Individual developerMaster 1–2 mature harnesses before building your ownUnderstanding Claude Code's / Aider's design decisions teaches you far more than writing a scaffold from scratch

Three more concrete rules of action:

  1. Exhaust the harness before switching models. Models iterate quarterly; your harness can iterate weekly. Leaderboard evidence shows harness optimization on a fixed model routinely delivers double-digit-percentage gains — far cheaper than waiting for the next model.
  2. Build evals before optimizing. Without your own task eval set, you can't tell whether an improvement came from the model or the harness — and you won't notice when "a new model turns an old trick into a negative optimization."
  3. Architect for model turnover. Design the "compensate for model weakness" parts of the harness (format constraints, step-by-step guidance) as removable modules; when a new model ships, your first instinct should be to try deleting them, not stacking new ones.

Further reading ​

References ​