Skip to content

Case Study: SWE-agent

At a glance SWE-agent from the Princeton team is a model example of “harness design as research”: it introduced the concept of the Agent-Computer Interface (ACI) and used commands and feedback formats purpose-built for language models to achieve results on SWE-bench far beyond non-interactive baselines.

Case Study: SWE-agent ​

SWE-agent is a paper by the Princeton NLP team (John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, Ofir Press, and others), released in May 2024 and accepted at NeurIPS 2024, together with its companion open-source system. The task setup is disarmingly simple: hand the agent a real GitHub issue and the entire code repository, and let it locate the problem, modify the code, and submit a fix on its own.

What makes it a classic case study is not that it "built yet another coding agent." It is that the authors turned something everyone took for granted into an object of study: the interface between the agent and the computer itself. The paper's title states the thesis outright — Agent-Computer Interfaces Enable Automated Software Engineering.

Seen through the harness lens, SWE-agent is the purest example of "academic-paper-driven harness design." It has almost no complex planner and no multi-agent orchestration; all of its energy went into a single question — what the model sees at each step, what actions it can take, and what feedback it receives when something fails. That is precisely the core concern of a harness.

Background: SWE-bench and the Ceiling of the Non-Interactive Approach ​

To understand SWE-agent, you first need to understand its benchmark, SWE-bench (also from the Princeton team; the ICLR 2024 paper). SWE-bench mines real issues and their corresponding merged PRs from 12 popular open-source Python projects (django, scikit-learn, sympy, among others), and asks a system to automatically produce patches that pass the associated tests.

Before SWE-agent, the dominant approach was to stuff the issue description and retrieved code snippets into the prompt in one shot and have the model emit a patch directly — a "non-interactive" approach. Its problems are obvious:

  • Real repositories run to hundreds of thousands of lines; no retriever can supply complete context;
  • The model cannot validate its own patch, so it never knows when it is wrong;
  • Fixing bugs is fundamentally an iterative process — locate, read the code, hypothesize, modify, run tests, adjust — and one-shot generation violates the very nature of that process.

SWE-agent took a different tack: instead of handing the model an exam paper, hand it a computer, and let it work in the repository the way a human engineer would. The question is — is the "computer" human engineers rely on (shell, editor, IDE) actually a good fit for language models? The paper's answer is no — and that is exactly where the ACI concept begins.

The Core Concept: ACI (Agent-Computer Interface) ​

ACI is the paper's most important conceptual contribution. It draws its analogy from HCI (Human-Computer Interface):

The paper's core argument

Humans benefit from purpose-designed software (IDEs, for example) on complex tasks. Language models are a new class of "end user" with their own capabilities and limitations, and they equally deserve interfaces designed specifically for them. Just as HCI studies how to help humans use computers well, ACI studies how to help LMs use computers well.

The weight of this analogy lies in what it does: it elevates "what should a harness look like" from an engineering tuning problem to a design problem that can be studied systematically. Why don't interfaces built for humans suit LMs? Drawing on the paper and on practice, we can list these concrete differences:

DimensionHuman engineerLanguage model
Visual processingComfortably skims syntax highlighting and scrolling filesReads token by token; long outputs waste budget and dilute attention
MemoryRetains file structure and variable names just seenHas only a context window; key information must be shown repeatedly
Error recoverySees a stack trace and intuits where things brokeNeeds formatted, parseable error feedback to correct reliably
Action granularityOne keystroke / one clickEach action is a full inference call — too fine a granularity and costs explode
Syntax toleranceA slipped keystroke gets flagged live by the editorOne malformed edit can throw off everything that follows

Take one example: vim is extremely efficient for humans and a disaster for LMs — modal state is invisible, edit results are not echoed back, and a single misstep leaves the model stuck. Flip it around, and an LM-friendly "editor" should be: stateless, echoing a window of results on every operation, with line numbers, showing a fixed maximum number of lines at a time, and running a syntax check automatically before an edit is committed. This is exactly what SWE-agent built.

The ACI's Concrete Design ​

SWE-agent's ACI consists of three parts: custom commands, feedback formats, and guardrails. We break down each part below, following the paper and the open-source repository's implementation.

Custom Commands: Redesigning the "Editor" for LMs ​

SWE-agent does not have the model manipulate files through a raw shell directly (though the shell is still available); instead it provides a set of custom commands. The representative ones:

CommandWhat it doesDesign intent
open <path> [line]Opens a file, showing a window centered on a given lineServes a fixed number of lines at a time (about 100 by default), with line numbers
goto <line>Jumps to a specific line in the current fileMakes "paging through a file" an explicit, predictable action
scroll_up / scroll_downScrolls the window up or downGranularity matched to the context window, not to human fingers
search_dir <term> [dir]Searches a directory for a termReturns a compressed list of filenames + line numbers, not raw grep output
search_file <term> [file]Searches within the current fileSame idea: output is a "hit list," not a flood of context
find_file <name> [dir]Finds files by nameA dedicated shortcut for a common operation
create <path>Creates a new file and enters edit modeFolds a multi-step shell operation into one action
edit <start>:<end> + replacement textReplaces the content of a given line rangePrecise, line-number-based editing, linted automatically before commit

Note the two design principles that run through all of it:

  1. Collapse high-frequency action sequences into single commands. For a human, "search, look at the results, open the file, jump to that line" is four small actions; for an LM, every hop is a full reasoning turn. The ACI compresses these sequences into one semantically explicit command, saving tokens and removing failure points.
  2. Commands are stateless and self-describing. Each command's inputs and outputs close within the current turn and depend on no hidden state (contrast vim's modes); the model never needs to "remember what mode the interface is currently in."

Feedback Format: Echoing the "Workspace State" After Every Action ​

The other half of the ACI is observation design. On every SWE-agent turn, what the model receives is not raw terminal output but a structured echo. The feedback from one open looks roughly like this:

text
[File: /repo/django/core/exceptions.py (120 lines total)]
  1: """
  2: Global Django exception and warning classes.
  3: """
  ...
 45: class FieldDoesNotExist(Exception):
 46:     """The requested model field does not exist"""
 47:     def __init__(self, model_field_name):
 48:         ...
 ...
(55 more lines above, 20 more lines below)

Key design points:

  • Fixed window size: no matter how big the file, only about 100 lines are shown at a time, with metadata on "how many lines above/below," so the model always has a sense of the file's scale;
  • Line numbers are always present: they anchor later line-based edits such as edit 45:52;
  • Long outputs are truncated: when command output runs long, head and tail are kept, the middle is cut, and the truncation amount is reported explicitly, so a single failing pytest run cannot blow up the context.

This is the polar opposite of the general-purpose shell's philosophy: a shell assumes users can handle output themselves; the ACI assumes output must be "tailored" for the model first.

Guardrails: Catching Errors Before Actions Take Effect ​

The clearest expression of the "design for the LM" spirit is SWE-agent's built-in guardrails. The canonical example is the lint check on the edit command: the replacement text the model submits goes through a syntax check first, and if it fails, the edit does not take effect — an explanatory error comes back instead:

text
Your proposed edit has introduced new syntax error(s).
Please read the error message carefully and then retry editing the file.

ERRORS:
- E999 IndentationError: unexpected indent (line 47)

This is how your edit would have looked if applied
-------------------------------------------------
[File: /repo/django/core/exceptions.py (120 lines total)]
 45: class FieldDoesNotExist(Exception):
 46:     """The requested model field does not exist"""
 47:         def __init__(self, model_field_name):
-------------------------------------------------

Behind this design is a lesson the paper states explicitly: let errors happen where they are cheap. If an edit containing a syntax error lands on disk, the model often does not notice until it runs the tests several turns later, and at that point tracing back "which step broke it" is expensive; intercepting the edit before it commits shortens the feedback loop to a single turn. The same treatment appears elsewhere: malformed command invocations get back a usage hint plus an example of the expected format, rather than failing silently.

The Overall Loop ​

Assembling these components, SWE-agent's runtime loop is a classic ReAct-style thought–action–observation cycle:

text
┌─────────────────────────────────────────────────────────────┐
│                         Task input                          │
│   GitHub issue description + repository environment +       │
│   ACI command manual (system prompt)                        │
└───────────────────────────┬─────────────────────────────────┘
                            ▼
┌─────────────────────────────────────────────────────────────┐
│   LM reasoning turn                                         │
│   Thought: I need to locate the validate() method mentioned │
│   in the error message first                                │
│   Action:  search_dir "def validate" django/db/models       │
└───────────────────────────┬─────────────────────────────────┘
                            ▼
┌─────────────────────────────────────────────────────────────┐
│   ACI execution layer (the harness itself)                  │
│   · Parse the action → invoke custom command / bash         │
│   · Trim output (windowing, truncation, line numbers)       │
│   · Guardrail checks (lint, format validation)              │
└───────────────────────────┬─────────────────────────────────┘
                            ▼
┌─────────────────────────────────────────────────────────────┐
│   Observation echoed back to the LM → next turn...          │
│   (until the model emits submit; the git diff is extracted  │
│   as the patch)                                             │
└─────────────────────────────────────────────────────────────┘

Notice that nothing on the model side involves special training — all of the "intelligence" lives in the ACI execution layer and the echo formats. That is precisely the harness's value proposition.

Results and What They Mean ​

The headline results reported in the SWE-agent paper (using GPT-4 Turbo as the base model):

  • Full SWE-bench test set: about 12.5% pass@1 (12.5% by the paper abstract's account, 12.47% by the repository README's) — the best score on SWE-bench at the time, and clearly above the best previous non-interactive methods;
  • SWE-bench Lite: about 23% (Lite is a 300-instance subset of the full set with a friendlier difficulty distribution);
  • HumanEvalFix: 87.7% pass@1 — also the best result at the time.

These numbers look unimpressive today, but their significance was never in the absolute values:

  1. It proved the "interactive agent route" is viable on real software engineering tasks. Before it, the mainstream on SWE-bench was one-shot retrieval + generation; after SWE-agent, the leaderboard came to be almost entirely occupied by "agent + custom tools" setups.
  2. It turned interface design into an ablatable experimental variable. The paper systematically analyzed how each ACI design choice (window size, error feedback, lint guardrails, and so on) affected behavior and success rate, giving later work empirical evidence for "which designs actually matter."
  3. It left behind a reusable conceptual framework. The term ACI has since been adopted by a great deal of follow-up work and has become standard vocabulary for discussing agent tool design.

Two caveats when reading these numbers

  • The paper's results are based on GPT-4 Turbo and the SWE-bench full set as it existed at the time; the leaderboard has since gone through SWE-bench Verified (an OpenAI human-curated 500-instance subset) and other revisions, and figures from different versions cannot be compared directly.
  • SWE-agent's score is a joint product of "harness + model." The paper's point is precisely to prove that under the same model, different ACI designs can produce results that differ by an order of magnitude.

Follow-Up Work: The Line That Grew Out of SWE-agent ​

SWE-agent is not an isolated system; it sits on a clear evolutionary chain, and that line is itself a chronicle of harness research:

  • SWE-bench (Oct 2023, ICLR 2024): the same team built the benchmark first, defining the task shape of "real issue → real patch";
  • SWE-agent (May 2024, NeurIPS 2024): introduced the ACI and established the interactive agent route;
  • EnIGMA (Sep 2024 preprint): carried the ACI ideas into cybersecurity. The team found that CTF (Capture The Flag) challenges require interactive programs such as gdb debuggers and remote service connections, while SWE-agent's ACI only supported the non-interactive mode of "send one command, get one output." EnIGMA introduced Interactive Agent Tools, letting an agent for the first time interact with programs that hold a persistent session, and it reached the best known results on three benchmarks — NYU CTF, Intercode-CTF, and CyBench — across 390 CTF challenges; on the full NYU CTF benchmark it solved 13.5% (27/200), more than three times the previous best agent. EnIGMA also documented, along the way, a new harness-level phenomenon — "soliloquizing": the model stops interacting with the environment and instead hallucinates an observation and keeps reasoning from it;
  • SWE-agent 1.0 / mini-SWE-agent (2025): the repository kept iterating. According to the project's announcements, SWE-agent 1.0 paired with Claude 3.7 set the best score on the full SWE-bench set in February 2025; mini-SWE-agent, released in July 2025, reaches 65% on SWE-bench Verified with about 100 lines of Python, and the project recommends it as a replacement for the original SWE-agent — a telling signal in itself: when the model is strong enough, the harness can be minimal (for more on this, see Model vs. Harness: Why the Harness Sets the Ceiling).

Paradigm Value: Harness Design as Research ​

SWE-agent's greatest contribution to agent research methodology was turning the harness from "invisible glue code" into a first-class object of study. It demonstrated a paradigm:

  1. Pose a design hypothesis: interfaces built for humans don't suit LMs (the analogy with HCI yields the ACI);
  2. State design principles: collapse high-frequency sequences into commands, keep them stateless and self-describing; window the feedback, keep line numbers, keep it parseable; intercept errors where they're cheap;
  3. Validate with controlled experiments: ablate each design dimension, quantify its effect on success rate, and analyze behavioral changes (for example, whether the model navigates the repository more efficiently);
  4. Distill the results into reusable assets: the ACI concept, the command design patterns, and the guardrail thinking were widely inherited by later work.

Set against many agent papers of the same era (focused on prompting tricks or planning algorithms), what makes SWE-agent distinctive is its subject of study: the contact surface between model and environment. That surface is precisely what every engineer building their own agent deals with daily, yet it is seldom discussed systematically.

Transferable Lessons for Engineers: Designing an ACI for Your Own Agent ​

Even if you aren't building a coding agent, SWE-agent's methodology transfers almost item for item. Here is an ACI design checklist you can put to use right away:

Command (tool) design

  • Watch your agent's high-frequency action sequences and merge stable 3–5 step chains into one composite tool (what search_dir is to grep + open + goto);
  • Keep tools stateless: each call's result is self-contained; never make the model "remember what mode the interface is in";
  • Fewer tool parameters beats more, and every parameter needs an explicit format example — malformed calls are the number-one source of agent failures.

Feedback (observation) design

  • Set a fixed budget for output: window it, keep head and tail, truncate the middle, and always report how much was cut;
  • Anchor the output: line numbers, IDs, paths, so later actions can reference things precisely;
  • Write error messages for the model, not for humans: say what went wrong, state the expected format, and give one correct example.

Guardrails

  • List the action categories that are hard to undo or hard to detect once they take effect (writing files, modifying databases, sending requests), and put cheap checks in front of them (syntax checks, schema validation, dry runs);
  • When a check fails, don't drop the action silently; return structured corrective feedback — every turn you cut from the feedback loop buys a real jump in success rate.

Experiment methodology

  • Treat every design decision as an ablatable variable: hold the model and task set fixed, change one ACI dimension at a time, and watch how success rate and turn count respond;
  • Beyond success rate, track behavioral metrics: average turns per task, proportion of invalid actions, proportion of repeated actions — these tend to expose interface problems before success rate does.

The one-sentence version

Treat the LM as your "user": design the interface for how it perceives (tokens, context window), design guardrails for how it fails (malformed output, hallucination, forgetfulness), then validate every design choice with ablation experiments.

Trade-offs ​

The ACI route is no free lunch either. In engineering terms, be clear-eyed about a few costs:

  • Custom interface vs. general capability: the more custom commands you add, the more the agent depends on your interface documentation, and porting to a new domain (like the CTF challenges EnIGMA faced) means redesigning; raw bash is awkward but universal. mini-SWE-agent's 100-line approach goes to the other extreme — nothing but bash, leaving everything to the model.
  • Guardrails vs. autonomy: intercepting guardrails prevent low-level mistakes but can also block "unconventional yet correct" operations. The paper's trade-off: guard the actions that are hard to recover from and slow to reveal their errors, leave the rest open.
  • Window size: too small an observation window, and the model burns extra turns paging through files; too large, and cost climbs while attention dilutes. The paper's roughly 100 lines is a sweet spot tuned by experiment — your task will need its own sweet spot.
  • The interface is a prior: the command set you design is, in essence, your encoding of a prior about "how this task should be solved." Get the prior right, and the agent goes twice as far for half the effort; get it wrong, and you have locked the agent inside a mistaken workflow. The more elaborate your ACI, the more vigilant you must be about this.

Further Reading ​

References ​