Skip to content

SWE-agent

At a glance The Princeton team's classic 2024 work introducing the Agent-Computer Interface (ACI): tool interfaces should be designed for the model, not for humans. Interface optimization lifted SWE-bench solve rates from 3.8% to 12.5%, and mini-SWE-agent's 100 lines of code made the reverse case that scaffolds have a shelf life.

SWE-agent ​

SWE-agent is one of 2024's most influential open-source coding agents, but its value is not "yet another agent that can fix bugs"—it lies in proposing and validating a transferable methodology: the Agent-Computer Interface (ACI)—studying the interface between model and computer as seriously as HCI studies human-computer interaction. This idea later seeped into the tool design of nearly every coding agent, including the editing tool shapes you see in Claude Code and Cursor.

Meanwhile, the SWE-agent team's own follow-up—mini-SWE-agent, which cut the system down to 100 lines—wrote the second half of this story with their own hands: once models got stronger, how much value remains in a carefully designed scaffold? This page covers both halves.

1. Academic Background: The "Exam Writers' Agent," From the Same Stable as SWE-bench ​

SWE-agent came out of the Princeton NLP group, with core members including John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao (first author of the ReAct paper), Karthik Narasimhan, and Ofir Press.

Understanding SWE-agent starts with understanding its relationship to SWE-bench:

  • SWE-bench (arXiv:2310.06770, ICLR 2024): the benchmark the same team released in October 2023, drawing 2,294 real GitHub issues from 12 well-known Python repos (Django, Flask, scikit-learn, matplotlib, etc.) and requiring models to generate patches that pass the repos' existing tests. It was the first "close to real software engineering" hard benchmark in the coding agent field and remains the de facto standard (see Advanced: Evaluation).
  • SWE-agent (arXiv:2405.15793, NeurIPS 2024, v1 posted May 2024): the solving system built by the exam-writing team itself. The paper title is SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.

"Exam writers building an agent" brings two unique values. First, they know best what the benchmark tests and where the difficulty lies—SWE-bench's pain is not writing code but locating, understanding, and minimally modifying code in an unfamiliar repo of hundreds of thousands of lines. Second, they naturally had the motivation to make SWE-agent a research platform rather than a product: open source, configurable, convenient for other people's controlled experiments. The community did indeed end up using it as "the public foundation for scaffold research."

Score at publication: on the full SWE-bench test set, GPT-4 Turbo + SWE-agent achieved 12.5% pass@1, versus only 3.8% for the previous best non-interactive RAG system—a more-than-3x jump achieved without switching models. The number looks small today, but in the first half of 2024 it directly proved the overwhelming advantage of the "agent form factor" over "one-shot retrieval + generation," and it pulled the community's attention from prompt to interface.

A Timeline Coincidence with Devin

In March 2024, Cognition released the Devin demo claiming to solve ~13.8% of SWE-bench issues—closed and unreproducible. One month later, SWE-agent matched a comparable score fully open source, and the community called it "the open-source Devin." This is also one reason it spread so fast in its early days.

2. The Core Innovation: The Agent-Computer Interface (ACI) ​

The Problem: Tools Built for Humans Don't Work Well for Models ​

The paper starts from an observation: human engineers have IDEs, mice, and visual working memory; tools are designed around "human strengths and weaknesses." An LM agent is an entirely new class of user—its capability and defect profile is completely different from ours:

DimensionHuman EngineerLM Agent
Visual scanningStrong—a glance covers a screen of codeNone—reads segment by segment through a text window
Short-term memoryLimited but reliable (remembers what was just seen)Only what's in the context window; gone once scrolled out
Precise editingShaky hands, wrong line, but the IDE has undoLine-number arithmetic is error-prone, with no safety net
Parallel windowsCan open ten tabsSingle-threaded serial, one action per step
Cost of errorLow (humans notice immediately)High (erroneous output pollutes all downstream reasoning)

What happens if you just hand the model a bare shell (bash, type anything)? The paper's ablations answer: on a 300-instance subset of SWE-bench Lite, the same GPT-4 Turbo scored 10.7 percentage points lower with the bare-shell baseline than with the full ACI. In other words, at that model capability level, interface design contributed more than upgrading a whole model generation.

The ACI's Four Design Decisions ​

SWE-agent's ACI consists of a set of custom commands and feedback formats. The core decisions:

1. A dedicated file viewer instead of cat. Catting a large file instantly floods the context, and most of the content is irrelevant. SWE-agent's viewer shows only about 100 lines at a time (with line numbers), with scroll_up / scroll_down / goto commands for paging. Ablations show that shrinking the window to 30 lines costs about 3.7 percentage points (too little information; paging back and forth wastes steps), while showing the whole file makes the model "lose focus." 100 lines was the empirical sweet spot of the GPT-4 era.

2. Restricted editing commands + a linter gatekeeper. Editing doesn't use sed/vim but a dedicated edit command: specify a line range and the replacement content. The key design is running a linter (e.g., pyflakes) automatically after every edit; a syntax error rejects the edit outright and returns the error message—bad code never lands on disk. This exploits the model's trait of "strong immediate error correction, weak self-awareness": error prevention is built into the interface itself. In the ablations, removing the editing interface (linter included) was the single largest loss, about 7.7 percentage points.

3. Restrained search feedback. A whole-directory string search returns only "which files have matches," not the context of every match. The team found that showing the model too many matching snippets confused it—the information density of feedback is a tuned parameter, not "the more the better."

4. Never return empty output. When a command succeeds but produces no output, return Your command ran successfully and did not produce any output. instead of an empty string. An empty response makes the model suspect the command didn't run, and it starts repeating attempts or doubting the environment—a trap a human would essentially never hit but models hit daily.

┌───────────────────────── The ACI Interaction Loop ─────────────────────────┐
│                                                                            │
│  LM ──► command (open/edit/scroll_down/search_dir/...)                     │
│           │                                                                │
│           ▼                                                                │
│      ACI layer: parse → validate → execute → linter gatekeeping            │
│           │                                                                │
│           ▼                                                                │
│      Structured feedback (100-line window / file lists / explicit          │
│      empty-output message)                                                 │
│           │                                                                │
│           └────► appended to history, feeding the next reasoning step      │
└────────────────────────────────────────────────────────────────────────────┘

A typical interaction looks like this (simplified):

Model:  open "django/core/management/base.py"
Env:    [File: /repo/django/core/management/base.py (612 lines total)]
        1  import os
        2  import sys
        ...
        100      self.stdout = OutputWrapper(stdout or sys.stdout)

Model:  edit 88:92
        <<replacement code>>
Env:    The file ... has been edited. Review the changes and make sure
        they are as expected (correct indentation, no duplicate lines, etc).
        Edit the file again if necessary.

Model:  python -m pytest tests/admin_scripts/tests.py -x
Env:    3 failed, 41 passed in 12.3s
        FAILED tests/admin_scripts/tests.py::TestRunserver::test_bug_xxxx
        ...

Note the tone of the feedback text—"Review the changes and make sure they are as expected"—a reminder written as if for a human, but the real reader is the model. The feedback copy is itself part of prompt engineering; every word shapes the model's next action. This is expanded more systematically in Components: Prompt Engineering.

What the ACI Is Not ​

To avoid misreading, three boundaries:

  • ACI ≠ tool count. SWE-agent's command set is not large; it wins because each command's feedback format was tuned over and over. Piling up 50 unpolished tools is the classic antipattern.
  • ACI ≠ hiding complexity. It doesn't "automatically fix" anything on the model's behalf; it just presents state in the form the model consumes most easily. The model still has to locate the bug and decide which lines to change.
  • The ACI's effectiveness depends on the underlying model. The paper also tested Claude, GPT-3.5, and others, concluding that the benefits of interface design transfer broadly, but the specific parameters (like window size) were tuned for GPT-4 and need retuning for other models.

The ACI's Essence Is "Error Prevention + Feedback Engineering"

Note that none of these four decisions is "teach the model stronger abilities"; all of them either reduce the probability of model error or make feedback exactly sufficient. They map one-to-one onto human UX design principles (affordance, constraints, immediate feedback)—which is precisely why the paper borrowed HCI concepts. When designing tools, asking "where is the model most likely to screw up" pays far more than asking "what more capability can I give the model." More general principles in Components: Tools & MCP.

3. Score Evolution: From 12.5% to Being Succeeded by Its Own "Mini" Version ​

The SWE-agent family's SWE-bench scores fall roughly into three phases:

DateSystemScoreNotes
2024-05GPT-4 Turbo + SWE-agent12.5% (full set)Paper figure; prior non-interactive SOTA was 3.8%
2024–2025Claude 3.5/3.7 Sonnet + SWE-agentRepeatedly topping the Lite/Verified boardsSWE-agent became a standard scaffold for each vendor's new models
2025-02SWE-agent v1.0.1Officially announced SOTA on SWE-bench FullRelease notes specifically note that gains on the non-Lite/Verified portion of the Full set were larger—"evaluating only Lite/Verified tells an incomplete story"
2025-07 onwardsVarious models + mini-SWE-agentClaude 4 Sonnet 64.9% (Verified)A 100-line scaffold matching heavyweight systems
2026-02Claude 4.5 Opus (high) + mini-SWE-agent v2.0.076.8% (Verified, $0.75 per instance)Currently in the top band of the official SWE-bench board

Two readings worth remembering. First, from 12.5% to 76.8%—a 6x rise in two years—but along the way the models changed several generations while the scaffold got simpler, not more complex; that is the most important lens for this phase. Second, the Verified subset is being "solved to death": after top models broadly passed 70%, the community has been moving to harder new benchmarks like SWE-bench Pro; when reading leaderboards, distinguish dataset versions (SWE-bench Full / Lite / Verified / Pro are four different things).

4. The Research Value of the Scaffold: A Natural Controlled Experiment ​

The SWE-agent family's biggest contribution to the research community may not be any particular system but the controlled-variable test bench it provides.

The official SWE-bench leaderboard has a uniquely designed track: "SWE-bench (bash only)"—every entry runs the same mini-SWE-agent scaffold, with the underlying model as the only variable. As of 2026, this board spans from low to high:

Same scaffold (mini-SWE-agent), only the model changes:
  Llama 4 Scout        9.1%
  GPT-4o              21.6%
  Gemini 2.5 Pro      53.6%
  Claude 4 Sonnet     64.9%
  GPT 5.2             72.8%
  Claude 4.5 Opus     76.8%   ← 2026-02

This is a very scarce setup in agent research. Compare how other products chase leaderboards—each with its own scaffold, prompts, retry strategies, and cost budgets, none public and all different—and you simply cannot answer "is the high score from a strong model or a strong harness?" A fixed-scaffold board directly yields the net capability ranking of models; conversely, fixing the model and swapping scaffolds (as in the SWE-agent paper's ablations) yields the scaffold's net contribution. Measuring the two axes of "model × scaffold" separately is the methodological legacy of the SWE-agent line of work.

Direct corollaries for practitioners:

  • When evaluating any agent system, always report the "model version + scaffold version" pair; a single number alone has no comparability.
  • To compare two models fairly, put them into the same minimal scaffold (say, fork mini-SWE-agent directly) instead of attaching each to its own home harness.
  • This is also what Advanced: Evaluation keeps stressing: the harness is part of the system under test, not part of the test equipment.

5. Architecture and Code Organization ​

The SWE-agent repo (github.com/SWE-agent/SWE-agent) is a codebase designed for "running experiments" rather than "running a product." In the v1.x era the structure was roughly:

sweagent/
├── agent/          # Agent main loop: step logic, history processors
│                   # compresses/trims overlong trajectories into the context window
├── environment/    # SWEEnv: spins up a container with the repo inside a Docker sandbox
│                   # provides a bash session, file read/write, command allowlist
├── tools/          # The ACI itself: one directory per tool
│                   # (file viewer, edit+lint, search, etc.; interfaces defined in YAML)
└── run.py          # Entry point: given an issue → start env → run loop → save trajectory

A few design points useful to self-learners:

  • Tools are configured, not hard-coded. Each ACI tool declares its command name, argument format, docstring, and invocation style in YAML. To experiment with a new interface, add a directory and write a YAML file—which is why the repo README describes itself as "built to make it easy to invent new ACIs."
  • History processors are first-class citizens. How to stuff an exploded conversation history back into the context window is abstracted into pluggable processors (trimming, summarization, folding observations, etc.). This idea is generalized further in Components: Context Engineering.
  • Environment execution was extracted into SWE-ReX. Sandbox execution (Docker/cloud/local) later became a separate repo, SWE-ReX, reused by SWE-agent and third parties—decoupling the "execution runtime" from "agent logic," a layering that later agent frameworks widely adopted.
  • Trajectories are fully persisted to disk. Every step's inputs and outputs are stored as structured files, with a trajectory viewer included. For observability and failure analysis, this beats any monitoring platform for directness.

⚠️ Important status note: the official docs now state clearly that SWE-agent is in maintenance-only mode, superseded by mini-SWE-agent. For learning the latest code or running the latest leaderboards, look at mini-SWE-agent; the SWE-agent repo's value is in paper reproduction and ACI/tool-set ablation experiments. This is a positioning matter, not a quality judgment.

6. Follow-On Work: Self-Disruption by the Team's Own Hands ​

The team's follow-up moves form the most instructive part of this case—they proactively answered the question every scaffold author should fear: if models keep getting stronger, how much are my carefully tuned interfaces still worth?

mini-SWE-agent: The 100-Line Answer ​

In July 2025, the team released mini-SWE-agent with a design philosophy exactly opposite to SWE-agent's:

  • Only bash, as a single tool—it doesn't even use the model's native tool-calling API, relying entirely on a text protocol.
  • A fully linear history, appending each step directly with no trimming or summarization—the trajectory is the messages.
  • Each action executes independently via subprocess.run, with no stateful shell session; sandboxing is just swapping subprocess.run for docker exec.
  • The core agent class is about 100 lines of Python.

The result: Claude 4 Sonnet + mini-SWE-agent scored 64.9% on SWE-bench Verified, on par with the heavyweight systems of the time; the v2 released in early 2026 pushed Claude 4.5 Opus to 76.8%, and the official README claims over 74% across the whole Verified set. The team's explanation is candid: many interface designs of the SWE-agent era (file viewer, edit guard) were compensating for GPT-4's capability deficits; once models got strong enough to write their own sed and control their own output length, those compensations became shackles and token waste.

mini-SWE-agent is still actively evolving: it has reached v2, is officially used as the standard scaffold for the SWE-bench (bash only) board, and its README says it is used for evaluation and training (as the base scaffold for RL/fine-tuning—simple enough that models don't overfit to a particular harness) by companies including Meta, NVIDIA, IBM, and Anyscale and by several universities. The project also claims adoption by Ramp's SWE-Bench evaluation and DataCurve's DeepSWE evaluation (note: this is the project's own claim; verify before citing).

It is also a directly usable command-line tool and Python library. Installation (from the official README):

bash
pip install mini-swe-agent
mini   # starts the CLI; work directly in your local terminal

The Python binding is equally minimal—getting the agent ready takes just these lines:

python
from minisweagent.agents.default import DefaultAgent
from minisweagent.models.litellm_model import LitellmModel
from minisweagent.environments.local import LocalEnvironment

# Model via litellm; environment is a local shell; swap in Docker by changing the Environment
agent = DefaultAgent(
    LitellmModel(model_name="anthropic/claude-sonnet-4-5"),
    LocalEnvironment(),
)
agent.run("Write a sudoku game")

Side by side, the division of labor between SWE-agent and mini-SWE-agent is clear:

SWE-agentmini-SWE-agent
StatusMaintenance modeActive development (v2)
ToolsCustom ACI command set (YAML-configurable)Bash only
History managementPluggable history processorsPurely linear appends
Best forACI/tool-set ablation research, paper reproductionDaily CLI use, leaderboards, RL/fine-tuning base scaffold
Code sizeA fully engineered repoCore agent class ~100 lines

The Academic Line: SWE-smith and the Data Side ​

The team's other line is SWE-smith (arXiv:2504.21798, NeurIPS 2025 Datasets & Benchmarks Spotlight): automatically turning any GitHub repo into an "SWE-gym"-style training environment, mass-producing verifiable software engineering tasks for generating tens of thousands of training trajectories. The logic is clean: SWE-bench (evaluation) → SWE-agent/mini (scaffold) → SWE-smith (data)—one team filled in all three infrastructure blocks of coding agent research.

On "Commercialization"

As of August 2026, SWE-agent/mini-SWE-agent remain academically led open-source projects (MIT license) with no separate commercial company; the SWE-bench website's acknowledgments include Open Philanthropy, AWS, Modal, a16z, OpenAI, Anthropic, and others. For updates on team members' moves or company formation, rely on the official site and repo announcements, not secondhand rumors.

7. Transferable Lessons from the ACI Idea ​

The most valuable takeaways from this case, in priority order:

  1. Treat the agent as the user; do UX for tools. Agents in any domain (database ops, data analysis, customer-service backends) should be asked the questions SWE-agent asked: whose intuition is this interface designed around? Is there noise in the feedback the model doesn't need? Empty outputs, error formats, pagination granularity—all are performance variables. This principle has generalized into tool design practice in the MCP era; see Components: Tools & MCP.

  2. Do error prevention before capability enablement. A linter gatekeeper (rejecting bad edits outright) beats any prompt "teaching the model to write better code." Audit your toolset: which errors can be structurally eliminated at the interface layer instead of being left to the model's self-discipline?

  3. Scaffolds have a half-life; treat them as consumables. The ACI's sweet spot (100-line file views, etc.) is a function of model capability; when the model changes generation it needs retuning—or, as mini-SWE-agent showed, outright deletion. Architecturally, keep the scaffold a replaceable thin layer; don't let it grow into an irreplaceable load-bearing wall. This is the empirical footnote to the "keep it simple" principle in Practice: Design Principles.

  4. Ablation experiments are the discipline of scaffold development. The most underrated part of the SWE-agent paper is that Table 2: every design decision has its own controlled number. Impose the same discipline on your own agent—before adding a tool, plan how you'll measure its marginal contribution.

  5. Benchmark authors building agents is the fastest path to understanding the benchmark. Conversely, when reading a benchmark paper, also read the exam-writing team's own baseline agent; you'll know where the scores' "water content" and ceiling lie sooner than 90% of users.

To reproduce this idea hands-on, the cheapest path is to fork mini-SWE-agent or write a 100-line version following Practice: Build Your Own Agent, then run a "linter vs. no linter" comparison on 10 SWE-bench Verified instances—within a day you'll have muscle-memory-level understanding of the ACI.

References ​