Skip to content

A Brief History

At a glance From ReAct and Toolformer in 2022 to the harness discourse of 2026 — every turning point in the five-year evolution of the agent harness: what people thought the solution was at the time, and what was later falsified or confirmed.

A Brief History ​

Nobody invented the agent harness. It was forced into existence by a chain of failed solutions. Every generation of agent systems answered the same set of questions — what goes in the context, how tools plug in, how the loop turns, who holds control — and each era's answers were partly overturned by the next.

This short history walks the timeline. At each stop we ask two questions: What did people at the time believe the solution was? And what was later falsified — or confirmed? These questions matter more than the events themselves, because several of those "what people believed" items are being re-marketed, with fresh packaging, right now.

text
2022.10  ReAct paper              alternating reasoning + acting; the proto agent loop
2022.10  LangChain open-sourced   the first harness framework
2023.02  Toolformer               tool use becomes a trainable capability
2023.03  GPT-4 / ChatGPT API      the economics of autonomous agents close
2023.03  BabyAGI / AutoGPT        the first boom and the first disillusionment
2023.06  OpenAI function calling  tool calling becomes an API primitive
2023.10  SWE-bench                the first reproducible yardstick
2024.03  Devin                    the product manifesto of the "AI software engineer"
2024.04  SWE-agent / OpenDevin    the open-source counterattack; ACI is coined
2024.11  Cursor agent mode / MCP  form converges + protocol standardization (same week)
2024.12  Building Effective Agents  the workflow-vs-agent debate gets its ruling
2025.02  Claude Code              the terminal-native coding agent takes final form
2025.03  METR long-horizon study  the frontier pivots to long-task reliability
2025.06  context engineering      the discourse shifts: prompt engineering exits
2025.09  Claude Agent SDK         "harness" becomes an official term

2022: the foundations — ReAct and Toolformer ​

ReAct: the loop is born ​

In October 2022, Yao et al. published ReAct (accepted to ICLR 2023): have the model alternate between generating a reasoning trace and an action; the action acts on the external environment, the observation flows back in, and the model keeps reasoning.

What people thought the solution was: a prompting technique — put a few "thought–action–observation" examples in the prompt and the model would follow suit. Like chain-of-thought, it was filed under prompt engineering.

What was later confirmed: ReAct's real legacy isn't those prompt examples but the alternating structure itself — "the model emits an action, the harness executes it, the observation flows back, the model decides again." The main loop of every coding agent today (see The Agent Loop) is a thicker version of this structure. What got internalized: the hand-written reasoning traces in ReAct were, in the era of reasoning models (o1, R1), trained into the model itself and no longer need prompt engineering to elicit — the first demonstration of "prompt tricks absorbed by the model," a pattern that repeats throughout this history.

Toolformer: the other road to tool use ​

In February 2023, Meta published Toolformer: teach the model, self-supervised, when to call APIs like calculators and search engines — a call counts as useful if it reduces the prediction loss of subsequent tokens, and only useful calls make it into the training data.

What people thought the solution was: two roads side by side — the ReAct camp said tool use comes from the harness (prompt + parsing + execution); the Toolformer camp said it comes from training (burning the capability into the weights).

What was later confirmed: the answer is "both." Tool-calling capability was indeed trained into models (post-function-calling models no longer need ReAct-style examples to call tools), but tool integration, execution, result feed-back, and failure handling live outside the model forever. Toolformer won the "where does the capability live" question; ReAct won the "how do you build the system" question.

A quieter thread from the same period

In October 2022, Harrison Chase open-sourced LangChain — at first just a Python library for chaining LLM calls. It was the first time "building scaffolding for a model" became framework engineering. Later it drew heavy criticism for over-abstracting, but the underlying judgment — "the harness deserves a framework" — was inherited by the whole industry.

2023: the AutoGPT summer — boom and bust ​

The fuse was economics ​

The ChatGPT API opened on March 1, 2023; GPT-4 launched on March 14. Together they meant individual developers could, for the first time at an affordable price, put a sufficiently strong model into a loop and burn tokens. Within weeks, autonomous agents sprang up everywhere:

  • March 28: Yohei Nakajima released BabyAGI — roughly a hundred-odd lines of Python implementing a closed loop of "create task → execute → re-prioritize";
  • March 30: Toran Bruce Richards released AutoGPT — give GPT-4 a goal and let it prompt itself onward: browsing the web, reading and writing files, executing code. It became the fastest-growing project in GitHub history at the time, blowing past 100k stars within weeks.

What people thought the solution was: full autonomy — give it a goal, walk away, and the agent plans, executes, and self-corrects until done. Two accompanying beliefs: first, that planning ability was overestimated — people assumed models could draft and execute multi-step plans; second, that prompt engineering was the core skill — write a good system prompt and most of the problem was solved.

What was falsified: nearly all users hit the same failure modes fast — the agent grinding on one failed step for half an hour, forgetting the original goal after context overflow, the plan becoming a sunk cost, API bills running wild. BabyAGI-style "regenerate the task queue after every step" meant the agent was always adjusting the plan instead of working (that lesson is dissected in Planning & Task Decomposition). By the second half of 2023, "AutoGPT-style fully autonomous agents" was, as a product roadmap, effectively bankrupt.

What was confirmed: three things survived that summer — the loop was right (a goal-driven agent loop really is the correct abstraction); tools were right (web access, files, and code execution really are capability amplifiers); and the human shouldn't walk away — full unattended autonomy was falsified, and that directly seeded the permission and approval designs of every later product (see Permissions, Safety & Human-in-the-Loop).

Two quiet mid-year events ​

Two things happened during the disillusionment phase that mattered more than AutoGPT itself:

June 2023: OpenAI shipped function calling. Tool calling went from a folk protocol — "agree on a format in the prompt, write regex parsing in the harness" — to a structured, API-level primitive. The reliability problem of harnesses parsing model output was absorbed by the vendor, and the tool system (see Tools & MCP) finally had a standard interface to build on. MCP is where that line ends.

July 2023: Lost in the Middle was published. The paper showed that models use information in the middle of a long context markedly worse than at the head and tail. It was early hard evidence that "more context is not always better" — the theoretical prelude to the context-engineering discourse two years later (see Context Engineering).

2024: SWE-bench and the coding-agent breakthrough ​

The yardstick precedes the breakthrough ​

In October 2023, the Princeton team released SWE-bench: real GitHub issues as tasks, with passing the real unit tests as the criterion. Its significance wasn't difficulty but reproducibility and comparability — for the first time, agent systems had a commonly accepted ruler. That ruler changed the dynamics of evolution: the harness went from "demo engineering" to "a thing that can be systematically optimized."

March–April 2024: six weeks that set the landscape ​

  • March 12: Cognition launched Devin, billed as "the first AI software engineer," scoring 13.86% on SWE-bench (the best GPT-4 baseline at the time was single-digit). The demo went viral — and quickly drew accusations of selective demonstration.
  • April: Princeton open-sourced SWE-agent (NeurIPS 2024), solving 12.29% of the full SWE-bench with GPT-4 — the same order of magnitude as Devin's launch number, but with fully public code. The paper coined the Agent-Computer Interface (ACI) concept: commands and feedback formats designed for the model affect scores as much as prompt engineering does. Tool interface design itself was elevated to a research subject — the academic legitimization of harness engineering (see Case Study: SWE-agent).
  • In the same window, weeks after Devin's launch, the open-source community produced OpenDevin (later renamed OpenHands), turning the "autonomous software engineer" into a public laboratory.

What people thought the solution was: Devin's narrative — package the agent as a "remote employee" with its own shell, browser, and editor, with humans as the client.

What was falsified: the fully anthropomorphized "AI employee" packaging. Developers didn't want a black-box coworker they had to hand control to; they wanted a tool embedded in their own workflow, interruptible at any moment. Devin's high pricing and the gap in early real-world testing (third-party evaluations from Answer.AI and others showed high real-task failure rates) proved the point.

What was confirmed: coding was the first vertical where the agent loop genuinely worked — and for harness-level reasons: the code environment is executable, verifiable, and gives instant feedback, so the "observe" phase of the loop is extremely high quality. That same year, general web-operation agents still hadn't broken through — the counterfactual proof that models hadn't magically learned to work; the code environment just happens to suit the loop.

The year-end ruling: workflow vs. agent ​

In December 2024, Anthropic published Building Effective Agents, drawing the line under a year-long debate:

A workflow is a system where "LLMs and tools are orchestrated through predefined code paths"; an agent is a system where "the LLM dynamically directs its own processes and tool usage."

The post's stance — start with the simplest thing that works; add complexity only when demonstrably necessary — became the default constitution of harness design afterwards. In effect it pronounced the failure of 2023-style heavyweight multi-agent frameworks and previewed 2025's convergence: a simple loop + good tools + good context beats an intricate predefined pipeline.

Late 2024–2025: form converges, discourse shifts ​

Two signals in the same week ​

In the last week of November 2024, two seemingly unrelated things happened:

  • November 24: Cursor shipped 0.43, bringing early agent mode to Composer: it picks its own context and runs its own terminal commands (see Case Study: Cursor);
  • November 25: Anthropic released the Model Context Protocol (MCP), an open standard unifying "how applications expose tools, data, and context to models" — often described as the USB-C port for AI applications.

One is a product form, the other an integration protocol — pointing at the same thing: the integration layer for tools and context is standardizing and generalizing. Three months later (February 24, 2025), Claude Code launched in research preview, nailing that converged form into the terminal (see Case Study: Claude Code).

The converged shape ​

By mid-2025, mainstream coding agents were strikingly uniform in design — a mirror image of 2023's AutoGPT:

text
      2023 AutoGPT                  2025 converged form
┌──────────────────────────┐   ┌──────────────────────────┐
│ Self-prompting loop,     │   │ Single agent loop +      │
│ no human present         │   │ human in the loop        │
│ Grand multi-step plans   │   │ Lightweight todo list,   │
│                          │   │ revised on a rolling basis│
│ Huge bespoke toolset     │   │ A few general-purpose    │
│ (dozens of plugins)      │   │ primitives (read/write/  │
│                          │   │ grep/bash)               │
│ Vector DB as "memory"    │   │ One markdown file +      │
│                          │   │ compaction               │
│ No permission concept,   │   │ Tool allowlist +         │
│ auto-execute everything  │   │ per-action approval      │
│ Replan every step        │   │ Model decides next step  │
└──────────────────────────┘   └──────────────────────────┘
  Intelligence lives in the       Intelligence stays with
  harness (the model can't        the model; the harness
  carry it)                       is kept thin and stable

The right column is today's default answer: keep the harness thin and stable, leave judgment to the model, leave control to the human. Notably, every item in the left column was considered "more advanced" at the time — the direction of evolution is subtraction, not addition.

The discourse shift: from prompt engineering to context engineering ​

On June 18, 2025, Shopify CEO Tobi Lütke posted: "I prefer the term context engineering — it better describes the core skill: the art of providing all the context needed to make the task solvable by the LLM." A week later Andrej Karpathy seconded it: the real work in production LLM applications is "the delicate art and science of filling the context window with just the right information."

The rename was not a rhetorical game. Prompt engineering presumes "the model answers a question in a vacuum"; context engineering admits "the model makes a chain of decisions inside a continuously evolving system" — and the harness must decide, at every step, what to inject, what to compress, what to isolate (see Context Engineering). The same month, Cognition published Don't Build Multi-Agents, calling context engineering "the first priority of engineers building AI agents." The 2023 folder labeled "prompt tricks" was, at this point, formally handed over to the harness.

2025–2026: the harness gets its name, the frontier pivots ​

How "harness" won ​

In September 2025, Anthropic renamed the Claude Code SDK to the Claude Agent SDK, describing it in official copy as "the same agent harness that powers Claude Code." The term was quickly adopted by others including the OpenAI Codex team (etymology and definition in What Is an Agent Harness?).

Naming sounds like a small thing; it's actually a cognitive repositioning: an official admission that an agent's capability is the joint product of the model and that surrounding system. By late 2025, the Confucius Code Agent paper supplied quantified evidence via controlled experiment (same model, only the scaffold swapped, a 9-point spread on SWE-bench Pro), and in early 2026 "Stop Comparing LLM Agents Without Disclosing the Harness" put "scores without harness disclosure aren't comparable" right in the title. The once-mocked "wrapper" now has its own name, papers, and metrics.

Frontier 1: subagents and multi-agent, the second act ​

On June 13, 2025, Anthropic published an engineering retrospective of its multi-agent research system: an orchestrator-worker architecture where the lead agent decomposes the problem, spawns parallel subagents (each with its own context window), and synthesizes results — a 90.2% improvement over a single agent on internal research evals, at roughly 15× the token cost.

The key difference from 2023's multi-agent wave: this time the subagent isn't an executor of "grand plans" but a tool for context isolation — quarantining the token-hungry middle of the process (searching, reading, trial and error) inside sub-contexts and bringing only conclusions back to the main line (see Subagents). Cognition had published against multi-agent the day before (subagents make mutually conflicting implicit decisions); Anthropic answered with a production system. That 24-hour public collision of views still has no standard verdict — it's the most active frontier in harness design today.

Frontier 2: long-horizon tasks ​

In March 2025, METR published Measuring AI Ability to Complete Long Tasks: measuring how long a task an agent can reliably complete, using "how long it takes a human" as the yardstick, finding that the 50%-success time horizon of frontier models roughly doubles every 7 months (the strongest at the time, Claude 3.7 Sonnet, was on the order of 50 minutes).

Long-horizon workloads make the hardest demands on a harness: the context will overflow (compression and summarization needed), sessions will be interrupted (persistence and resume needed — see Memory Systems), goals will drift (externalized plan artifacts needed), and one context window won't be enough (subagent divide-and-conquer needed). Long tasks are the touchstone that stresses every harness component to its limit simultaneously — and that's why this site gives observability its own section: an agent that runs for hours without trajectory auditing is a black box.

History's boomerang

Note that 2023's ghost is still in the room: the every-7-months-doubling curve is re-igniting the "fully autonomous AI employee" narrative. Back then, what failed was model capability, and what was falsified was "no harness design needed." Neither conclusion has expired.

The throughline: what changes, what doesn't ​

Looking back over these five years, a clear two-layer motion:

Layer by layer, retired into the model (once the harness's job, later internalized by training): hand-written reasoning examples (→ reasoning models), tool-call format conventions (→ function calling), instruction-following prompt tricks (→ alignment training). Every model generation eats away part of the harness's thickness — an elaborate harness tuned for an old model often becomes dead weight on the new one.

Four question sets that never change, which every generation must answer anew:

QuestionThe 2023 answerThe 2025–26 answer
Context: what the model sees each stepStuff it all in + vector retrievalDynamic assembly, compression, isolation (context engineering)
Tools: what the model can doDozens of bespoke pluginsA few general primitives + MCP standard integration
Loop: how the task advancesSelf-prompting + replanning every stepSingle loop + externalized todo + event-driven revision
Control: when the human steps inFully autonomous, no humanTiered permissions + approvals at critical junctures

Models change, and will keep changing; but these four questions won't go away — because their root isn't model capability, it's the structural gap between the model and the real environment: finite context windows, changing environments, actions with consequences, and responsibility that someone must own. As long as that gap exists, a harness is needed to bridge it. That's why this site exists: to study agents is to study these invariant questions. Next, read Anatomy of the Harness to take today's answers apart layer by layer, or Model vs. Harness to slice the changing/invariant boundary finer.

Further reading ​