Appearance
Classic Papers, Annotated
Almost every component in today's mainstream coding agents — Claude Code, Cursor, OpenHands — has a prototype in these papers. Reading them isn't about knowing history; it's about seeing clearly which designs were validated by experiments, and which are just engineering inertia.
This page is a close reading of the 7 papers that influenced harness design the most. Each follows the same structure: the problem → the core mechanism → the key experimental results → the implications for harness design. The last part matters most — a paper's value isn't how many benchmark points it racked up, but how it changed the way everyone after built agents.
The Map: Papers and Their Harness Components
| Paper | Year | What it defined | Harness components |
|---|---|---|---|
| MRKL | 2022 | Modular architecture of router + expert tools | Tool System |
| ReAct | 2022 | The Thought–Action–Observation loop | Agent Loop |
| Toolformer | 2023 | A model teaching itself when to call tools | Tool System |
| Reflexion | 2023 | Verbal reflection on failure and episodic memory | Memory System |
| Generative Agents | 2023 | Memory stream + reflection + planning | Memory, Planning |
| Voyager | 2023 | Skill library and lifelong learning | Skills, Memory |
| SWE-agent | 2024 | Agent-Computer Interface | Tool System, Case Study: SWE-agent |
1. ReAct: Interleaving Reasoning and Action (2022)
Paper info
ReAct: Synergizing Reasoning and Acting in Language Models — Shunyu Yao et al. (Princeton University & Google Research Brain), released October 2022 (arXiv:2210.03629), accepted at ICLR 2023.
The Problem
In 2022, chain-of-thought (CoT) prompting had just been shown to boost reasoning dramatically, but CoT is a closed system: the model can only reason from the knowledge in its parameters, and once it misremembers a fact, the error cascades all the way down (hallucination cascades). Meanwhile, the action-taking agents of the day (imitation learning / RL methods on ALFWorld, say) did execute actions in the environment, but had no explicit reasoning traces — they couldn't plan, couldn't handle surprises, and couldn't explain what they were doing.
ReAct's question is refreshingly direct: can reasoning and action live in the same generation sequence, each feeding the other?
Core Mechanism
With a single few-shot prompt template, the model alternates between generating three kinds of text:
Thought 1: I need to find the developer of "Girls' Frontline", search first.
Action 1: search["Girls' Frontline developer"]
Observation 1: Girls' Frontline was developed by Sunborn Network (MICA Team)...
Thought 2: I need to confirm when Sunborn Network's founding company was registered, search again.
Action 2: lookup["Sunborn Network founding date"]
Observation 2: Sunborn Network was founded in 2015...
Thought 3: I have both facts now, I can answer.
Action 3: finish["2015"]There are only three key design choices, and each deserves a closer look:
- Thought is free text, not structured state. At any moment the model can maintain, in natural language, "what's my current plan, where am I stuck, what's next." It's a remarkably cheap implementation of "working memory" — zero extra infrastructure.
- The Action vocabulary is sparse and task-specific. HotpotQA has just three actions,
search / lookup / finish; ALFWorld uses the environment's action set. A small action space is what lets the model learn when to stop. - Observations are written back by the environment, and the model trusts its context unconditionally. That's both the source of the capability and the root of every prompt-injection problem that came after.
Key Findings
- Knowledge tasks (HotpotQA, FEVER, with PaLM-540B + a Wikipedia API): ReAct sharply reduced CoT's hallucination problem and matched or beat CoT; the ReAct + CoT-SC combination was best on FEVER.
- Decision-making tasks (ALFWorld, WebShop): ReAct pulled far ahead — a 34-point absolute gain in success rate on ALFWorld over the best imitation/RL baseline, and roughly 10 points on WebShop, at a time when other methods hadn't even reached half of human expert performance.
Implications for Harness Design
- The minimal skeleton of the agent loop was fixed here. Every tool-calling loop today (including native function calling) is an industrialized ReAct: Thought got internalized into model reasoning, Action got formalized into JSON-schema tool calls, and Observation gets truncated/compressed and fed back. For a thorough understanding of this loop, see Agent Loop.
- Writing reasoning traces into context is free context engineering. ReAct showed that letting the model "think, act, and take notes as it goes" beats both doing without thinking and thinking without doing — lesson one of Context Engineering: context isn't just reference material for the model to read; it's the model's own scratchpad.
- The action space should be small and specialized. ReAct completed its tasks with three to five actions, echoing what SWE-agent would later formalize as the ACI (see paper 7): the more the interface feels designed for the model, the higher the success rate — rather than throwing the whole human shell at it.
2. MRKL: A Modular Neuro-Symbolic Architecture (2022)
Paper info
MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning — Ehud Karpas et al. (AI21 Labs), released May 2022 (arXiv:2205.00445).
The Problem
MRKL answers a structural question: LLMs know a little about everything, but they fumble arithmetic, their facts go stale, and their reasoning drifts. Instead of hoping a bigger model will learn everything, why not treat the model as a dispatcher and hand deterministic work to deterministic systems?
Core Mechanism
MRKL (Modular Reasoning, Knowledge and Language — pronounced "miracle") is a three-stage architecture:
User question
│
┌──────▼──────┐
│ Router │ ← An LLM plays the router,
│ (Jurassic-1)│ deciding which expert to call
└──────┬──────┘
┌───────────┼────────────┐
▼ ▼ ▼
┌──────────┐ ┌──────────┐ ┌──────────┐
│Calculator│ │ Knowledge│ │ Generic │ ← Expert modules
│(discrete │ │ base/API │ │ LLM │
│ symbolic)│ │ expert │ │ fallback │
└──────────┘ └──────────┘ └──────────┘- Router: an LLM that routes each input to the right expert. The paper sketches a way to train the router, but also notes that prompt-based routing already works.
- Expert modules: a calculator, weather/exchange-rate APIs, database retrieval, plus a "generic LLM fallback expert." Each expert handles only the deterministic tasks it's good at.
- A neuro-symbolic pairing: the neural side understands intent, the symbolic side executes deterministic computation — each does what it does best.
Key Findings
The paper reports that an MRKL system built from a 6B-parameter Jurassic-1 plus a calculator expert beat 175B GPT-3 on arithmetic word problems — a model 30x smaller overtaking the giant by bolting on a deterministic tool. In 2022 this was counterintuitive: the prevailing narrative was still "scale is everything."
Implications for Harness Design
- MRKL was the first complete definition of a "tool system": who decides to call (the router), what to call (the experts), and how results come back (handed to the LLM for integration). Today's tool-calling protocols (function calling, MCP) are just this architecture, standardized. See Tool System for details.
- Routing is the most fragile link in this architecture. The paper admits it itself: add an expert and you have to retrain or retune the router. That pain persists today — once tools multiply, a model picking the wrong tool or passing wrong arguments is a routing problem at heart. See Common Pitfalls.
- The "generic LLM fallback expert" is engineering honesty: not every input can be routed to a deterministic expert, and the fallback guarantees the system stays complete. Designing a harness today means answering the same question: what does the system do when no tool can handle the input?
3. Toolformer: A Model Teaching Itself to Use Tools (2023)
Paper info
Toolformer: Language Models Can Teach Themselves to Use Tools — Timo Schick et al. (Meta AI), released February 2023 (arXiv:2302.04761), accepted at NeurIPS 2023.
The Problem
MRKL's router needed hand-specified rules or labeled data; ReAct relied on carefully hand-written few-shot examples. Could the model teach itself where to insert tool calls, which tool to call, and how to weave results back in? Toolformer turned tool use from a "prompt engineering problem" into a "training data generation problem."
Core Mechanism
Starting from GPT-J (6.7B), in four steps:
- Sample candidate calls: with a few-shot prompt, the model casually drops API calls into ordinary text (e.g.,
The population of Pittsburgh is QA("What is the population of Pittsburgh?") → 302,971). - Execute the calls: actually run those 5 tools — a QA system, a calculator, Wikipedia search, machine translation, and a calendar.
- Filter by usefulness: the cleverest step in the paper. For each candidate call, compare the continuation with the tool result against the continuation without it, and see which gives lower loss on the text that follows. Only calls that genuinely reduce the prediction loss are kept.
- Fine-tune on the filtered data: the model thereby learns "when it's worth interrupting itself to go check."
Note that the criterion in step 3 is purely self-supervised: no human labeling, and "useful" is defined simply as "helped the model better predict the text that came next."
Key Findings
- The 6.7B Toolformer outperformed 175B GPT-3 zero-shot across several downstream tasks, especially ones demanding precise facts or arithmetic.
- The model's tool-calling behavior showed an emergent self-awareness: it tends to insert calls at numbers, dates, and factual noun phrases — exactly where it's most likely to go wrong.
Implications for Harness Design
- When to call a tool can be learned; it doesn't all have to be hard-coded into prompts. Toolformer is the common ancestor of every later "tool-use data synthesis" effort — model vendors training function calling today still follow the same core pipeline: synthesize candidates → execute → filter → fine-tune.
- "Does it reduce loss?" is a transferable engineering instinct for a filter. When you build memory writes or skill consolidation into a harness, you likewise need an objective criterion for "is this lesson worth keeping?" — instead of accepting everything that comes along. The idea resurfaces in Voyager's skill verification (paper 6).
- A model will expose its own uncertainty. The distribution of call locations Toolformer learned is effectively a map of "where the model is unreliable." When designing observability, the density and placement of tool calls are signals worth recording — see Observability.
4. Reflexion: Reinforcement Learning with Language (2023)
Paper info
Reflexion: Language Agents with Verbal Reinforcement Learning — Noah Shinn et al. (Northeastern University & MIT), released March 2023 (arXiv:2303.11366), accepted at NeurIPS 2023.
The Problem
ReAct made agents that can do things, but that don't get better after failing: same trap next time, same fall. Traditional reinforcement learning solves this, but updating weights is expensive, slow, and off the table for closed models. Reflexion's question: can an agent learn from failure without touching weights, using nothing but natural language?
Core Mechanism
Reflexion turns the whole RL triad into words:
┌────────────┐
│ Actor │ ← Generates a trajectory (e.g., a ReAct loop)
└─────┬──────┘
▼ The environment returns a success/failure signal (scalar or text)
┌────────────┐
│ Evaluator │ ← Judges whether this attempt succeeded
└─────┬──────┘
▼ On failure
┌────────────┐
│Self-Reflect│ ← An LLM writes a natural-language reflection:
│ │ "Last time I failed because I kept going
│ │ without checking the return value..."
└─────┬──────┘
▼
┌────────────┐
│ Episodic │ ← The reflection goes into a buffer and is injected
│ memory │ into the prompt on the next attempt, steering
└────────────┘ the new trajectoryAll three components are played by the same LLM with different prompts. The "gradient" becomes text: not a parameter update, but a lesson written for your next self, stored in episodic memory.
Key Findings
- HumanEval: Reflexion hit 91% pass@1, beating bare GPT-4's 80% — a "reflection loop" lifted the same model's score by 11 points.
- ALFWorld: ReAct + Reflexion completed 130 of 134 tasks (about 97%), well ahead of plain ReAct.
- HotpotQA: reflection clearly beat a naive retry baseline, showing the gains come from "what was learned," not "how many tries."
Implications for Harness Design
- Failure is a more valuable signal than success — but only if you write it down. The most practical lesson for Memory System design: episodic memory shouldn't just store "what happened"; it should store "why the last attempt failed and what not to do this time."
- The evaluator must be reliable, or reflection just adds to the hallucination. Reflexion's gains rest on "the environment can produce a trustworthy success/failure signal" (unit tests, task-completion checks). On tasks without a reliable verifier, self-reflection easily degenerates into self-soothing — a point Common Pitfalls hammers repeatedly.
- The loop needs a limit. Reflexion caps retries by default; an unbounded "reflect-retry" loop burns tokens and never converges. In agent loop design today, max steps, max retries, and budget caps are required equipment — see Agent Loop.
5. Generative Agents: Memory Stream, Reflection, and Planning (2023)
Paper info
Generative Agents: Interactive Simulacra of Human Behavior — Joon Sung Park et al. (Stanford University & Google Research), released April 2023 (arXiv:2304.03442), Best Paper at UIST 2023.
The Problem
Let 25 agents "live" for two in-game days in Smallville, a sandbox town: wake up, make breakfast, go to work, chat, spread the news, even spontaneously organize a Valentine's Day party. For behavior to be believable, an agent has to remember what happened yesterday, form opinions of the others, and plan for tomorrow — which happens to be the extreme version of the "memory + planning" twin problems of harness design.
Core Mechanism
The architecture is a trio, all built around one core data structure — the memory stream, a long list that logs every one of the agent's experiences in natural language:
- Retrieval: before each decision, the most relevant memories are pulled from the stream by a weighted blend of three scores — recency (exponential decay), relevance (embedding cosine similarity), and importance (an LLM-assigned 1–10 rating).
- Reflection: when accumulated importance crosses a threshold, the agent stops to "think about its life": recent memories go to the LLM, which synthesizes higher-level insights (e.g., "Klaus has been working on the library project; he seems passionate about it"), and the insights are written back into the stream.
- Planning: top-down recursive decomposition — first a rough plan for the day, then broken into hour-level blocks, then into 5–15 minute actions; unexpected events trigger local replanning.
Key Findings
- Within two days, the 25 agents produced social behavior nobody programmed: an invitation to the Valentine's party spread from 1 agent to 8, and 5 agents showed up at the right time and place; word about who was running for town mayor also circulated through the town.
- The ablations matter: remove any one of reflection, planning, or retrieval, and human raters' scores on "behavioral believability" drop significantly — all three are load-bearing architecture, not decoration.
Implications for Harness Design
- Memory isn't a log; the retrieval strategy is the real object. "Store everything + retrieve weighted by recency/relevance/importance + periodically compress into higher-level abstractions" — this paradigm directly defined how agent long-term memory is built today. See Memory System for details.
- Reflection is a compression algorithm for memory. Raw experiences grow without bound; reflection distills many observations into reusable judgments — complementary to Reflexion's lesson memory: one distills a worldview, the other distills mistakes.
- Recursive decomposition planning (day → hour → minute) is hierarchical planning in its minimal form, and the shape most later Planning systems adopted. Mind the cost: every step burns LLM calls, and running Smallville for two days cost thousands of dollars in API fees — an elegant architecture is not the same as a controllable one.
6. Voyager: A Lifelong Learning Agent That Accumulates Skills (2023)
Paper info
Voyager: An Open-Ended Embodied Agent with Large Language Models — Guanzhi Wang et al. (NVIDIA, Caltech, UT Austin, and others), released May 2023 (arXiv:2305.16291).
The Problem
Earlier Minecraft agents (mostly RL methods) could only handle narrow, predefined tasks; they couldn't explore openly and couldn't accumulate reusable capability. Voyager's question: can an agent set its own problems, write its own code to solve them, and store the solved skills so it keeps compounding?
Core Mechanism
Three components, all driven through GPT-4 as a black box, with no weight updates at all:
- Automatic curriculum: based on the current state (position, inventory, completed tasks), GPT-4 proposes "the next appropriately difficult task" — chop trees first, then craft a wooden pickaxe, then mine stone.
- Iterative prompting: write code → run it → get environment feedback and errors → self-verify whether the task is done → on failure, rewrite with the critique attached, until success or giving up.
- Skill library: verified code skills go into a vector database, keyed by the embedding of the skill's description, with the code itself as the value. On a related task later, it retrieves the top-5 skills as context or calls them directly.
Voyager's action space is JavaScript code (via the Mineflayer API), not discrete button presses — an early statement of "code as action."
Key Findings
- Against the SOTA of the day (ReAct, Reflexion, AutoGPT-style baselines), Voyager: 3.3x more unique items discovered, 2.3x the travel distance, and up to 15.3x faster at unlocking key tech-tree milestones.
- Ablations show the skill library is where the compounding comes from: Voyager with its library keeps getting faster; empty the library and the agent falls back to "writing everything from scratch every time."
- Zero-shot generalization: dropped into a new world, Voyager with its old skill library works out of the box and is markedly stronger than starting from scratch.
Implications for Harness Design
- The skill library = executable memory. This is Voyager's most influential idea: memory doesn't have to be text — it can be code or procedures verified to actually run. Claude Code's Skills and the "tool consolidation" features in various agents share the same lineage. See Skills for details.
- Verify before storing. Voyager writes a skill into the library only after it "passes self-verification" — the embodied version of Toolformer's "keep it only if it helps" filter. An unverified skill library is a junkyard; the more you retrieve from it, the more it misleads.
- Code, not an API list, as the action space buys composability: code can loop, branch, and reuse old skills. CodeAct (2024) later validated this judgment systematically — see Further Reading at the end.
- The automatic curriculum is a practical answer to "task planning": instead of one grand, unreachable goal, the agent keeps proposing sub-goals just within reach. Directly useful for designing planners for long-horizon tasks — see Planning.
7. SWE-agent: The Interface Is the Capability (2024)
Paper info
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering — John Yang et al. (Princeton University), released May 2024 (arXiv:2405.15793), accepted at NeurIPS 2024.
The Problem
By 2024, GPT-4-class models were already strong at writing code, yet on the real-world software engineering benchmark SWE-bench (fixing real GitHub issues), the best non-interactive retrieval systems solved only 3.8%. SWE-agent's bet: the bottleneck isn't the model, it's the interface — the toolset, feedback formats, and interaction flow designed for the model (which the authors call the Agent-Computer Interface, ACI) determine results more than the model itself does.
Core Mechanism
SWE-agent isn't "let the model loose on bash"; it builds a purpose-designed interface around the question "what does an LM get wrong?":
- File viewer: after
open, only a roughly 100-line window is shown, withscroll_up / scroll_down / gotoalongside — preventing the attention drift that comes from stuffing a whole file into the model at once. - Constrained editor: the
editcommand forces an exact line range to replace, avoiding the errors of big free-form diffs; edits are lint-checked automatically, and syntax errors are fed back immediately. - Search commands:
search_dir / search_filereturn trimmed results, not raw grep output. - Feedback formatting: environment output is pruned into a form the model can easily read; when the model emits malformed output, the error message it gets back is itself carefully designed to teach self-correction.
In one line: they designed an IDE for an LLM the way you'd design one for a human engineer.
Key Findings
- GPT-4 Turbo + SWE-agent solved 12.47% (286/2294) of the full SWE-bench test set and 18.00% (54/300) of the Lite subset — the interactive-system SOTA of its day, far beyond the 3.8% of the earlier non-interactive systems.
- Ablation: swap the carefully designed ACI for a "bare shell" baseline and the success rate drops 10.7 points — interface design contributed most of the improvement.
- Portability: the ACI designed for GPT-4 Turbo still solved 10.5% when moved to Claude 3 Opus — good interface design generalizes across models.
Implications for Harness Design
- The ACI is the purest expression of the harness idea: same model, different interface, a 3x-plus difference in success rate. The hardest experimental evidence for "model capability = model × harness" comes from this paper. Full analysis in Case Study: SWE-agent.
- An interface designed for humans is not an interface designed for models. grep's full output, bash's error formats, an editor's cursor interactions — all built around human cognition. The model-facing versions should: cap information density, structure the feedback, and make every error message a tutorial. The principle applies to all Tool System design.
- Constraints are features, not limitations. Limiting the view to 100 lines at a time and requiring a line range for edits look like they take freedom away from the model, but they really outsource attention management to the harness. Same core claim as Context Engineering: less noise in the context beats more cleverness in the model.
A Timeline View
2022.05 MRKL ──────────── Router + experts: the tool system takes shape
2022.10 ReAct ─────────── Reasoning interwoven with action: the agent loop is set
2023.02 Toolformer ────── Tool calling goes from prompt trick to trainable skill
2023.03 Reflexion ─────── Failure verbalized as memory: learning from mistakes
2023.04 Generative Agents Memory stream + reflection + planning
2023.05 Voyager ───────── The skill library: executable lifelong memory
2024.05 SWE-agent ─────── ACI: interface design as system capabilityRead the 7 together and a clear throughline emerges: the research focus keeps shifting from "what's inside the model" to "what's wrapped around it". Before ReAct, people assumed capability lived in the parameters; after SWE-agent, it's hard for any researcher to deny — the harness itself is a primary variable in system capability.
Want to go deeper? These are worth reading too
- CodeAct (arXiv:2402.01030, ICML 2024): makes a systematic case that "executable Python code as a unified action space" beats JSON/text actions, with higher success rates in fewer steps — the generalization of Voyager's "code as action" idea.
- WebGPT (arXiv:2112.09332, 2021): a QA agent trained with human feedback to interact with a browser — the forerunner of browsing agents.
- Tree of Thoughts (arXiv:2305.10601, NeurIPS 2023): extends single-chain reasoning into a searchable tree of thoughts, shaping a generation of planner designs.
- HuggingGPT (arXiv:2303.17580, NeurIPS 2023): an LLM as controller orchestrating specialist models — MRKL's idea scaled up to the model ecosystem.
For the full literature map, see Frontier Papers.
Further Reading
- Agent Loop — the engineering realization of the ReAct loop
- Tool System — tool architecture from MRKL to MCP
- Memory System — memory streams, reflection, and skill libraries in one view
- Planning — recursive decomposition and automatic curricula today
- Skills — the skill library idea turned into product
- Case Study: SWE-agent — a complete case analysis of ACI design
- Frontier Papers — what came after these classics
References
- ReAct: Synergizing Reasoning and Acting in Language Models (arXiv:2210.03629)
- MRKL Systems (arXiv:2205.00445)
- Toolformer: Language Models Can Teach Themselves to Use Tools (arXiv:2302.04761)
- Reflexion: Language Agents with Verbal Reinforcement Learning (arXiv:2303.11366)
- Generative Agents: Interactive Simulacra of Human Behavior (arXiv:2304.03442)
- Voyager: An Open-Ended Embodied Agent with Large Language Models (arXiv:2305.16291)
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering (arXiv:2405.15793)
- CodeAct (arXiv:2402.01030)