Appearance
The Paper Map
Agent Harness — an agent's skeletal frame — wasn't the invention of any single paper. It is a systemic picture that dozens of papers assembled piece by piece over three years: ReAct supplied the interleaved reasoning-and-acting main loop, Toolformer proved tool use could be trained into the model, SWE-bench dragged evaluation into real software engineering, and SWE-agent put interface design in a paper title for the first time. This page organizes that work into a map by topic, so you can figure out where a new paper lands on the grid and what to read first.
How to use this map
If you only have two hours, go straight to Reading Paths and pick one. Every paper is tagged with its arXiv ID, and the publication years and author institutions have been verified against the original publication pages. Within each topic, papers are ordered chronologically so you can watch the ideas evolve.
Paradigm Shift Timeline
text
2021 2022 2023 2024
│ │ │ │
HumanEval CoT ─────────────► ReAct sets the Agent Loop paradigm │
(Codex) SayCan Toolformer / Gorilla / ToolLLM │
Inner Monologue Generative Agents / MemGPT │
STaR Voyager / Reflexion / Self-Refine │
WebArena / AgentBench / GAIA ──────►│
SWE-bench ─────────────────────────►│
OSWorld
SWE-agent (ACI)
CodeAct
AutoCodeRover / Agentless
OpenHandsThree clear phase shifts:
- 2022: from generating answers to interleaved decisions. CoT got models to reason before answering; ReAct wove reasoning and external actions together, and the Agent Loop took its first recognizable shape.
- 2023: the harness components get invented one by one. Tools (Toolformer, Gorilla), memory (Generative Agents, MemGPT), skill libraries and self-reflection (Voyager, Reflexion), and real-environment evaluation (WebArena, SWE-bench) all surfaced within roughly the same year — by its end, the anatomy of the harness was essentially drawn.
- 2024: system design and evaluation discipline take center stage. SWE-agent introduced the Agent-Computer Interface (ACI) concept, Agentless cross-examined the need for agents from the outside with an agent-free pipeline, and OSWorld pushed evaluation into real operating systems. The research focus shifted from making models think better to fencing them in better — which is exactly the heart of harness research.
Reasoning and Acting Paradigms
This group of papers defines the agent's heartbeat: what the model thinks at each step, what it does, and how the two alternate.
| Paper | Institution | Year | Contribution in one sentence |
|---|---|---|---|
| Chain-of-Thought Prompting Elicits Reasoning in Large Language Models | Google Research | 2022 | Uses few-shot exemplars to elicit intermediate reasoning steps, establishing the think-before-you-answer baseline capability |
| ReAct: Synergizing Reasoning and Acting in Language Models | Princeton / Google Brain | 2022 | Interleaves reasoning traces with external actions — the template for nearly every modern agent loop |
| Tree of Thoughts: Deliberate Problem Solving with Large Language Models | Princeton / Google DeepMind | 2023 | Expands single-chain reasoning into a searchable, backtrackable tree of thoughts, lifting the Game of 24 success rate from 4% to 74% |
| Executable Code Actions Elicit Better LLM Agents (CodeAct) | UIUC | 2024 | Replaces JSON tool calls with executable Python code as a unified action space, improving success rates by up to 20% |
Independent judgment
ReAct's real legacy isn't the prompt template — it's that the observe-think-act turn structure was cemented as the harness's main-loop protocol. CodeAct reminds us that the representation of the action space (JSON versus code) is itself a harness design decision, and a consequential one. Both papers deserve a word-by-word read.
Tool Learning
How does a model learn to call tools — through prompt orchestration, or by internalizing them through training?
| Paper | Institution | Year | Contribution in one sentence |
|---|---|---|---|
| Toolformer: Language Models Can Teach Themselves to Use Tools | Meta AI | 2023 | Self-supervised: the model learns when to call an API and what arguments to pass, with only a handful of examples per API |
| HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face | Zhejiang University / Microsoft Research Asia | 2023 | An LLM acts as the controller dispatching expert models on Hugging Face, with language as the universal interface |
| Gorilla: Large Language Model Connected with Massive APIs | UC Berkeley | 2023 | Retrieval-augmented fine-tuning beats GPT-4 on API call generation while easing API hallucination |
| ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs | Tsinghua University / ModelBest | 2023 | The ToolBench dataset plus a depth-first-search decision tree give open-source models command of over ten thousand real APIs |
Tool learning cuts both ways for the harness: the better the harness writes its tool descriptions, the more accurately the model calls them (see Tools & MCP); and the better the model gets trained at calling tools, the less orchestration the harness has to do.
Self-Improvement
How does an agent get better at tasks without gradient updates?
| Paper | Institution | Year | Contribution in one sentence |
|---|---|---|---|
| STaR: Bootstrapping Reasoning With Reasoning | Stanford | 2022 | Fine-tunes the model on its own rationales that lead to correct answers, bootstrapping reasoning ability |
| Reflexion: Language Agents with Verbal Reinforcement Learning | Northeastern / MIT / Princeton | 2023 | Replaces weight updates with verbal reflection, reaching 91% HumanEval pass@1 (as reported in the paper) |
| Self-Refine: Iterative Refinement with Self-Feedback | CMU / Allen Institute for AI | 2023 | A single LLM doubles as generator, feedback provider, and refiner, iterating on outputs without any training |
| Voyager: An Open-Ended Embodied Agent with Large Language Models | NVIDIA / Caltech / UT Austin | 2023 | Automatic curriculum plus an executable code skill library plus self-verification in Minecraft, delivering lifelong learning |
A note on reading Reflexion
The 91% HumanEval figure Reflexion reports was obtained under conditions that allow iterative attempts and test feedback — it is not on the same track as the 80% single-shot baseline. That gap is precisely the harness's value, but cite the numbers with their conditions spelled out.
Memory
What an agent remembers, how it retrieves it, and how it compresses it determine how long a task it can run.
| Paper | Institution | Year | Contribution in one sentence |
|---|---|---|---|
| Generative Agents: Interactive Simulacra of Human Behavior | Stanford / Google Research | 2023 | An observe-reflect-plan memory stream architecture from which 25 simulated townsfolk emerged with believable social behavior |
| MemGPT: Towards LLMs as Operating Systems | UC Berkeley | 2023 | Borrows tiered storage and "virtual context management" from operating systems, using interrupts to page memory in and out on its own |
Voyager's skill library can also be read as a kind of memory — except that what it stores isn't facts but things the agent knows how to do. All three routes (retrieval-based memory streams, OS-style paging, programmatic skill libraries) have counterparts in today's production harnesses; see Memory Systems.
Planning
Separating "what to do next" from the model itself is the most persistent debate in agent research.
| Paper | Institution | Year | Contribution in one sentence |
|---|---|---|---|
| Do As I Can, Not As I Say: Grounding Language in Robotic Affordances (SayCan) | Google Research | 2022 | Constrains LLM planning with affordance scores over skills, so plans land on things the robot can actually do |
| Inner Monologue: Embodied Reasoning through Planning with Language Models | Google Research | 2022 | Closed-loop language feedback (success detection, scene description, human input) acts as an inner monologue to improve embodied planning |
| On the Planning Abilities of Large Language Models | Arizona State University | 2023 | A critical evaluation built on International Planning Competition domains: LLMs autonomously produce executable plans only about 3% of the time |
| LLM+P: Empowering Large Language Models with Optimal Planning Proficiency | UT Austin | 2023 | The LLM translates the natural-language problem into PDDL, hands it to a classical planner for the optimal solution, and translates the result back |
Taken together, these papers tell one complete story: LLMs plan poorly on their own (Valmeekam et al.), but the picture improves dramatically once you outsource planning to a symbolic system (LLM+P) or constrain it in a closed loop with environment feedback (SayCan, Inner Monologue). This is the academic source of the harness's division-of-labor idea, mirrored in Planning & Task Decomposition.
Evaluation Benchmarks
Without benchmarks there is no harness engineering. The benchmarks below define the yardstick for whether an agent works.
| Benchmark | Institution | Year | Contribution in one sentence |
|---|---|---|---|
| HumanEval (Evaluating Large Language Models Trained on Code) | OpenAI | 2021 | 164 hand-written programming problems plus execution-based judging — the de facto starting point of code-generation evaluation |
| WebArena: A Realistic Web Environment for Building Autonomous Agents | CMU | 2023 | A self-hosted environment of fully functional real websites; the strongest GPT-4 agent of its day managed only 14.41% (humans: 78.24%) |
| AgentBench: Evaluating LLMs as Agents | Tsinghua University | 2023 | Multi-dimensional evaluation across 8 interactive environments, systematically charting the agent-capability gap between commercial and open-source models |
| SWE-bench: Can Language Models Resolve Real-World GitHub Issues? | Princeton | 2023 | 2,294 real GitHub issues, pushing evaluation from writing functions to fixing real repositories — the industry's gold standard |
| GAIA: a benchmark for General AI Assistants | Meta AI / Hugging Face et al. | 2023 | 466 real-world questions that are easy for humans and hard for AI, calling for browsing, reasoning, and tool combinations |
| OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments | University of Hong Kong / CMU / Salesforce et al. | 2024 | 369 open-ended tasks in real operating systems with execution-based evaluation; the best model scores only 12.24% |
How to read a benchmark paper
Don't stop at the leaderboard numbers. Look at three things: where the tasks come from (realism), how success is judged (execution-based or match-based), and what the failure cases look like (which decides what the harness should supply). The "Claude 2 solved only 1.96%" figure in the SWE-bench paper matters because it proved the problem back then wasn't model intelligence — it was the mode of interaction.
Software Engineering Agents
Software engineering is the field where harness thinking has landed most completely, and it is the center of gravity of this site.
| Paper | Institution | Year | Contribution in one sentence |
|---|---|---|---|
| SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering | Princeton | 2024 | Introduces the Agent-Computer Interface (ACI) concept, showing that interface design alone significantly moves agent performance |
| AutoCodeRover: Autonomous Program Improvement | National University of Singapore | 2024 | Organizes code search around program structure (AST/classes/methods) rather than a flat file view, at a markedly lower cost than its contemporaries |
| Agentless: Demystifying LLM-based Software Engineering Agents | UIUC | 2024 | A three-stage localize-repair-verify agent-free pipeline that reached the top ranks of SWE-bench Lite at a fraction of the cost |
| OpenHands: An Open Platform for AI Software Developers as Generalist Agents | UIUC / CMU et al. | 2024 | A community-driven open platform: code execution sandboxes, multi-agent coordination, and integrated evaluation across 15 benchmarks |
Why Agentless is this site's most important cautionary tale
Agentless deliberately forgoes the agent loop — the model never freely decides the next step, a fixed pipeline does all the work — and still beat every open-source software agent of its time. The lesson is not that "agents are useless"; it is that in harness design, every degree of freedom must have a reason. Autonomy piled on blindly gets paid for in reliability, cost, and debuggability. This dovetails with Design Principles and Common Pitfalls on this site.
Surveys
To build the full picture quickly, these two surveys complement each other.
| Paper | Institution | Year | Contribution in one sentence |
|---|---|---|---|
| A Survey on Large Language Model based Autonomous Agents | Renmin University of China | 2023 | Proposes a unified construction-application-evaluation framework for organizing agent research, with a continuously maintained paper repository |
| The Rise and Potential of Large Language Model Based Agents: A Survey | Fudan University | 2023 | A brain-perception-action three-part framework, plus agent societies and their philosophical roots; a long read |
Further Reading
- What Is an Agent Harness? — the components from the papers reorganized under the harness concept
- Anatomy of the Harness — mapping this paper map onto a system architecture diagram
- The Agent Loop — the ReAct paradigm worked out in engineering terms
- Context Engineering — a close reading of each paper's prompt and context strategy
- Planning & Task Decomposition — the contemporary echoes of the planning papers
- Memory Systems — the contemporary echoes of the memory papers
- Core Papers: Close Readings — paper-by-paper close readings of a select few
- Frontier Work — new work after 2024
- SWE-agent Case Study — the complete case for the ACI idea
- OpenHands Case Study — the open-platform case
References
- ReAct (arXiv:2210.03629)
- Chain-of-Thought (arXiv:2201.11903)
- Tree of Thoughts (arXiv:2305.10601)
- CodeAct (arXiv:2402.01030)
- Toolformer (arXiv:2302.04761)
- HuggingGPT (arXiv:2303.17580)
- Gorilla (arXiv:2305.15334)
- ToolLLM (arXiv:2307.16789)
- STaR (arXiv:2203.14465)
- Reflexion (arXiv:2303.11366)
- Self-Refine (arXiv:2303.17651)
- Voyager (arXiv:2305.16291)
- Generative Agents (arXiv:2304.03442)
- MemGPT (arXiv:2310.08560)
- SayCan (arXiv:2204.01691)
- Inner Monologue (arXiv:2207.05608)
- On the Planning Abilities of LLMs (arXiv:2302.06706)
- LLM+P (arXiv:2304.11477)
- HumanEval / Codex (arXiv:2107.03374)
- WebArena (arXiv:2307.13854)
- AgentBench (arXiv:2308.03688)
- SWE-bench (arXiv:2310.06770)
- GAIA (arXiv:2311.12983)
- OSWorld (arXiv:2404.07972)
- SWE-agent (arXiv:2405.15793)
- AutoCodeRover (arXiv:2404.05427)
- Agentless (arXiv:2407.01489)
- OpenHands (arXiv:2407.16741)
- A Survey on LLM based Autonomous Agents (arXiv:2308.11432)
- The Rise and Potential of LLM Based Agents: A Survey (arXiv:2309.07864)