A Brief History
From ReAct and Toolformer in 2022 to the harness discourse of 2026 — every turning point in the five-year evolution of the agent harness: what people thought the solution was at the time, and what was later falsified or confirmed.
AGENT HARNESS HANDBOOK
Beyond the model: the system that makes agents actually get work done — agent loop · context engineering · tools · memory · subagents · safety boundaries
40+ articles aren't a library you must read cover to cover — they're routes you can combine on demand
Interviewing within 1–3 months: go straight to JD benchmarking, resume rewrites, and the interview question bank — every page maps to what big tech actually tests.
View path →Have time to build a solid foundation: move week by week from the guide through hands-on practice, with a clear checkpoint at the end of every week.
View path →Working builders don't read it cover to cover: a problem → page index for looking things up on the spot — treat it as a reference book.
View path →If you only read ten, make it these
One diagram to understand the overall architecture of an Agent Harness: how the user, the environment, the agent loop, the eight core components, and the LLM fit together — with a full data-flow walkthrough and role-based reading routes.
GuideThis site's 40-odd pages aren't a library to read cover to cover — they're a route map you can combine on demand: a two-week Job-Hunt Sprint for anyone interviewing within three months, a six-week Career Switch for people who want their foundations done right, and a problem → page Desk Reference index for builders already on the job.
GuideAn agent harness is the complete system wrapped around an LLM: context, tools, the loop, memory, and guardrails. This article gives a precise definition, traces the word's etymology, untangles harness from model, agent, framework, scaffolding, and wrapper, and presents the evidence that the harness sets the ceiling on what an agent can do.
GuideContext is an agent's scarcest resource, and the harness's core job is deciding what the model sees at every step. Anatomy of a real context window, the four failure modes (poisoning, distraction, confusion, clash), compression and truncation strategies, KV cache hit-rate economics, and the trade-off between RAG and just-in-time retrieval.
Core ComponentsThe agent loop is the heart of Harness—how a single while loop assembles context, calls the model, executes tools, handles errors, and decides when to stop. A complete teardown, from the ReAct paper to a production-grade implementation.
Core ComponentsA deep dive into Claude Code's harness design — the minimal Unix-style toolset, the planning mechanisms behind TodoWrite and plan mode, CLAUDE.md memory files, subagents, hooks, MCP, and permission modes, plus why it chose the terminal over the IDE. The site's component framework, assembled in full inside a real product.
Case StudiesA close reading of 7 classic papers that shaped how agent harnesses are designed — ReAct, MRKL, Toolformer, Reflexion, Generative Agents, Voyager, and SWE-agent — covering their mechanisms, key results, and what they mean for system design.
PapersImplement a minimal agent harness from zero in one ~130-line, framework-free Python file — dissecting the five stages of context assembly, model calls, tool parsing and execution, result write-back, and human approval — then mounting todo planning, memory compression, subagents, and observability one at a time along a "symptom → component" path.
Practice2026 年 8 月国内外 22 家公司 44 个 Agent/LLM 工程在招岗位完整清单:官方全文、官方标题级、聚合站快照三级来源标注,附总览速查表与失效免责说明。
Careers一份策展式而非罗列式的 Agent Harness 阅读地图:只收录一手、有立场的工程博客、开源实现、协议规范与论文入口,每条都注明为什么值得读、什么时候读。
ResourcesBrowse by section, or search directly
From ReAct and Toolformer in 2022 to the harness discourse of 2026 — every turning point in the five-year evolution of the agent harness: what people thought the solution was at the time, and what was later falsified or confirmed.
One diagram to understand the overall architecture of an Agent Harness: how the user, the environment, the agent loop, the eight core components, and the LLM fit together — with a full data-flow walkthrough and role-based reading routes.
This site's 40-odd pages aren't a library to read cover to cover — they're a route map you can combine on demand: a two-week Job-Hunt Sprint for anyone interviewing within three months, a six-week Career Switch for people who want their foundations done right, and a problem → page Desk Reference index for builders already on the job.
The same model inside different agent harnesses (scaffolds) can perform several times apart on real tasks. Hard numbers from SWE-bench to argue that, at equal model capability, the harness is the dominant variable in agent performance — plus a look at how model progress rewrites harness design in return.
An agent harness is the complete system wrapped around an LLM: context, tools, the loop, memory, and guardrails. This article gives a precise definition, traces the word's etymology, untangles harness from model, agent, framework, scaffolding, and wrapper, and presents the evidence that the harness sets the ceiling on what an agent can do.
Context is an agent's scarcest resource, and the harness's core job is deciding what the model sees at every step. Anatomy of a real context window, the four failure modes (poisoning, distraction, confusion, clash), compression and truncation strategies, KV cache hit-rate economics, and the trade-off between RAG and just-in-time retrieval.
Why agent evals are hard, trace recording and replay, the major benchmarks (SWE-bench, Terminal-Bench, τ-bench, WebArena/WorkArena), how to use LLM-as-judge and where its biases lie, and how to build an evals-driven harness iteration workflow.
A teardown of the Agent Harness memory system: the three-layer memory model, the rise of the memory-file pattern, when to write and how to forget, the division of labor between memory and RAG, and three representative implementations—Claude Code, ChatGPT, and Letta.
Which model to assign to which step is a design decision that directly shapes both capability and the bill — the capability/price/latency three-way trade-off across flagship and small models, three routing modes (rule-based, cascade, and learned routers), fallback and retry discipline, cost knobs such as prompt caching and batch processing, and the eval discipline that every model swap must run regressions.
A deep dive into the trust mechanics of an agent harness: the spectrum of permission modes from full manual approval to yolo, Claude Code's four permission modes and hook guardrails, the paired filesystem-and-network sandbox, the prompt injection and lethal trifecta threat model, and the design principles behind human-in-the-loop gates and read/write tool separation.
A deep dive into how AI agents plan — a comparison of three modes (no planning, plan-and-execute, interleaved planning), how the todo list works as explicit working memory, and whether explicit planning still matters in the era of reasoning models.
The three ways agents acquire domain knowledge — static injection (system prompts and rules files), dynamic retrieval (RAG), and capability packages (Agent Skills) — with a close look at Anthropic's progressive disclosure design, rule precedence and conflict resolution, and how to pick the right approach for the job.
What subagents really are — context isolation and parallelization, not "multiple AIs holding a meeting." Claude Code's delegate-and-summarize mechanism, the orchestrator-worker architecture and hard numbers behind Anthropic's multi-agent research system, Cognition's counter-argument in "Don't Build Multi-Agents," and a decision framework for choosing between a strong single-threaded agent and a multi-agent setup.
The agent loop is the heart of Harness—how a single while loop assembles context, calls the model, executes tools, handles errors, and decides when to stop. A complete teardown, from the ReAct paper to a production-grade implementation.
Tools are the interface between a model and its environment, and the most underrated part of harness design: this article breaks down what deserves to be a tool, how schemas and return values shape model behavior, the choice-overload problem of having too many tools, what the MCP protocol does and doesn't solve, the tradeoffs of general-purpose computer use tools, and the peculiar form of prompt-only tools like TodoWrite.
Aider is the open-source CLI pair-programming agent best suited for studying harness internals — the trade-offs of its edit format, its tree-sitter repo map, its lint/test feedback loops, and its deep git integration, each backed by verifiable benchmark data.
A deep dive into Claude Code's harness design — the minimal Unix-style toolset, the planning mechanisms behind TodoWrite and plan mode, CLAUDE.md memory files, subagents, hooks, MCP, and permission modes, plus why it chose the terminal over the IDE. The site's component framework, assembled in full inside a real product.
Codex is not "a model that answers directly" — it's a closed-loop system driven by a harness. This article breaks down, step by step, what happens and why at every stage of its interaction, drawing the clearest possible responsibility boundary between the harness and the LLM.
A teardown of ByteDance's Coze harness design — the "platform-style" route that turns the loop, tools, memory, planning, and subagents into visual product features, the opposite pole from Claude Code's CLI primitives, and a look at where the low-code harness ceiling sits.
An anatomy of the harness design behind an IDE-native agent — how Cursor supplies context for large codebases with Merkle tree indexing and embedding retrieval, feeds the model editor state only an IDE can provide, solves the engineering problem of getting edits right with a specially trained apply model, and how Tab completion and agent mode divide the work inside a single harness.
An anatomy of the harness behind Devin, the "fully autonomous software engineer" — the long-horizon architecture built from a shell/browser/editor/planner combination, the launch demos and the debunking controversy of 2024, the first-hand lessons on context engineering and model selection published on Cognition's engineering blog, and the boldest bet yet on the autonomy-vs-controllability spectrum of agent productization.
A teardown of Dify's harness design — packaging the agent loop, the RAG pipeline, tool plugins, and a model-adaptation layer into "self-hostable middleware" — and where it sits on the spectrum between LangGraph (a code framework) and Coze (hosted SaaS).
A deep dive into LangGraph's harness design — the "framework camp" approach that models agent control flow explicitly as a state graph, how its philosophy compares with the Claude Code-style free-running loop, and when to reach for a framework versus writing your own loop.
An anatomy of the harness behind the "general AI agent" Manus — the one-cloud-VM-per-task isolation architecture, the six context engineering lessons the team published and the cost logic behind them, the Wide Research experiment running a hundred agents in parallel, and the full arc from viral launch to Meta acquisition to a Chinese regulatory block — how a harness-first team used context engineering to fight model commoditization.
A deep dive into the harness architecture of the open-source agent platform OpenHands (formerly OpenDevin): the event stream, Docker sandbox runtime, CodeAct's unified action space, microagents and pluggable design — plus its academic value as a "harness research platform."
SWE-agent from the Princeton team is a model example of “harness design as research”: it introduced the concept of the Agent-Computer Interface (ACI) and used commands and feedback formats purpose-built for language models to achieve results on SWE-bench far beyond non-interactive baselines.
A close reading of 7 classic papers that shaped how agent harnesses are designed — ReAct, MRKL, Toolformer, Reflexion, Generative Agents, Voyager, and SWE-agent — covering their mechanisms, key results, and what they mean for system design.
A survey of seven frontier threads in 2025–2026 agent harness research—long-horizon execution, context engineering, memory infrastructure, multi-agent collaboration, self-improving agents, endogenous training, and evaluation discipline—with each paper's problem, method, and implications for harness design.
Common questions that come up when reading Agent Harness papers: whether success-rate numbers can be compared across papers, the conditions behind Reflexion's 91%, and how to read a benchmark paper.
Reading orders for three kinds of readers: a 4-hour quick start, a 2-week deep dive, and a 1-week route into engineering practice. Each path notes what to read, why, and where to head once you finish.
Entry point to the paper close-reading section: pick a reading path that matches your goal, or jump straight into the paper map, classic close readings, and the frontier.
An academic map of Agent Harness research: representative papers grouped by topic, plus a timeline of paradigm shifts, to help you place any new paper on the grid.
The complete hands-on path to building agent evals: what the unit, trajectory, and outcome layers each assert; how eval datasets grow out of real failures and why 20 cases are enough to start; where programmatic assertions end and LLM-as-judge begins; how to wire evals into CI for regression while trading off cost against sampling—and finally, a minimal eval script for mini_harness.
Implement a minimal agent harness from zero in one ~130-line, framework-free Python file — dissecting the five stages of context assembly, model calls, tool parsing and execution, result write-back, and human approval — then mounting todo planning, memory compression, subagents, and observability one at a time along a "symptom → component" path.
Which foundation should you build an agent on? This article lays out writing your own loop, code frameworks, and platforms along a single spectrum, compares LangGraph, CrewAI, AutoGen, OpenAI Agents SDK, Claude Agent SDK, Dify, and Coze on abstraction level, control, and lock-in risk, and closes with a scenario-based decision tree.
Ten high-frequency anti-patterns in agent harness development — framework-first design, tool sprawl, context hoarding, no termination conditions, swallowed errors, full-auto everything, prompt silver bullets, demo-driven development, blind spots on cost and latency, and premature multi-agent — each with symptoms, root cause, and the fix.
Actionable harness design principles distilled from the practice of Anthropic, Cognition, 12-Factor Agents, and mainstream coding agents — each with a rationale and a positive/negative example.
Three progressively harder agent portfolio projects for engineers moving into agent roles—a mini harness with a full evals suite, a production-grade MCP server, and a multi-agent / long-horizon task agent—each mapped to the most frequently requested job-description keywords, annotated with its technical points, acceptance criteria, and STAR resume bullets, with effort estimates at a light and a full tier.
Split a 130-line minimal harness into three independently runnable versions: v1 is just the agent loop skeleton, v2 adds a tool registry, output truncation, and an approval gate, and v3 rounds it out with todo planning and session-summary compaction—every version runs as-is, and every component maps to an observable symptom.
A deep dive into the config file agents read: what it really is — a persistent prompt layer, not enforced configuration — what to write and what to leave out, why shorter is more effective, how to organize it in layers and share it with a team, how it divides work with skills and hooks, plus a template you can adopt as-is.
基于 44 份真实在招 JD 的高频要求归纳出的 10 道 Agent Harness 系统设计题:每题附考察点、好答案骨架与踩坑回答,并逐一映射回手册对应章节。
从 44 份真实在招 JD 反推 Agent 工程岗的能力模型:七项核心能力的 JD 证据、简历好差写法对比、面试考点,三类候选人画像的对标分析与补课路径,附逐条自查清单。
2026 年 8 月国内外 22 家公司 44 个 Agent/LLM 工程在招岗位完整清单:官方全文、官方标题级、聚合站快照三级来源标注,附总览速查表与失效免责说明。
基于 44 个国内外真实在招 Agent/LLM 工程岗 JD 的词频统计,拆解高频知识点与国内外的考察差异,并把每个知识点逐一映射回本手册的对应章节,给出「掌握到什么程度算够」的标尺。
Reverse-engineering an Agent Harness learning path from 44 real, live job postings across 22 companies in global and Chinese markets (August 2026): the role landscape, the full picture of skill demands, and the industry signal that 'Agent Harness Engineer' has become an official job title.
The handbook's built-in community forum: one single feed for quick posts, with optional #question #check-in #musings #wishlist topic tags. Anonymous and registration-free.
一份策展式而非罗列式的 Agent Harness 阅读地图:只收录一手、有立场的工程博客、开源实现、协议规范与论文入口,每条都注明为什么值得读、什么时候读。
Agent Harness 领域核心术语表:智能体、挽具、上下文工程、工具、记忆、安全与评测等 50+ 术语的准确定义与辨析。
一份带评注的系统提示词逆向档案:收录 Claude Code、Cursor、Devin、Manus、v0、Perplexity 六款产品的公开逆向材料,逐条标注可信度,并用 harness 组件框架拆解每份提示词背后的设计取舍。