Appearance
Build a Minimal Harness from Scratch
In Common Pitfalls, the first anti-pattern is "over-frameworking": wanting to change one behavior means reading through three layers of abstraction, until the framework's abstractions hold you hostage. Dex Horthy, author of 12-Factor Agents, takes a more radical stance in the original repo: don't touch a framework yet — get the agent running with a raw API plus a while loop. Factor 8 is literally titled "Own your control flow."
This page turns that advice into code. The harness we'll write is about 130 lines of Python with zero framework dependencies — zero LLM SDKs, even. We first get the structure running with a rule-based MockLLM; swapping in a real model is just replacing one function signature. The full code lives in the repo at .scratch/mini_harness.py; this page is its layer-by-layer dissection.
This page's stance
Every harness component should "earn its place." We start with a deliberately bare-bones version, then — like adding medicines to a prescription — mount each component only when its symptom appears. This matches Anthropic's stance in Building effective agents: find the simplest thing that works, add complexity only when demonstrably necessary.
The minimal loop, in full
Strip away every product-grade nicety and the agent loop does exactly seven things (the pseudocode in What Is an Agent Harness? sketched this; here's the runnable version):
text
┌──────────────────────────────────────────────────────┐
│ Minimal agent loop │
│ │
│ ① Assemble ──> ② Call the ──> ③ Parse the │
│ context model output │
│ ▲ │ │
│ │ ┌────┴────┐ │
│ │ a tool call declares done │
│ │ │ │ │
│ ⑤ Write result <── ④ Execute │ ▼ │
│ back the tool return │
│ (with approval the │
│ gate) answer │
│ │
│ Outside the loop: step limit as a backstop, │
│ feeding errors back, output truncation │
└──────────────────────────────────────────────────────┘We'll write this into code in four steps. Each step solves exactly one problem — no premature optimization.
Step 1: the Agent Loop skeleton
python
import json
MAX_STEPS = 20 # max loop turns; keeps the agent from spinning out of control
SYSTEM = """You are a coding agent. On each step output exactly one JSON object:
- To call a tool: {"tool": "tool name", "args": {...}}
- To finish the task: {"done": "summary"}
First draft a todo plan, then execute step by step, deciding the next step
from the previous step's tool result."""python
def agent_loop(llm, task: str):
messages = [{"role": "user", "content": "Task: " + task}]
for step in range(1, MAX_STEPS + 1):
reply = llm(messages) # ② the model sees the full context, outputs a decision
try:
action = json.loads(reply) # ③ the harness parses the decision
except json.JSONDecodeError:
messages.append({"role": "user", "content": "Format error: output a single JSON object"})
continue
if "done" in action: # the model declares completion; the loop exits
return action["done"]
messages.append({"role": "assistant", "content": reply})
result = run_tool(action["tool"], action.get("args", {})) # ④ execute
messages.append({"role": "user", "content": f"[Tool {action['tool']} returned]\n{result}"})
return f"Reached max steps ({MAX_STEPS}); force-stopped" # the backstop brakeThree design decisions deserve a pause:
One: the protocol is JSON, not function calling. We have the model output a JSON object each step instead of relying on each API's native tool-calling fields. There's a real cost (the model may emit invalid JSON, so we catch json.JSONDecodeError and feed the format error back for a retry), but the payoff is structural transparency: nothing SDK-specific exists anywhere in the loop, so changing models or API vendors doesn't touch the harness. Production implementations use native tool calling plus JSON schema validation, but the mechanism is fully isomorphic — "the model emits a structured decision; the harness parses and executes it." This is exactly the protocol ReAct (arXiv:2210.03629) established: the alternation of thought and action is, at bottom, an agreement between the model's output format and the harness's executor.
Two: errors are context, not exceptions. When json.loads fails, we don't throw our way out — we push "format error" back into messages as a message. Likewise, exceptions raised by tools inside run_tool are caught and converted into text results. This is an iron law of harness engineering: an error the model can see is the only recoverable error. The inverse of silent failure (pitfall #5) is "every failure returns to the context, structured" — the same thing 12-Factor Agents' Factor 9 calls "Compact errors into context window."
Three: the stop condition must be the harness's hard constraint. The MAX_STEPS line looks unremarkable, but it's the only brake guaranteed to work from outside the model. The model's "I'm done" can be trusted — somewhat. Runaway loops burning tokens with no termination condition are a real, frequent incident category. Production harnesses add cost caps, time limits, and timeout detection, but step one is always the hard ceiling of a for loop.
Step 2: defining and registering tools
Tools are the model's only channel for changing the world in this loop. Note how almost recklessly naive our tool implementations are:
python
import os, subprocess
def tool_read(path: str) -> str:
if not os.path.exists(path):
return f"Error: file does not exist {path}"
with open(path, encoding="utf-8") as f:
return f.read()
def tool_write(path: str, content: str) -> str:
with open(path, "w", encoding="utf-8") as f:
f.write(content)
return f"Wrote {path} ({len(content)} chars)"
def tool_bash(command: str) -> str:
r = subprocess.run(command, shell=True, capture_output=True, text=True, timeout=30)
return (r.stdout + r.stderr).strip() or "(no output)"Three conventions matter more than the implementations:
- The signature is the interface. Each tool is a pure function: a few simple-typed arguments in, one string out. That return string is the model's "observation" — tool output always becomes text in the context anyway, so keeping it textual means any tool is plug-and-play.
- Errors return strings; they don't throw.
tool_readreturns"Error: file does not exist …"for a missing file. This turns "read a nonexistent path" into an observation the model can learn from, not a crash. - The registry is the single source of authority. Tools register in one dict; the loop only consults the table — never
eval:
python
TOOLS = { # name -> (function, is_dangerous)
"echo": (tool_echo, False),
"read": (tool_read, False),
"todo": (tool_todo, False),
"write": (tool_write, True), # Writing files and running commands change the world;
"bash": (tool_bash, True), # humans must hold veto power over them
}
def run_tool(name: str, args: dict) -> str:
if name not in TOOLS:
return f"Error: unknown tool {name}" # a hallucinated tool name gets a safe refusal
fn, dangerous = TOOLS[name]
...
try:
return truncate(str(fn(**args)))
except Exception as e: # a tool failure isn't the end of the world; feed the exception back
return f"Tool execution failed: {e}"The name not in TOOLS line is the first brick in the safety boundary: when the model hallucinates a tool name like delete_everything, the harness answers with one harmless line of text instead of a catastrophe. For tool counts, argument design, and why "tool explosion" is an anti-pattern, see Tools & MCP.
Step 3: the first line of defense on tool output — truncation
That truncate() inside run_tool is this page's first "not-a-toy" component:
python
MAX_OUTPUT_CHARS = 1000 # tool output truncation threshold
def truncate(text: str, limit: int = MAX_OUTPUT_CHARS) -> str:
"""Trim very long output head and tail: the beginning and end of files
and logs usually carry the most information; the middle can be dropped."""
if len(text) <= limit:
return text
head, tail = text[: limit // 2], text[-limit // 2 :]
return f"{head}\n... [truncated: {len(text) - limit} chars omitted] ...\n{tail}"Why is this essential? Because tool_bash("cat huge.log") will pour tens of MB of text into the context — blowing the window in one shot and torching the bill. Head-and-tail trimming is a crude but effective heuristic: the start and end of logs and files usually have the highest information density. The seed of all of context engineering is buried here — the harness decides what the model sees, and deciding what the model does not see matters just as much. Note that the truncation notice itself ("N chars omitted") is written into the return value: the model needs to know the data was cut, or it will confidently make mistakes based on incomplete observations.
Step 4: add interruption and approval
The last puzzle piece pulls the human from outside the loop back into it. That dangerous boolean in the TOOLS table finally earns its keep:
python
def run_tool(name: str, args: dict) -> str:
fn, dangerous = TOOLS[name]
if dangerous and os.environ.get("AUTO_APPROVE") != "1":
print(f"\n⚠️ Agent requests a dangerous operation: {name}({json.dumps(args)[:80]})")
if input("Allow? (y/n) ").strip().lower() != "y":
return "The user declined this operation"
...Only two design points:
- Tier by irreversibility, not by gut feel about tool names.
read/todowrite nothing — let them through;write/bashchange the world — stop them. This is the minimal implementation of "tiered interception by irreversibility" from Permissions, Safety & Human-in-the-Loop. Real coding agents make the tiers far finer: bash commands classified by allowlist/regex,writeconfined to the working directory, network access approved separately — but the axis of judgment is the same one: can this operation be undone? - A refusal is also an observation. When the user presses
n, the return value is"The user declined this operation", not an interrupt. The model can then change approach, explain its intent first, or give up — human–machine negotiation is modeled as an ordinary interaction inside the loop. TheAUTO_APPROVE=1environment variable is the escape hatch for batch runs, off by default.
At this point, the 130-line harness has every essential structure: context, loop, tools, approval, backstop. Run it with the MockLLM (printf 'y\ny\n' | python3 mini_harness.py) and you'll see it execute its prewritten script in order: build a todo → read the file → run wc -l → write the report → declare done, with one prompt for each of the two dangerous operations.
The incremental path: what to add, and when
Once the skeleton runs, the temptation is to pile every component on at once. That's wrong. The right posture is prescribe only when the symptom appears — every component should map to a pain point that actually occurs in production. Here's the symptom-to-prescription map, each expanded below:
| Symptom you observe | Component to mount | Deep dive |
|---|---|---|
| On long tasks the agent forgets the original goal and redoes finished work | Todo planning (explicit working memory) | Planning & Task Decomposition |
| As sessions grow, the agent gets dumber and more expensive | Context management beyond output truncation; memory compression | Context Engineering, Memory Systems |
| Single-threaded exploration pollutes the main context; attention scatters | Subagents | Subagents & Multi-Agent Orchestration |
| When something breaks, you can't answer "where did it go off the rails" | Structured logging and trajectory recording | Evaluation & Observability |
| More and more dangerous operations; your finger cramps from y/n | Fine-grained permissions and a rules engine | Permissions, Safety & Human-in-the-Loop |
Increment 1: todo planning
Our tool_todo is already a simplified version: the model submits its task list as a tool call, and the rendered result is written back into the context. It has zero runtime logic — the harness doesn't check it, nag about it, or enforce it — its entire mechanism of action is cognitive offloading: the list lives in the context, visible at every decision round. Claude Code's TodoWrite is the mature form of this idea (a three-state machine, full-overwrite semantics, discipline enforced via the system prompt); the full dissection is in Planning & Task Decomposition. When to mount: tasks reliably exceed three steps, or multiple requests run in parallel. On single-step tasks it's pure ceremony.
Increment 2: memory compression and context management
truncate only solves the size of a single output; it can't solve session-level bloat: after dozens of tool calls, early key information (the user's original request, the reasoning behind mid-course decisions) is buried. The next step is layering: rolling summaries (compress old turns into a summary, keep the last N turns verbatim), externalized notes (the agent proactively writes key findings to a file and reads them back when needed), and cross-session persistent memory files. These are the core topics of Context Engineering and Memory Systems respectively. When to mount: you observe "longer sessions = dumber and more expensive" — the context-garbage accumulation of pitfall #3.
Increment 3: subagents
When the task spawns "highly exploratory but tangential" sub-questions ("look up how this library is used," "get to the bottom of this error"), having the main agent search personally pours dozens of retrieval trajectories into the main context. The fix is spawning a subagent: give it a narrow task and a small context, let it run its own loop, and return only conclusions to the main loop. The main agent's context stays clean — a subagent is essentially a "tool call" that may loop internally many times. In implementation it's a recursive call to agent_loop plus a result aggregation — but when to split, how deep, and whether subagents share memory are real questions; see Subagents & Multi-Agent Orchestration. When to mount: exploration trajectories begin visibly diluting the main line's attention — and note this typically comes later than intuition expects; going multi-agent too early is pitfall #10.
Increment 4: observability
Our loop already contains the most primitive observation: a print of each step's model output and tool result. When that stops being enough (the classic scenario: a user reports "the agent did something dumb in the middle of the night," and you can't answer "at which step did it go off the rails"), swap the print for structured trajectory recording: per-step input hashes, model decisions, tool calls and results, latency and token counts, persisted as replayable JSONL. Trajectories are the elementary particles of harness debugging and the data source for evals; see Evaluation & Observability. When to mount: the first time you need to answer "why" and can't.
When to stop
This incremental path has no end — you could keep adding sandboxes, parallel tool calls, checkpoint-resume, evals. So the scarcer skill is knowing when to stop. Three meta-rules from this site:
- Complexity is bought with symptoms, not imagined into existence. For every component you add, be able to name the real symptom it treats. Anthropic's own words: find the simplest thing that works; add complexity only when necessary.
- Keep tools general and primitive. Claude Code's toolset has stayed at "read, write, grep, bash" for a long time (see Case Study: Claude Code), leaving intelligence to the model rather than the harness — the thicker the harness, the more must be rewritten when the next model ships.
- Write evals before adding capabilities. Before adding a feature to the harness, have a minimal eval set that can measure whether it helps — otherwise you can't tell "improvement" from "fidgeting" (pitfall #8: demo-driven development).
More decision frameworks live in Harness Design Principles; if you want to see these principles embodied in real products, the case-study section has two contrasting samples: SWE-agent (the exemplar of a minimal scaffold) and OpenHands (a full-featured open-source harness).
The next step after building it
Hook this minimal harness to a real model (replace the MockLLM — about ten lines of API calls), then pick one of your own small tasks and run it three times: first bare, second with only the todo tool, third with truncation and approval added. Compare the three trajectories, and your understanding of "what each component actually buys" will surpass reading ten articles.
Further reading
- The Agent Loop — this page's loop at production scale: stop conditions, error handling, concurrency
- Context Engineering — truncation is just the entrance; summarization, offloading, and injection are the full picture
- Tools & MCP — tool counts, naming, argument design, and MCP
- Planning & Task Decomposition — the complete dissection of the TodoWrite mechanism
- Memory Systems — from in-session checklists to cross-session persistent memory
- Subagents & Multi-Agent Orchestration — context isolation and task delegation
- Permissions, Safety & Human-in-the-Loop — approval tiers, sandboxes, and irreversible operations
- Harness Design Principles — eight rules for deciding what to add and what not to
- Common Pitfalls & Anti-Patterns — the cautionary tale for every step on this page
References
- 12-Factor Agents (humanlayer/12-factor-agents) — the source of "Own your control flow" and "Compact errors into context window"
- Anthropic: Building effective agents — the original phrasing of "start simple; add complexity only when necessary"
- ReAct: Synergizing Reasoning and Acting in Language Models (arXiv:2210.03629) — the original paper on the thought–action alternation protocol
- Case Study: Claude Code — the product design of a minimal toolset