Appearance
Progressive Tutorial: Three Running Versions
Building a Minimal Harness from Scratch walks through the 130 lines of .scratch/mini_harness.py layer by layer—that was "reading comprehension." This page takes a different approach to the same code: break the finished product back into three progressive versions, each one a Python file you can run directly. Run each version by hand, watch the output, compare the traces, and the reason each component exists will surface on its own.
All three files live in the repo's examples/ directory:
| File | Added over the previous version | Symptom it addresses |
|---|---|---|
v1_minimal_loop.py | (Starting point) the five stages of a minimal agent loop | No loop, no agent |
v2_with_tools.py | Tool registry, output truncation, approval gate | Tools that are hard to maintain, context flooded shut, dangerous operations with nobody at the switch |
v3_with_planning_memory.py | Todo planning tool, session-summary compaction | Goal drift on long tasks; the longer the session, the dumber and pricier it gets |
All three files depend only on the Python standard library and need no API key—the model is played by a MockLLM driven by a pre-written script. That's not laziness; it's a deliberate piece of teaching methodology: get the harness's structure working first, and swapping in a real model is just a matter of changing one function signature. A mock model lets you iterate on harness behavior at zero cost, which is also the minimal form of the "regression-test against fixed traces" idea in Evals in Practice.
How to Run
bash
# v1: no interaction at all—just run it
python3 examples/v1_minimal_loop.py
# v2: stops twice waiting for approval—pipe in two y's; or just auto-approve outright
printf 'y\ny\n' | python3 examples/v2_with_tools.py
AUTO_APPROVE=1 python3 examples/v2_with_tools.py
# v3: a longer script (9 steps) that genuinely triggers one round of context compaction
AUTO_APPROVE=1 python3 examples/v3_with_planning_memory.pyAll three scripts write their demo files into a temp directory created by tempfile.mkdtemp(), so they never pollute your workspace. The module docstring at the top of each file spells out "what this version adds," so you can read it side by side with the source.
Where to plug in a real model
A comment at the end of each file marks the hook: replace MockLLM with any callable whose signature is (messages) -> str (about ten lines of urllib.request for an OpenAI-compatible API) and not one other line of the harness changes. The same goes for v3's summarize()—swap in a single LLM summarization call and you have the very same mechanism as Claude Code's /compact.
v1: The Bare-Skeleton Minimal Loop
v1_minimal_loop.py is about 110 lines. The entire harness is one agent_loop function plus two inline tools (echo, read). It runs the complete five-stage cycle:
text
┌──────────────────────────────────────────────────────────────┐
│ v1: minimal agent loop │
│ │
│ ① Assemble context ──> ② Call model ──> ③ Parse JSON decision │
│ ▲ │ │
│ │ ┌──────┴──────┐ │
│ │ tool call done │
│ │ │ │ │
│ ⑤ write result back <─────────────── ④ run tool│ ▼ │
│ (if/elif dispatch) │
│ return the final answer │
│ │
│ Outside the loop: MAX_STEPS hard cap, malformed │
│ output is fed back to the model │
└──────────────────────────────────────────────────────────────┘The core code is just this (abridged):
python
def agent_loop(llm, task: str):
messages = [
{"role": "system", "content": SYSTEM},
{"role": "user", "content": "Task: " + task},
]
for step in range(1, MAX_STEPS + 1):
reply = llm(messages) # ② the one and only "model moment"
action = json.loads(reply) # ③ parse (on failure, feed the format error back)
if "done" in action:
return action["done"] # the model declares it's done
messages.append({"role": "assistant", "content": reply})
if action["tool"] == "echo": # ④ execute (v1 dispatches inline)
result = tool_echo(**action.get("args", {}))
elif action["tool"] == "read":
result = tool_read(**action.get("args", {}))
else:
result = f"Error: unknown tool {action['tool']}"
messages.append({"role": "user", "content": f"[Tool {action['tool']} returned]\n{result}"}) # ⑤ write backRun it and you'll see it finish in three steps: read notes.txt → echo the count → declare done. This version covers every foundational concept of The Agent Loop and nothing else—that's intentional.
What makes v1 teachable is what it's missing. Stare at the code and the output for a minute and three symptoms practically announce themselves:
- Tool dispatch is a stretch of
if/elif—unreadable once you reach ten tools, and when the model hallucinates a tool name that doesn't exist, the rejection logic is scattered through the dispatch; tool_readwrites back exactly what it reads, verbatim—reada few dozen megabytes of log and the context is flooded shut in one call;- Would you dare add world-changing tools like
writeorbashnow? You wouldn't. Add them and you've got an unattended disaster.
Hold onto those three symptoms—they are the reason v2 exists.
v2: Tool Registry, Truncation, and Approval
v2_with_tools.py hangs three components off v1's loop, and agent_loop itself changes by only one line (inline dispatch becomes a run_tool table lookup):
Component one: the tool registry. Tools are registered centrally in the TOOLS dict; the loop only consults the table—it never evals:
python
TOOLS = { # name -> (function, dangerous or not)
"echo": (tool_echo, False),
"read": (tool_read, False),
"write": (tool_write, True), # can change the world,
"bash": (tool_bash, True), # so a human must hold veto power
}
def run_tool(name: str, args: dict) -> str:
if name not in TOOLS:
return f"Error: unknown tool {name}" # a hallucinated tool name, safely rejected
...name not in TOOLS is the first brick in the safety boundary. The registry is also the mount point for every tool-level policy that comes later (approval tiers, argument validation, per-tool truncation limits)—see Tools & MCP for the deep dive.
Component two: output truncation. truncate() keeps the head and the tail and writes "N characters omitted" into the return value—the model has to know the data was cut, or it will confidently go wrong based on a partial picture of the world. This is the seed of Context Engineering: the harness decides what the model sees, and what you don't let it see matters just as much.
Component three: the approval gate. That dangerous boolean in the registry takes effect here:
python
if dangerous and os.environ.get("AUTO_APPROVE") != "1":
print(f"\n⚠️ Agent requests a dangerous operation: {name}(...)")
if input("Allow? (y/n) ").strip().lower() != "y":
return "The user rejected this operation" # a rejection is an observation, not an interruptionTiering goes by irreversibility rather than by gut feel on tool names: read-only tools pass, world-changing ones get stopped. A rejection is modeled as an ordinary tool-result write-back, so the model can change course based on it—the human-machine negotiation stays inside the loop. This is the minimal implementation of "tier interventions by irreversibility" from Permissions, Safety & Human-in-the-Loop.
Why these three components ship in the same version
Because together they are the minimal safety kit for "letting an agent touch the real world": the registry governs what can be called, truncation governs how much it can see, and approval governs whether it can act. Adding write/bash without these three is handing a toddler a pair of scissors. The reverse holds too: pairing v1's two read-only tools with all three is pure over-engineering—that's what "timing" means.
v3: Todo Planning and Session-Summary Compaction
v3_with_planning_memory.py deals with a different class of symptom—loss of control along the time dimension: goals drift as tasks stretch out, sessions bloat. Its two new components correspond to the Planning & Task Decomposition and Memory Systems chapters.
Component one: the todo planning tool. The model submits its task list as an ordinary tool call, and the harness renders it back into context:
python
def tool_todo(items: list) -> str:
lines = [f"[{'x' if i.get('done') else ' '}] {i['task']}" for i in items]
return "Plan updated:\n" + "\n".join(lines)Note that it has no runtime logic whatsoever—the harness doesn't check, doesn't nag, doesn't enforce. Its entire mechanism is cognitive offloading: writing the list forces the model to think through the task in structured form first, the list lives in context as an anchor for every round of decisions, and the state-update step turns "review progress" into a fixed rhythm of the loop. Claude Code's TodoWrite is the mature form of this idea (a three-state state machine, full overwrite, discipline written into the system prompt); for the full breakdown see Planning & Task Decomposition.
Component two: session-summary compaction. truncate governs the size of a single tool output; it can't govern session-level bloat. v3 adds one event-driven check at the end of the loop:
python
if len(messages) > COMPACT_THRESHOLD: # the only new line
messages = compact(messages)compact() follows a "keep the head, compress the middle, keep the tail" policy: the system prompt and the original task stay untouched, old turns are compressed into a single summary, and the most recent KEEP_RECENT turns keep their original text. The demo summarize() is a rule-based implementation (one line per turn: which tool was called and the start of its result); a production implementation swaps in a single LLM call—Claude Code's /compact is exactly this mechanism. The trigger for compaction has to be owned by the harness; counting on the model to compact its own context out of its own initiative is unreliable, as Context Engineering explains in depth.
You'll see compaction happen live when you run v3
The demo script runs 9 steps; the message count crosses the threshold (14 messages) after step 7, and the terminal prints a line like 🗜️ [context compaction] 16 messages → 9 (8 old turns compacted). Turn COMPACT_THRESHOLD up and run it again—the compaction disappears. Same harness, behavior set by a parameter: that's exactly what it means for a harness to be a "system" rather than a "script."
The Three Versions Side by Side: Complexity Is Paid for in Symptoms
| Dimension | v1: minimal loop | v2: + tool system | v3: + planning & memory |
|---|---|---|---|
| Lines of code (approx.) | 110 | 140 | 190 |
| Components added | — | Registry, truncation, approval | Todo tool, summary compaction |
| Tasks it can handle | Single-step, read-only | Multi-step, world-changing | Long tasks, long sessions |
| Symptoms that motivated the additions | No agent | Tools spiraling, context flooded, operations ungated | Goal drift, session bloat |
| Corresponding chapters | The Agent Loop | Tools, Permissions, Context Engineering | Planning, Memory |
All three versions share one design philosophy (the same Anthropic "Building effective agents" principle this site cites again and again): find the simplest thing that works, and add complexity only when it is genuinely necessary. Every component must be able to name the real, actually-observed symptom it treats—complexity is paid for in symptoms, not in imagination. This "wait for the symptom, then prescribe" framework is systematized in Design Principles, and the cautionary tales are collected in Common Pitfalls.
Where to go next, hands-on
- Swap
MockLLMfor a real model (the hook is in the comment at the end of each file) and compare the traces of all three versions on the same task; - Manufacture failures on purpose: make
readopen a file that doesn't exist, pressnat the approval prompt, setCOMPACT_THRESHOLDto 8—then watch how the agent recovers from each kind of failure; - Write a fixed set of scripts for each of the three versions as regression tests, then try modifying the harness (say, adding a whitelist to
bash) and use trace comparison to check whether your "improvement" was real progress or just churn—that's the on-ramp to Evals in Practice.
Further Reading
- Building a Minimal Harness from Scratch—the layer-by-layer walkthrough of the same code; this page and that one are two sides of the same coin
- The Agent Loop—what v1's loop grows into in a production-grade implementation
- Tools & MCP—beyond the registry: tool count, naming, argument design, and MCP
- Permissions, Safety & Human-in-the-Loop—the full tiering system behind v2's approval gate
- Context Engineering—beyond truncation and compaction: injection, externalization, isolation
- Planning & Task Decomposition—the complete teardown of the TodoWrite mechanism and three planning modes
- Memory Systems—from within-session summaries to persistent cross-session memory
- Evals in Practice—how to measure how good the three harnesses you just built actually are
- Case Study: SWE-agent—a real product sample of the minimal-scaffold route
References
- Anthropic: Building effective agents—the original statement of "find the simplest thing that works; add complexity only when necessary"
- 12-Factor Agents (humanlayer/12-factor-agents)—the source of "Own your control flow" and "Compact errors into context window"
- ReAct: Synergizing Reasoning and Acting in Language Models (arXiv:2210.03629)—the original paper on the interleaved think-act protocol