Skip to content

Progressive Tutorial: Three Running Versions

At a glance Split a 130-line minimal harness into three independently runnable versions: v1 is just the agent loop skeleton, v2 adds a tool registry, output truncation, and an approval gate, and v3 rounds it out with todo planning and session-summary compaction—every version runs as-is, and every component maps to an observable symptom.

Progressive Tutorial: Three Running Versions ​

Building a Minimal Harness from Scratch walks through the 130 lines of .scratch/mini_harness.py layer by layer—that was "reading comprehension." This page takes a different approach to the same code: break the finished product back into three progressive versions, each one a Python file you can run directly. Run each version by hand, watch the output, compare the traces, and the reason each component exists will surface on its own.

All three files live in the repo's examples/ directory:

FileAdded over the previous versionSymptom it addresses
v1_minimal_loop.py(Starting point) the five stages of a minimal agent loopNo loop, no agent
v2_with_tools.pyTool registry, output truncation, approval gateTools that are hard to maintain, context flooded shut, dangerous operations with nobody at the switch
v3_with_planning_memory.pyTodo planning tool, session-summary compactionGoal drift on long tasks; the longer the session, the dumber and pricier it gets

All three files depend only on the Python standard library and need no API key—the model is played by a MockLLM driven by a pre-written script. That's not laziness; it's a deliberate piece of teaching methodology: get the harness's structure working first, and swapping in a real model is just a matter of changing one function signature. A mock model lets you iterate on harness behavior at zero cost, which is also the minimal form of the "regression-test against fixed traces" idea in Evals in Practice.

How to Run ​

bash
# v1: no interaction at all—just run it
python3 examples/v1_minimal_loop.py

# v2: stops twice waiting for approval—pipe in two y's; or just auto-approve outright
printf 'y\ny\n' | python3 examples/v2_with_tools.py
AUTO_APPROVE=1 python3 examples/v2_with_tools.py

# v3: a longer script (9 steps) that genuinely triggers one round of context compaction
AUTO_APPROVE=1 python3 examples/v3_with_planning_memory.py

All three scripts write their demo files into a temp directory created by tempfile.mkdtemp(), so they never pollute your workspace. The module docstring at the top of each file spells out "what this version adds," so you can read it side by side with the source.

Where to plug in a real model

A comment at the end of each file marks the hook: replace MockLLM with any callable whose signature is (messages) -> str (about ten lines of urllib.request for an OpenAI-compatible API) and not one other line of the harness changes. The same goes for v3's summarize()—swap in a single LLM summarization call and you have the very same mechanism as Claude Code's /compact.

v1: The Bare-Skeleton Minimal Loop ​

v1_minimal_loop.py is about 110 lines. The entire harness is one agent_loop function plus two inline tools (echo, read). It runs the complete five-stage cycle:

text
┌──────────────────────────────────────────────────────────────┐
│                    v1: minimal agent loop                    │
│                                                              │
│ ① Assemble context ──> ② Call model ──> ③ Parse JSON decision │
│      ▲                                             │         │
│      │                                      ┌──────┴──────┐  │
│      │                                     tool call  done   │
│      │                                           │       │  │
│  ⑤ write result back <─────────────── ④ run tool│       ▼  │
│                                  (if/elif dispatch)          │
│                                      return the final answer │
│                                                              │
│  Outside the loop: MAX_STEPS hard cap, malformed             │
│  output is fed back to the model                             │
└──────────────────────────────────────────────────────────────┘

The core code is just this (abridged):

python
def agent_loop(llm, task: str):
    messages = [
        {"role": "system", "content": SYSTEM},
        {"role": "user", "content": "Task: " + task},
    ]
    for step in range(1, MAX_STEPS + 1):
        reply = llm(messages)                      # ② the one and only "model moment"
        action = json.loads(reply)                 # ③ parse (on failure, feed the format error back)
        if "done" in action:
            return action["done"]                  # the model declares it's done
        messages.append({"role": "assistant", "content": reply})
        if action["tool"] == "echo":                    # ④ execute (v1 dispatches inline)
            result = tool_echo(**action.get("args", {}))
        elif action["tool"] == "read":
            result = tool_read(**action.get("args", {}))
        else:
            result = f"Error: unknown tool {action['tool']}"
        messages.append({"role": "user", "content": f"[Tool {action['tool']} returned]\n{result}"})  # ⑤ write back

Run it and you'll see it finish in three steps: read notes.txt → echo the count → declare done. This version covers every foundational concept of The Agent Loop and nothing else—that's intentional.

What makes v1 teachable is what it's missing. Stare at the code and the output for a minute and three symptoms practically announce themselves:

  1. Tool dispatch is a stretch of if/elif—unreadable once you reach ten tools, and when the model hallucinates a tool name that doesn't exist, the rejection logic is scattered through the dispatch;
  2. tool_read writes back exactly what it reads, verbatim—read a few dozen megabytes of log and the context is flooded shut in one call;
  3. Would you dare add world-changing tools like write or bash now? You wouldn't. Add them and you've got an unattended disaster.

Hold onto those three symptoms—they are the reason v2 exists.

v2: Tool Registry, Truncation, and Approval ​

v2_with_tools.py hangs three components off v1's loop, and agent_loop itself changes by only one line (inline dispatch becomes a run_tool table lookup):

Component one: the tool registry. Tools are registered centrally in the TOOLS dict; the loop only consults the table—it never evals:

python
TOOLS = {  # name -> (function, dangerous or not)
    "echo":  (tool_echo,  False),
    "read":  (tool_read,  False),
    "write": (tool_write, True),   # can change the world,
    "bash":  (tool_bash,  True),   # so a human must hold veto power
}

def run_tool(name: str, args: dict) -> str:
    if name not in TOOLS:
        return f"Error: unknown tool {name}"   # a hallucinated tool name, safely rejected
    ...

name not in TOOLS is the first brick in the safety boundary. The registry is also the mount point for every tool-level policy that comes later (approval tiers, argument validation, per-tool truncation limits)—see Tools & MCP for the deep dive.

Component two: output truncation. truncate() keeps the head and the tail and writes "N characters omitted" into the return value—the model has to know the data was cut, or it will confidently go wrong based on a partial picture of the world. This is the seed of Context Engineering: the harness decides what the model sees, and what you don't let it see matters just as much.

Component three: the approval gate. That dangerous boolean in the registry takes effect here:

python
    if dangerous and os.environ.get("AUTO_APPROVE") != "1":
        print(f"\n⚠️  Agent requests a dangerous operation: {name}(...)")
        if input("Allow? (y/n) ").strip().lower() != "y":
            return "The user rejected this operation"   # a rejection is an observation, not an interruption

Tiering goes by irreversibility rather than by gut feel on tool names: read-only tools pass, world-changing ones get stopped. A rejection is modeled as an ordinary tool-result write-back, so the model can change course based on it—the human-machine negotiation stays inside the loop. This is the minimal implementation of "tier interventions by irreversibility" from Permissions, Safety & Human-in-the-Loop.

Why these three components ship in the same version

Because together they are the minimal safety kit for "letting an agent touch the real world": the registry governs what can be called, truncation governs how much it can see, and approval governs whether it can act. Adding write/bash without these three is handing a toddler a pair of scissors. The reverse holds too: pairing v1's two read-only tools with all three is pure over-engineering—that's what "timing" means.

v3: Todo Planning and Session-Summary Compaction ​

v3_with_planning_memory.py deals with a different class of symptom—loss of control along the time dimension: goals drift as tasks stretch out, sessions bloat. Its two new components correspond to the Planning & Task Decomposition and Memory Systems chapters.

Component one: the todo planning tool. The model submits its task list as an ordinary tool call, and the harness renders it back into context:

python
def tool_todo(items: list) -> str:
    lines = [f"[{'x' if i.get('done') else ' '}] {i['task']}" for i in items]
    return "Plan updated:\n" + "\n".join(lines)

Note that it has no runtime logic whatsoever—the harness doesn't check, doesn't nag, doesn't enforce. Its entire mechanism is cognitive offloading: writing the list forces the model to think through the task in structured form first, the list lives in context as an anchor for every round of decisions, and the state-update step turns "review progress" into a fixed rhythm of the loop. Claude Code's TodoWrite is the mature form of this idea (a three-state state machine, full overwrite, discipline written into the system prompt); for the full breakdown see Planning & Task Decomposition.

Component two: session-summary compaction. truncate governs the size of a single tool output; it can't govern session-level bloat. v3 adds one event-driven check at the end of the loop:

python
        if len(messages) > COMPACT_THRESHOLD:   # the only new line
            messages = compact(messages)

compact() follows a "keep the head, compress the middle, keep the tail" policy: the system prompt and the original task stay untouched, old turns are compressed into a single summary, and the most recent KEEP_RECENT turns keep their original text. The demo summarize() is a rule-based implementation (one line per turn: which tool was called and the start of its result); a production implementation swaps in a single LLM call—Claude Code's /compact is exactly this mechanism. The trigger for compaction has to be owned by the harness; counting on the model to compact its own context out of its own initiative is unreliable, as Context Engineering explains in depth.

You'll see compaction happen live when you run v3

The demo script runs 9 steps; the message count crosses the threshold (14 messages) after step 7, and the terminal prints a line like 🗜️ [context compaction] 16 messages → 9 (8 old turns compacted). Turn COMPACT_THRESHOLD up and run it again—the compaction disappears. Same harness, behavior set by a parameter: that's exactly what it means for a harness to be a "system" rather than a "script."

The Three Versions Side by Side: Complexity Is Paid for in Symptoms ​

Dimensionv1: minimal loopv2: + tool systemv3: + planning & memory
Lines of code (approx.)110140190
Components added—Registry, truncation, approvalTodo tool, summary compaction
Tasks it can handleSingle-step, read-onlyMulti-step, world-changingLong tasks, long sessions
Symptoms that motivated the additionsNo agentTools spiraling, context flooded, operations ungatedGoal drift, session bloat
Corresponding chaptersThe Agent LoopTools, Permissions, Context EngineeringPlanning, Memory

All three versions share one design philosophy (the same Anthropic "Building effective agents" principle this site cites again and again): find the simplest thing that works, and add complexity only when it is genuinely necessary. Every component must be able to name the real, actually-observed symptom it treats—complexity is paid for in symptoms, not in imagination. This "wait for the symptom, then prescribe" framework is systematized in Design Principles, and the cautionary tales are collected in Common Pitfalls.

Where to go next, hands-on

  1. Swap MockLLM for a real model (the hook is in the comment at the end of each file) and compare the traces of all three versions on the same task;
  2. Manufacture failures on purpose: make read open a file that doesn't exist, press n at the approval prompt, set COMPACT_THRESHOLD to 8—then watch how the agent recovers from each kind of failure;
  3. Write a fixed set of scripts for each of the three versions as regression tests, then try modifying the harness (say, adding a whitelist to bash) and use trace comparison to check whether your "improvement" was real progress or just churn—that's the on-ramp to Evals in Practice.

Further Reading ​

References ​