Skip to content

Build a Minimal Agent from Scratch

At a glance Hand-write a working minimal Agent with nothing but the LLM API—no frameworks. Start from a 50-line Agent Loop, then layer on tool error feedback, a step cap, todo planning, and context compression, with complete runnable code and the expected traces at every step.

Build a Minimal Agent from Scratch ​

The biggest trap when learning Agents: you pick up a pile of frameworks, but when someone asks "what actually is an Agent?", you can't answer. This page takes the most direct route—no LangChain, no framework of any kind, just the LLM API and a while loop, building a working Agent from zero. By the end you'll understand that what an Agent framework packages up is essentially the few hundred lines you're about to write, plus a pile of engineering details.

The page iterates through four versions, each solving one real pain point:

v0  Minimal Agent Loop             → runs, but can spin forever and can crash
v1  Tool error feedback + step cap → self-healing, can't run away
v2  todo planning tool             → long tasks stay on track
v3  Context compression            → long sessions don't blow the context window

You only need three things: Python 3.10+, pip install openai (1.x), and an API key for a model that supports tool calling. The code uses the OpenAI Chat Completions format, the de facto standard—it runs unchanged against OpenAI's official API as well as OpenAI-compatible endpoints like DeepSeek, Qwen, and Kimi; just swap the base_url.

1. First, Get Clear on the Minimal Core of an Agent ​

Strip away all the framework packaging and an Agent's core is a loop:

        ┌──────────────────────────────────────────┐
        │                                          │
        ▼                                          │
  ┌───────────┐   tool_calls   ┌────────────┐      │
  │    LLM    │───────────────▶│ Run tools  │      │
  │ (with     │                │ (your code)│      │
  │  context) │◀───────────────┴────────────┘      │
  └───────────┘   feed tool results back into      │
        │           messages                       │
        └── no more tool calls = final answer → exit the loop ──┘

That is the entire essence of the Agent Loop. Each iteration does only four things:

  1. Send the model the "conversation history + tool schemas";
  2. The model either returns final text or one or more tool_calls (function name + JSON arguments);
  3. Your code executes the requested functions and appends the results to the history as role: "tool" messages;
  4. Back to step 1, until the model stops calling tools.

A key insight

LLMs never "execute" tools. They only generate structured text saying "I want to call read_file with arguments {"path": "notes.txt"}". It's your Python code that actually runs it—which means a tool's permission boundary, timeouts, and error handling are all your responsibility, not the model's. This is exactly why Tools & MCP and Agent Security keep hammering on boundary control.

2. v0: A 50-line Agent Loop ​

Here is the complete, runnable v0. It has two real tools: read_file (read a file) and run_shell (run whitelisted read-only commands). Note the two security decisions—these are baseline requirements, not optional polish:

  • read_file hard-restricts paths to workspace/ to prevent directory traversal (a model-supplied ../../etc/passwd gets rejected);
  • run_shell uses a command allowlist plus shell=False (arguments split via shlex.split and passed as an array), ruling out injection and write operations.
python
# mini_agent_v0.py — a minimal Agent that actually runs
# Dependencies: pip install openai; environment variable OPENAI_API_KEY
import json
import shlex
import subprocess
from pathlib import Path

from openai import OpenAI

client = OpenAI()  # compatible endpoint example: OpenAI(base_url="https://api.deepseek.com", api_key="...")
MODEL = "gpt-5-mini"  # any model that supports tool calling will do

# Tools may only operate inside this directory, so the model can't wander the disk
WORKSPACE = Path("./workspace").resolve()

SYSTEM_PROMPT = """You are a file-operations assistant.
- Use read_file to view file contents and run_shell to run read-only commands (ls/cat/grep/find/wc, etc.).
- All paths are relative to the working directory workspace/.
- Gather enough information before answering; once you have enough, give your final answer directly instead of calling more tools.
"""

# Tool schemas: the model relies on descriptions to decide when to call and how to fill in arguments—clear writing is productivity
TOOLS = [
    {
        "type": "function",
        "function": {
            "name": "read_file",
            "description": "Read the contents of a text file inside the working directory",
            "parameters": {
                "type": "object",
                "properties": {
                    "path": {"type": "string", "description": "File path relative to workspace"},
                },
                "required": ["path"],
            },
        },
    },
    {
        "type": "function",
        "function": {
            "name": "run_shell",
            "description": "Run a read-only shell command (allowlist: ls/cat/grep/find/pwd/wc/head/tail/echo)",
            "parameters": {
                "type": "object",
                "properties": {
                    "command": {"type": "string", "description": "The command to run"},
                },
                "required": ["command"],
            },
        },
    },
]

ALLOWED_CMDS = {"ls", "cat", "grep", "find", "pwd", "wc", "head", "tail", "echo"}


def safe_path(p: str) -> Path:
    """Confine model-supplied paths to WORKSPACE to prevent directory traversal"""
    full = (WORKSPACE / p).resolve()
    if not str(full).startswith(str(WORKSPACE)):
        raise ValueError(f"path outside workspace: {p}")
    return full


def read_file(path: str) -> str:
    p = safe_path(path)
    if not p.is_file():
        return f"error: file not found: {path}"
    return p.read_text(encoding="utf-8", errors="replace")[:4000]  # truncate to protect the context


def run_shell(command: str) -> str:
    argv = shlex.split(command)
    if not argv or argv[0] not in ALLOWED_CMDS:
        return f"error: command rejected (not in allowlist): {command}"
    try:
        out = subprocess.run(argv, cwd=WORKSPACE, capture_output=True, text=True, timeout=10)
        return (out.stdout + out.stderr)[:4000] or "(no output)"
    except subprocess.TimeoutExpired:
        return "error: command timed out (10s)"
    except FileNotFoundError:
        return f"error: command not found: {argv[0]}"


DISPATCH = {"read_file": read_file, "run_shell": run_shell}


def run(task: str) -> str:
    messages = [
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": task},
    ]
    while True:  # v0 deliberately adds no guardrails—watch what happens first
        resp = client.chat.completions.create(model=MODEL, messages=messages, tools=TOOLS)
        msg = resp.choices[0].message
        # The model's raw reply must go into the history; tool messages pair with it via tool_call_id
        messages.append(msg.model_dump(exclude_none=True))

        if not msg.tool_calls:  # stop condition: no more tool calls = final answer
            return msg.content

        for tc in msg.tool_calls:  # a single turn may call several tools in parallel
            args = json.loads(tc.function.arguments)
            result = DISPATCH[tc.function.name](**args)
            print(f"  [tool] {tc.function.name}({args}) -> {result[:60]}...")
            messages.append({"role": "tool", "tool_call_id": tc.id, "content": result})


if __name__ == "__main__":
    WORKSPACE.mkdir(exist_ok=True)
    (WORKSPACE / "notes.txt").write_text("Goal this month: get the mini agent running\nBudget: 3,000 yuan\n", encoding="utf-8")
    print(run("What budget is written in the notes.txt in workspace?"))

A typical run (details vary each time, but the skeleton is the same):

$ python mini_agent_v0.py
  [tool] run_shell({'command': 'ls'}) -> notes.txt...
  [tool] read_file({'path': 'notes.txt'}) -> Goal this month: get the mini agent running...
The budget written in notes.txt is 3,000 yuan.

A few points worth pausing on:

  • msg.model_dump(exclude_none=True) can't be skipped or casually changed. The Chat Completions protocol requires the assistant's tool_calls message and the subsequent role: "tool" messages to pair up one-to-one via tool_call_id; miss one or get the order wrong and you get a 400. This is the most common beginner error by far.
  • The stop condition is simply "the model stops calling tools." You don't need an explicit finish tool—the loop ends naturally when the model outputs plain text. This is also why the system prompt says "once you have enough information, answer directly."
  • Don't pass temperature to reasoning models like the GPT-5 series—the API errors out on it. Omit the parameter and let the model use its default.
  • Cost: mid-2026 pricing for mini/nano-tier models is on the order of a few dozen cents per million input tokens; running every experiment on this page typically costs less than a cent. Run freely.

Why Chat Completions instead of the newer Responses API

OpenAI's docs now push the Responses API (flattened tool schemas, built-in tool search, and other new features), but Chat Completions is the lowest common denominator across all OpenAI-compatible endpoints—learn this message structure once and you can switch vendors without changing code. For new production projects, evaluate the Responses API.

3. v1: Tool Error Feedback and a Step Cap ​

v0 handles simple tasks fine; harder tasks expose two fatal flaws:

  1. Any exception crashes the whole program. The model emits malformed JSON arguments, passes a nonexistent parameter name, or calls a tool that doesn't exist—json.loads or DISPATCH[name] throws, the loop dies, and all prior work is lost.
  2. while True has no brakes. The model can fall into a "call tool → dislike the result → retry with tweaked arguments" loop, burning tokens until the end of time.

v1's fix: add a unified dispatcher that converts every exception into a string fed back to the model, plus a step cap:

python
MAX_STEPS = 12  # if a task isn't done within 12 steps, it has probably gone off the rails


def run_tool(name: str, raw_args: str) -> str:
    """Single entry point for tool execution: every exception becomes a string fed back to the model instead of a crash"""
    fn = DISPATCH.get(name)
    if fn is None:
        return f"error: unknown tool {name}; available tools: {list(DISPATCH)}"
    try:
        args = json.loads(raw_args)
    except json.JSONDecodeError:
        return f"error: arguments are not valid JSON: {raw_args[:200]}"
    try:
        return fn(**args)
    except TypeError as e:
        return f"error: argument mismatch: {e}"
    except Exception as e:
        return f"error: tool internal error {type(e).__name__}: {e}"


def run(task: str) -> str:
    messages = [
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": task},
    ]
    for step in range(MAX_STEPS):
        resp = client.chat.completions.create(model=MODEL, messages=messages, tools=TOOLS)
        msg = resp.choices[0].message
        messages.append(msg.model_dump(exclude_none=True))

        if not msg.tool_calls:
            return msg.content

        for tc in msg.tool_calls:
            result = run_tool(tc.function.name, tc.function.arguments)  # exceptions are absorbed here
            print(f"  [step {step}] {tc.function.name} -> {result[:60]}...")
            messages.append({"role": "tool", "tool_call_id": tc.id, "content": result})

    return f"(Hit the {MAX_STEPS}-step limit and stopped. Last message: {messages[-1]})"

Only two changes, but the behavioral difference is dramatic. Deliberately set a trap: delete notes.txt and ask again, and you'll see a trace like this:

  [step 0] read_file -> error: file not found: notes.txt...
  [step 1] run_shell -> notes.txt doesn't exist? Let me check with ls...
  [step 2] read_file -> error: path outside workspace: ../notes.txt...
  [step 3] (final answer) notes.txt doesn't currently exist...

That's the power of "error feedback": models can read error messages and change strategy on their own. error: file not found and it runs ls; error: path outside workspace and it switches back to relative paths. You don't have to teach it anything—you just have to write the errors clearly.

Error messages are written for the model, not for humans

error: unknown tool xxx; available tools: [read_file, run_shell] works far better than KeyError: 'xxx'—the former tells the model exactly what to do next. The first principle of tool return design: make it possible for the model to recover from errors. This is "errors are information" from Agent Design Principles put into practice.

The step cap is the other safety net: it's not a UX concern, it's a cost and safety concern. A runaway Agent burning tens of dollars in an hour is a real, documented incident. When MAX_STEPS runs out, v1 carries the last state out so you can see where it got stuck.

4. v2: Adding a todo Planning Tool ​

Once a task gets complex (say, "summarize all the TODO comments in the .py files under workspace into a report"), v1 shows classic symptoms: forgetting the goal halfway through, re-reading the same file, or submitting hastily after reading a single file. The cause is plain—the plan lives only in the model's "short-term memory," and every turn it has to re-infer what it's supposed to be doing from a sea of context.

The fix is an external "notepad" for the model: a todo tool. It's standard equipment in coding agents like Claude Code and Cursor, yet almost absurdly simple to implement—it doesn't actually do anything; it just writes the task list into the context:

python
TODO_STATE: list[dict] = []


def todo_write(todos: list) -> str:
    """Fully overwrite the task list. todos: [{"content": "...", "status": "pending|in_progress|completed"}]"""
    TODO_STATE.clear()
    TODO_STATE.extend(todos)
    mark = {"pending": " ", "in_progress": "~", "completed": "x"}
    lines = [f"[{mark.get(t.get('status'), ' ')}] {t.get('content', '')}" for t in TODO_STATE]
    return "Current task list:\n" + "\n".join(lines)


TODO_TOOL = {
    "type": "function",
    "function": {
        "name": "todo_write",
        "description": "Update the task list. List the steps before starting work, and mark each item completed as you finish it",
        "parameters": {
            "type": "object",
            "properties": {
                "todos": {
                    "type": "array",
                    "items": {
                        "type": "object",
                        "properties": {
                            "content": {"type": "string"},
                            "status": {"type": "string", "enum": ["pending", "in_progress", "completed"]},
                        },
                        "required": ["content", "status"],
                    },
                },
            },
            "required": ["todos"],
        },
    },
}

TOOLS.append(TODO_TOOL)
DISPATCH["todo_write"] = todo_write

And add one line to the system prompt:

- For tasks that take more than 3 steps, first lay out a plan with todo_write, and update the list's status after each completed step.

The expected trace becomes:

  [step 0] todo_write -> Current task list: [~] Find all .py files [ ] Extract TODO comments [ ] Compile the report
  [step 1] run_shell(find . -name "*.py") -> ./a.py ./b.py...
  [step 2] todo_write -> [x] Find all .py files [~] Extract TODO comments [ ] Compile the report...
  [step 3] run_shell(grep -rn TODO .) -> ./a.py:3:# TODO: handle empty files...
  [step 4] todo_write -> [x] [x] [x] ...
  [step 5] (final answer) found 2 files with 3 TODOs in total...

Why does a tool that "does nothing" improve performance so much? Two mechanisms:

  • Externalized goals. The plan shifts from "re-imagined by the model every turn" to "sitting in plain text in the context," so every turn the model can read "here's where I am right now." This is the minimal implementation of plan-and-execute from the Planning section.
  • Attention anchoring. Transformers attend more strongly to the beginning and end of the context; a repeatedly refreshed checklist keeps pushing "the current step" toward the recent end, counteracting the lost-in-the-middle effect in long contexts.

Rule of thumb: when is a todo tool worth it?

Add it when a task is expected to exceed 3-5 steps, or when intermediate artifacts need to be referenced across multiple steps; skip it for single-question-answer tasks—it only adds token overhead and one extra tool call. More tools isn't better: choosing between tools also consumes the model's "attention budget."

5. v3: Context Compression ​

Run long tasks and you'll hit this wall sooner or later: tool results pile into messages one by one, and after a few dozen turns the token count approaches the model's context window limit—you either get errors or a ridiculous bill. v3 adds the final mechanism: once a threshold is exceeded, automatically summarize the early conversation.

python
MAX_CONTEXT_TOKENS = 6000  # small threshold for the demo; in production set it to 70-80% of the model's real limit
KEEP_RECENT = 6            # keep the last N messages verbatim after compaction


def approx_tokens(messages: list) -> int:
    """Rough token estimate: ~4 characters ≈ 1 token. Use tiktoken if you need precision—this is enough for teaching"""
    return sum(len(json.dumps(m, ensure_ascii=False)) for m in messages) // 4


def compact(messages: list) -> list:
    """Compress the early conversation into a summary: keep system + summary + the last N messages verbatim"""
    boundary = len(messages) - KEEP_RECENT
    # Never cut between an assistant(tool_calls) message and its tool replies, or the protocol errors out.
    # Walk back until you land on a non-tool message.
    while boundary > 1 and messages[boundary].get("role") == "tool":
        boundary -= 1
    old, recent = messages[1:boundary], messages[boundary:]

    resp = client.chat.completions.create(
        model=MODEL,
        messages=[
            {"role": "system", "content": "You are a context compressor. Condense the conversation history into a briefing for the next Agent. "
                                          "You must preserve: the user's goal, completed steps, key file paths and conclusions, and outstanding todos."},
            {"role": "user", "content": "Please condense the following conversation:\n" + json.dumps(old, ensure_ascii=False)[:8000]},
        ],
    )
    summary = resp.choices[0].message.content
    print(f"  [compact] {approx_tokens(messages)} tokens -> compressed to summary")
    return [messages[0], {"role": "user", "content": f"[Summary of earlier context]\n{summary}"}] + recent

Wire one line into the top of each loop iteration:

python
    for step in range(MAX_STEPS):
        if approx_tokens(messages) > MAX_CONTEXT_TOKENS:
            messages = compact(messages)
        resp = client.chat.completions.create(model=MODEL, messages=messages, tools=TOOLS)
        # ...the rest is identical to v1

Note the while backtrack inside compact—this one comes from real battle scars: if a tool message is left at the start of recent while its paired assistant tool_calls message got compressed away, the API complains about an unpaired tool_call_id. The compaction point must fall on a message-pairing boundary.

Two honest caveats:

  • Compression is lossy. Claude Code users have actually run into this: after compaction, project instructions were "forgotten," and "next step" suggestions inside the summary were even executed automatically as if they were user instructions (see the issue in the references). That's why our implementation never compresses the system prompt; genuinely critical facts (conclusions, paths, decisions) should be written into files, not left only in conversation history—which is exactly the point of v2's todo and the memory file in the exercises below.
  • Don't set the threshold right at the limit. Leave 20-30% headroom for summary generation and subsequent output. Compacting at the wall creates a deadlock: "too long so compress, but too long to compress" (Claude Code stepped on this too).

The industrial version of this machinery is far more complex—layered summaries, selective retention, importance-based eviction—all territory of Context Engineering, but you've now implemented the core idea with your own hands.

6. Exercise List: What to Add Next ​

This ~150-line skeleton is still far from a "product-grade" Agent, but every extension is a real learning experience. In recommended order:

  1. Dangerous operation confirmation: give run_shell a write-command list that requires a human y/n before execution. Figuring out "which operations need human review" is the core question of Human-in-the-Loop.
  2. Memory file: have the Agent write important conclusions into memory.md and inject the file's content into the system prompt each turn. Compare the "compress conversation" and "external memory" routes—see Memory Systems.
  3. Sub-agents: add a spawn_agent(task) tool—internally just another call to run(), with only the final result brought back to the main conversation. Feel why "a sub-agent's context isolation protects the main conversation from pollution"; details in Multi-Agent Architectures.
  4. Retries and rate limiting: add exponential backoff around API calls (sleep and retry on 429/5xx). Without this, your Agent won't survive a day on a real network.
  5. Streaming output: wire up stream=True and print token by token. A fundamental change to the user experience, implemented in a dozen or so lines.
  6. Write evals: collect 10 tasks plus expected results, and run them after every prompt change to watch the pass rate. Prompt tuning without evals is astrology; the method is in Evals in Practice.
  7. Hook up MCP: swap one read_file for a call to a real MCP server. You'll find that an MCP client does exactly the work of your hand-written dispatch—the protocol is just standardized; see Tools & MCP.

Finish 1-3 and you'll have grasped 80% of the skeleton of products like Claude Code—the remaining 20% is polish, see Claude Code Case Study. Package the project with a README and publish it, and you have a legitimate portfolio project.

7. Common Bugs and Debugging Techniques ​

Sorted by how often you'll step on them:

SymptomRoot causeFix
400: tool_call_id did not have responseThe assistant's tool_calls never entered the history, or a tool reply is missing/misorderedCheck the msg.model_dump() line; mind the pairing boundary when compressing (see v3)
The model calls the same tool over and overThe tool errored but the message doesn't say what to do, or the result lacks the information it wantedImprove the error wording (v1); add an "after N failures, change approach" instruction to the system prompt
The model skips tools and just makes things upThe system prompt doesn't require verification first, or the tool descriptions read like decorationState explicitly "read the file before answering"; write descriptions like prompts
temperature parameter errorsGPT-5/o-series reasoning models don't support that parameterRemove it and use the default
Context overflowTool output isn't truncated—one cat dumps tens of thousands of charactersTruncate all tool output (4000 chars everywhere on this page); add v3 compression
ls/cat "command not found" on WindowsThe allowlisted commands are Unix-onlySwitch to a cross-platform implementation (pure-Python list_dir/read_file), or allow dir/type
JSON argument parsing failsThe model (especially weaker ones) produced truncated/invalid JSONv1's run_tool already catches it; if it happens often, switch to a stronger model or enable strict mode

Three debugging techniques cover most needs:

  1. Print the trace of every step (all the [tool] output on this page). The first job of Agent debugging is making the invisible loop visible—the primitive form of Observability.
  2. Dump the full messages to JSONL, one line per step. When something breaks, read the last few entries instead of guessing—it's a hundred times better than guessing, and you accumulate an eval dataset along the way.
  3. Mock the LLM to test the tool executor: a fake client returning fixed tool_calls lets you unit-test your dispatch, security boundaries, and error feedback. The tool layer is pure functions—don't spend API money testing it.

At this point you're holding a working Agent skeleton that can run, self-heal, plan, and won't blow the context. More importantly: from now on, when you read any framework's source code, you won't see magic—you'll see "oh, I've written this part."

References ​