Appearance
Build a Minimal Agent from Scratch
The biggest trap when learning Agents: you pick up a pile of frameworks, but when someone asks "what actually is an Agent?", you can't answer. This page takes the most direct route—no LangChain, no framework of any kind, just the LLM API and a while loop, building a working Agent from zero. By the end you'll understand that what an Agent framework packages up is essentially the few hundred lines you're about to write, plus a pile of engineering details.
The page iterates through four versions, each solving one real pain point:
v0 Minimal Agent Loop → runs, but can spin forever and can crash
v1 Tool error feedback + step cap → self-healing, can't run away
v2 todo planning tool → long tasks stay on track
v3 Context compression → long sessions don't blow the context windowYou only need three things: Python 3.10+, pip install openai (1.x), and an API key for a model that supports tool calling. The code uses the OpenAI Chat Completions format, the de facto standard—it runs unchanged against OpenAI's official API as well as OpenAI-compatible endpoints like DeepSeek, Qwen, and Kimi; just swap the base_url.
1. First, Get Clear on the Minimal Core of an Agent
Strip away all the framework packaging and an Agent's core is a loop:
┌──────────────────────────────────────────┐
│ │
▼ │
┌───────────┐ tool_calls ┌────────────┐ │
│ LLM │───────────────▶│ Run tools │ │
│ (with │ │ (your code)│ │
│ context) │◀───────────────┴────────────┘ │
└───────────┘ feed tool results back into │
│ messages │
└── no more tool calls = final answer → exit the loop ──┘That is the entire essence of the Agent Loop. Each iteration does only four things:
- Send the model the "conversation history + tool schemas";
- The model either returns final text or one or more
tool_calls(function name + JSON arguments); - Your code executes the requested functions and appends the results to the history as
role: "tool"messages; - Back to step 1, until the model stops calling tools.
A key insight
LLMs never "execute" tools. They only generate structured text saying "I want to call read_file with arguments {"path": "notes.txt"}". It's your Python code that actually runs it—which means a tool's permission boundary, timeouts, and error handling are all your responsibility, not the model's. This is exactly why Tools & MCP and Agent Security keep hammering on boundary control.
2. v0: A 50-line Agent Loop
Here is the complete, runnable v0. It has two real tools: read_file (read a file) and run_shell (run whitelisted read-only commands). Note the two security decisions—these are baseline requirements, not optional polish:
read_filehard-restricts paths toworkspace/to prevent directory traversal (a model-supplied../../etc/passwdgets rejected);run_shelluses a command allowlist plusshell=False(arguments split viashlex.splitand passed as an array), ruling out injection and write operations.
python
# mini_agent_v0.py — a minimal Agent that actually runs
# Dependencies: pip install openai; environment variable OPENAI_API_KEY
import json
import shlex
import subprocess
from pathlib import Path
from openai import OpenAI
client = OpenAI() # compatible endpoint example: OpenAI(base_url="https://api.deepseek.com", api_key="...")
MODEL = "gpt-5-mini" # any model that supports tool calling will do
# Tools may only operate inside this directory, so the model can't wander the disk
WORKSPACE = Path("./workspace").resolve()
SYSTEM_PROMPT = """You are a file-operations assistant.
- Use read_file to view file contents and run_shell to run read-only commands (ls/cat/grep/find/wc, etc.).
- All paths are relative to the working directory workspace/.
- Gather enough information before answering; once you have enough, give your final answer directly instead of calling more tools.
"""
# Tool schemas: the model relies on descriptions to decide when to call and how to fill in arguments—clear writing is productivity
TOOLS = [
{
"type": "function",
"function": {
"name": "read_file",
"description": "Read the contents of a text file inside the working directory",
"parameters": {
"type": "object",
"properties": {
"path": {"type": "string", "description": "File path relative to workspace"},
},
"required": ["path"],
},
},
},
{
"type": "function",
"function": {
"name": "run_shell",
"description": "Run a read-only shell command (allowlist: ls/cat/grep/find/pwd/wc/head/tail/echo)",
"parameters": {
"type": "object",
"properties": {
"command": {"type": "string", "description": "The command to run"},
},
"required": ["command"],
},
},
},
]
ALLOWED_CMDS = {"ls", "cat", "grep", "find", "pwd", "wc", "head", "tail", "echo"}
def safe_path(p: str) -> Path:
"""Confine model-supplied paths to WORKSPACE to prevent directory traversal"""
full = (WORKSPACE / p).resolve()
if not str(full).startswith(str(WORKSPACE)):
raise ValueError(f"path outside workspace: {p}")
return full
def read_file(path: str) -> str:
p = safe_path(path)
if not p.is_file():
return f"error: file not found: {path}"
return p.read_text(encoding="utf-8", errors="replace")[:4000] # truncate to protect the context
def run_shell(command: str) -> str:
argv = shlex.split(command)
if not argv or argv[0] not in ALLOWED_CMDS:
return f"error: command rejected (not in allowlist): {command}"
try:
out = subprocess.run(argv, cwd=WORKSPACE, capture_output=True, text=True, timeout=10)
return (out.stdout + out.stderr)[:4000] or "(no output)"
except subprocess.TimeoutExpired:
return "error: command timed out (10s)"
except FileNotFoundError:
return f"error: command not found: {argv[0]}"
DISPATCH = {"read_file": read_file, "run_shell": run_shell}
def run(task: str) -> str:
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": task},
]
while True: # v0 deliberately adds no guardrails—watch what happens first
resp = client.chat.completions.create(model=MODEL, messages=messages, tools=TOOLS)
msg = resp.choices[0].message
# The model's raw reply must go into the history; tool messages pair with it via tool_call_id
messages.append(msg.model_dump(exclude_none=True))
if not msg.tool_calls: # stop condition: no more tool calls = final answer
return msg.content
for tc in msg.tool_calls: # a single turn may call several tools in parallel
args = json.loads(tc.function.arguments)
result = DISPATCH[tc.function.name](**args)
print(f" [tool] {tc.function.name}({args}) -> {result[:60]}...")
messages.append({"role": "tool", "tool_call_id": tc.id, "content": result})
if __name__ == "__main__":
WORKSPACE.mkdir(exist_ok=True)
(WORKSPACE / "notes.txt").write_text("Goal this month: get the mini agent running\nBudget: 3,000 yuan\n", encoding="utf-8")
print(run("What budget is written in the notes.txt in workspace?"))A typical run (details vary each time, but the skeleton is the same):
$ python mini_agent_v0.py
[tool] run_shell({'command': 'ls'}) -> notes.txt...
[tool] read_file({'path': 'notes.txt'}) -> Goal this month: get the mini agent running...
The budget written in notes.txt is 3,000 yuan.A few points worth pausing on:
msg.model_dump(exclude_none=True)can't be skipped or casually changed. The Chat Completions protocol requires the assistant'stool_callsmessage and the subsequentrole: "tool"messages to pair up one-to-one viatool_call_id; miss one or get the order wrong and you get a 400. This is the most common beginner error by far.- The stop condition is simply "the model stops calling tools." You don't need an explicit
finishtool—the loop ends naturally when the model outputs plain text. This is also why the system prompt says "once you have enough information, answer directly." - Don't pass
temperatureto reasoning models like the GPT-5 series—the API errors out on it. Omit the parameter and let the model use its default. - Cost: mid-2026 pricing for mini/nano-tier models is on the order of a few dozen cents per million input tokens; running every experiment on this page typically costs less than a cent. Run freely.
Why Chat Completions instead of the newer Responses API
OpenAI's docs now push the Responses API (flattened tool schemas, built-in tool search, and other new features), but Chat Completions is the lowest common denominator across all OpenAI-compatible endpoints—learn this message structure once and you can switch vendors without changing code. For new production projects, evaluate the Responses API.
3. v1: Tool Error Feedback and a Step Cap
v0 handles simple tasks fine; harder tasks expose two fatal flaws:
- Any exception crashes the whole program. The model emits malformed JSON arguments, passes a nonexistent parameter name, or calls a tool that doesn't exist—
json.loadsorDISPATCH[name]throws, the loop dies, and all prior work is lost. while Truehas no brakes. The model can fall into a "call tool → dislike the result → retry with tweaked arguments" loop, burning tokens until the end of time.
v1's fix: add a unified dispatcher that converts every exception into a string fed back to the model, plus a step cap:
python
MAX_STEPS = 12 # if a task isn't done within 12 steps, it has probably gone off the rails
def run_tool(name: str, raw_args: str) -> str:
"""Single entry point for tool execution: every exception becomes a string fed back to the model instead of a crash"""
fn = DISPATCH.get(name)
if fn is None:
return f"error: unknown tool {name}; available tools: {list(DISPATCH)}"
try:
args = json.loads(raw_args)
except json.JSONDecodeError:
return f"error: arguments are not valid JSON: {raw_args[:200]}"
try:
return fn(**args)
except TypeError as e:
return f"error: argument mismatch: {e}"
except Exception as e:
return f"error: tool internal error {type(e).__name__}: {e}"
def run(task: str) -> str:
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": task},
]
for step in range(MAX_STEPS):
resp = client.chat.completions.create(model=MODEL, messages=messages, tools=TOOLS)
msg = resp.choices[0].message
messages.append(msg.model_dump(exclude_none=True))
if not msg.tool_calls:
return msg.content
for tc in msg.tool_calls:
result = run_tool(tc.function.name, tc.function.arguments) # exceptions are absorbed here
print(f" [step {step}] {tc.function.name} -> {result[:60]}...")
messages.append({"role": "tool", "tool_call_id": tc.id, "content": result})
return f"(Hit the {MAX_STEPS}-step limit and stopped. Last message: {messages[-1]})"Only two changes, but the behavioral difference is dramatic. Deliberately set a trap: delete notes.txt and ask again, and you'll see a trace like this:
[step 0] read_file -> error: file not found: notes.txt...
[step 1] run_shell -> notes.txt doesn't exist? Let me check with ls...
[step 2] read_file -> error: path outside workspace: ../notes.txt...
[step 3] (final answer) notes.txt doesn't currently exist...That's the power of "error feedback": models can read error messages and change strategy on their own. error: file not found and it runs ls; error: path outside workspace and it switches back to relative paths. You don't have to teach it anything—you just have to write the errors clearly.
Error messages are written for the model, not for humans
error: unknown tool xxx; available tools: [read_file, run_shell] works far better than KeyError: 'xxx'—the former tells the model exactly what to do next. The first principle of tool return design: make it possible for the model to recover from errors. This is "errors are information" from Agent Design Principles put into practice.
The step cap is the other safety net: it's not a UX concern, it's a cost and safety concern. A runaway Agent burning tens of dollars in an hour is a real, documented incident. When MAX_STEPS runs out, v1 carries the last state out so you can see where it got stuck.
4. v2: Adding a todo Planning Tool
Once a task gets complex (say, "summarize all the TODO comments in the .py files under workspace into a report"), v1 shows classic symptoms: forgetting the goal halfway through, re-reading the same file, or submitting hastily after reading a single file. The cause is plain—the plan lives only in the model's "short-term memory," and every turn it has to re-infer what it's supposed to be doing from a sea of context.
The fix is an external "notepad" for the model: a todo tool. It's standard equipment in coding agents like Claude Code and Cursor, yet almost absurdly simple to implement—it doesn't actually do anything; it just writes the task list into the context:
python
TODO_STATE: list[dict] = []
def todo_write(todos: list) -> str:
"""Fully overwrite the task list. todos: [{"content": "...", "status": "pending|in_progress|completed"}]"""
TODO_STATE.clear()
TODO_STATE.extend(todos)
mark = {"pending": " ", "in_progress": "~", "completed": "x"}
lines = [f"[{mark.get(t.get('status'), ' ')}] {t.get('content', '')}" for t in TODO_STATE]
return "Current task list:\n" + "\n".join(lines)
TODO_TOOL = {
"type": "function",
"function": {
"name": "todo_write",
"description": "Update the task list. List the steps before starting work, and mark each item completed as you finish it",
"parameters": {
"type": "object",
"properties": {
"todos": {
"type": "array",
"items": {
"type": "object",
"properties": {
"content": {"type": "string"},
"status": {"type": "string", "enum": ["pending", "in_progress", "completed"]},
},
"required": ["content", "status"],
},
},
},
"required": ["todos"],
},
},
}
TOOLS.append(TODO_TOOL)
DISPATCH["todo_write"] = todo_writeAnd add one line to the system prompt:
- For tasks that take more than 3 steps, first lay out a plan with todo_write, and update the list's status after each completed step.The expected trace becomes:
[step 0] todo_write -> Current task list: [~] Find all .py files [ ] Extract TODO comments [ ] Compile the report
[step 1] run_shell(find . -name "*.py") -> ./a.py ./b.py...
[step 2] todo_write -> [x] Find all .py files [~] Extract TODO comments [ ] Compile the report...
[step 3] run_shell(grep -rn TODO .) -> ./a.py:3:# TODO: handle empty files...
[step 4] todo_write -> [x] [x] [x] ...
[step 5] (final answer) found 2 files with 3 TODOs in total...Why does a tool that "does nothing" improve performance so much? Two mechanisms:
- Externalized goals. The plan shifts from "re-imagined by the model every turn" to "sitting in plain text in the context," so every turn the model can read "here's where I am right now." This is the minimal implementation of plan-and-execute from the Planning section.
- Attention anchoring. Transformers attend more strongly to the beginning and end of the context; a repeatedly refreshed checklist keeps pushing "the current step" toward the recent end, counteracting the lost-in-the-middle effect in long contexts.
Rule of thumb: when is a todo tool worth it?
Add it when a task is expected to exceed 3-5 steps, or when intermediate artifacts need to be referenced across multiple steps; skip it for single-question-answer tasks—it only adds token overhead and one extra tool call. More tools isn't better: choosing between tools also consumes the model's "attention budget."
5. v3: Context Compression
Run long tasks and you'll hit this wall sooner or later: tool results pile into messages one by one, and after a few dozen turns the token count approaches the model's context window limit—you either get errors or a ridiculous bill. v3 adds the final mechanism: once a threshold is exceeded, automatically summarize the early conversation.
python
MAX_CONTEXT_TOKENS = 6000 # small threshold for the demo; in production set it to 70-80% of the model's real limit
KEEP_RECENT = 6 # keep the last N messages verbatim after compaction
def approx_tokens(messages: list) -> int:
"""Rough token estimate: ~4 characters ≈ 1 token. Use tiktoken if you need precision—this is enough for teaching"""
return sum(len(json.dumps(m, ensure_ascii=False)) for m in messages) // 4
def compact(messages: list) -> list:
"""Compress the early conversation into a summary: keep system + summary + the last N messages verbatim"""
boundary = len(messages) - KEEP_RECENT
# Never cut between an assistant(tool_calls) message and its tool replies, or the protocol errors out.
# Walk back until you land on a non-tool message.
while boundary > 1 and messages[boundary].get("role") == "tool":
boundary -= 1
old, recent = messages[1:boundary], messages[boundary:]
resp = client.chat.completions.create(
model=MODEL,
messages=[
{"role": "system", "content": "You are a context compressor. Condense the conversation history into a briefing for the next Agent. "
"You must preserve: the user's goal, completed steps, key file paths and conclusions, and outstanding todos."},
{"role": "user", "content": "Please condense the following conversation:\n" + json.dumps(old, ensure_ascii=False)[:8000]},
],
)
summary = resp.choices[0].message.content
print(f" [compact] {approx_tokens(messages)} tokens -> compressed to summary")
return [messages[0], {"role": "user", "content": f"[Summary of earlier context]\n{summary}"}] + recentWire one line into the top of each loop iteration:
python
for step in range(MAX_STEPS):
if approx_tokens(messages) > MAX_CONTEXT_TOKENS:
messages = compact(messages)
resp = client.chat.completions.create(model=MODEL, messages=messages, tools=TOOLS)
# ...the rest is identical to v1Note the while backtrack inside compact—this one comes from real battle scars: if a tool message is left at the start of recent while its paired assistant tool_calls message got compressed away, the API complains about an unpaired tool_call_id. The compaction point must fall on a message-pairing boundary.
Two honest caveats:
- Compression is lossy. Claude Code users have actually run into this: after compaction, project instructions were "forgotten," and "next step" suggestions inside the summary were even executed automatically as if they were user instructions (see the issue in the references). That's why our implementation never compresses the system prompt; genuinely critical facts (conclusions, paths, decisions) should be written into files, not left only in conversation history—which is exactly the point of v2's todo and the memory file in the exercises below.
- Don't set the threshold right at the limit. Leave 20-30% headroom for summary generation and subsequent output. Compacting at the wall creates a deadlock: "too long so compress, but too long to compress" (Claude Code stepped on this too).
The industrial version of this machinery is far more complex—layered summaries, selective retention, importance-based eviction—all territory of Context Engineering, but you've now implemented the core idea with your own hands.
6. Exercise List: What to Add Next
This ~150-line skeleton is still far from a "product-grade" Agent, but every extension is a real learning experience. In recommended order:
- Dangerous operation confirmation: give
run_shella write-command list that requires a humany/nbefore execution. Figuring out "which operations need human review" is the core question of Human-in-the-Loop. - Memory file: have the Agent write important conclusions into
memory.mdand inject the file's content into the system prompt each turn. Compare the "compress conversation" and "external memory" routes—see Memory Systems. - Sub-agents: add a
spawn_agent(task)tool—internally just another call torun(), with only the final result brought back to the main conversation. Feel why "a sub-agent's context isolation protects the main conversation from pollution"; details in Multi-Agent Architectures. - Retries and rate limiting: add exponential backoff around API calls (sleep and retry on 429/5xx). Without this, your Agent won't survive a day on a real network.
- Streaming output: wire up
stream=Trueand print token by token. A fundamental change to the user experience, implemented in a dozen or so lines. - Write evals: collect 10 tasks plus expected results, and run them after every prompt change to watch the pass rate. Prompt tuning without evals is astrology; the method is in Evals in Practice.
- Hook up MCP: swap one
read_filefor a call to a real MCP server. You'll find that an MCP client does exactly the work of your hand-written dispatch—the protocol is just standardized; see Tools & MCP.
Finish 1-3 and you'll have grasped 80% of the skeleton of products like Claude Code—the remaining 20% is polish, see Claude Code Case Study. Package the project with a README and publish it, and you have a legitimate portfolio project.
7. Common Bugs and Debugging Techniques
Sorted by how often you'll step on them:
| Symptom | Root cause | Fix |
|---|---|---|
400: tool_call_id did not have response | The assistant's tool_calls never entered the history, or a tool reply is missing/misordered | Check the msg.model_dump() line; mind the pairing boundary when compressing (see v3) |
| The model calls the same tool over and over | The tool errored but the message doesn't say what to do, or the result lacks the information it wanted | Improve the error wording (v1); add an "after N failures, change approach" instruction to the system prompt |
| The model skips tools and just makes things up | The system prompt doesn't require verification first, or the tool descriptions read like decoration | State explicitly "read the file before answering"; write descriptions like prompts |
temperature parameter errors | GPT-5/o-series reasoning models don't support that parameter | Remove it and use the default |
| Context overflow | Tool output isn't truncated—one cat dumps tens of thousands of characters | Truncate all tool output (4000 chars everywhere on this page); add v3 compression |
ls/cat "command not found" on Windows | The allowlisted commands are Unix-only | Switch to a cross-platform implementation (pure-Python list_dir/read_file), or allow dir/type |
| JSON argument parsing fails | The model (especially weaker ones) produced truncated/invalid JSON | v1's run_tool already catches it; if it happens often, switch to a stronger model or enable strict mode |
Three debugging techniques cover most needs:
- Print the trace of every step (all the
[tool]output on this page). The first job of Agent debugging is making the invisible loop visible—the primitive form of Observability. - Dump the full
messagesto JSONL, one line per step. When something breaks, read the last few entries instead of guessing—it's a hundred times better than guessing, and you accumulate an eval dataset along the way. - Mock the LLM to test the tool executor: a fake
clientreturning fixedtool_callslets you unit-test your dispatch, security boundaries, and error feedback. The tool layer is pure functions—don't spend API money testing it.
At this point you're holding a working Agent skeleton that can run, self-heal, plan, and won't blow the context. More importantly: from now on, when you read any framework's source code, you won't see magic—you'll see "oh, I've written this part."
References
- OpenAI Function Calling official guide — the authoritative reference for tool schemas, the tool_calls protocol, strict mode, and the Responses API.
- How to Build an AI Agent from Scratch in Python (2026) — an English tutorial with the same approach as this page, about 60 lines of core loop.
- sergenes/mini_agent — a minimal open-source Agent using only the OpenAI SDK and a while loop, with a companion Medium series.
- codereindeer-dev/minimal-agent — a single-file Python Agent that adds one concept per commit (tools, memory, sub-agents); a good cross-reference for the exercise list.
- TheSeydiCharyyev/build-your-own-agent — a curated index of "build-your-own-x" style Agent component tutorials.
- Claude Code issue: CLAUDE.md instructions lost after compaction — a real case of context compression's lossiness, backing up v3's warning.
- Alibaba Cloud: calling Qwen via the OpenAI-compatible mode — reference for switching this page's code to a domestic model endpoint.