Appearance
Build an Agent from Scratch
Reading ten agent tutorials is worth less than getting a "working" agent running with your own hands. This article takes the minimal viable path and builds, from scratch, a little assistant that can call tools to get things done — "check the weather and add it to my schedule" — with runnable code at every step, so you can see for yourself that the core of an agent is nothing more than a loop.
An agent is a software system that uses a large language model (LLM) as its decision-making hub, can call external tools to complete tasks, and can observe tool results to keep acting. It differs from "a plain LLM call" in exactly one way: the model doesn't just answer questions — it decides "what to do next" — look something up, run code, call an API, write a file — then looks at the result and decides what to do after that.
Why write it by hand? Because "agent" is the term most easily wrapped in marketing mystique, yet its core is remarkably plain: a while loop. Frameworks (LangGraph, OpenAI Agents SDK) merely industrialize that loop — adding state management, memory, parallelism, and observability. Once you have implemented the loop yourself, you will instantly understand what any framework is doing under the hood. For the conceptual big picture, see Agent Core Concepts and Manus and Agent Applications.
Here is the overall map first:
User request
│
▼
┌────────────────────────────┐
│ LLM decision layer (ReAct) │ ← loop: think → pick which tool to call
└────────────┬───────────────┘
│ call tool (function call)
▼
┌────────────────────────────┐
│ Tool layer (Tools) │ get_weather / add_event / ...
└────────────┬───────────────┘
│ feed the result back to the LLM
▼
┌────────────────────────────┐
│ Memory layer (Memory) │ conversation history / vector memory / files
└────────────┬───────────────┘
│ result satisfies the user's request
▼
Final answerOur implementation path has five steps: define the tools → the tool-calling loop → add memory → add planning → the complete demo.
Prerequisites
This article assumes you already know: what an agent is (Agent Core Concepts), the basic capability limits of LLMs (Large Language Models), and how to write instructions for a model (Prompt Engineering). If you want to first see how a ReAct loop reasons in theory, read it alongside Transformers and Attention to understand how "context" carries the entire loop's memory.
1. Define the tools: the function calling schema
To make an LLM "know how to use" tools, you give it a machine-readable tool manual (a schema) that spells out each tool's name, what it does, and which parameters it takes. The LLM never executes any code — it only outputs a structured "tool call request", which your program parses and executes.
① Tool implementation (the real logic)
python
import json
from pathlib import Path
def get_weather(city: str, date: str = "today") -> str:
"""Get the weather for a city. Mock data for demo purposes; swap in a weather API for real projects."""
import random
conditions = ["Sunny", "Cloudy", "Light rain", "Windy"]
return json.dumps({
"city": city, "date": date,
"weather": random.choice(conditions),
"temp": f"{random.randint(18, 34)}°C",
}, ensure_ascii=False)
def add_event(date: str, time: str, title: str) -> str:
"""Write an event to the calendar. Uses a local JSON file for the demo; wire up a calendar API for real projects."""
path = Path("events.json")
events = json.loads(path.read_text(encoding="utf-8")) if path.exists() else []
events.append({"date": date, "time": time, "title": title})
path.write_text(json.dumps(events, ensure_ascii=False, indent=2), encoding="utf-8")
return f"OK, added '{title}' to {date} {time}"② function calling schema (OpenAI style; other providers are similar)
python
TOOLS = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the weather for a city on a given day. Returns the condition and temperature.",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "City name, e.g. Beijing, Shanghai"},
"date": {"type": "string", "description": "Date (YYYY-MM-DD); defaults to today"},
},
"required": ["city"],
},
},
},
{
"type": "function",
"function": {
"name": "add_event",
"description": "Write an event to the calendar.",
"parameters": {
"type": "object",
"properties": {
"date": {"type": "string", "description": "Event date (YYYY-MM-DD)"},
"time": {"type": "string", "description": "Event time (HH:MM)"},
"title": {"type": "string", "description": "Event title"},
},
"required": ["date", "time", "title"],
},
},
},
]Tool descriptions make or break the agent
The most critical part of the schema is not the name but the description and each parameter's description — the LLM relies on descriptions to decide "when to use the tool and how to fill in the arguments." Vague descriptions (such as "handle weather") cause the model to skip the tool when it should use it, or to fill in arguments at random. Descriptions should state the trigger conditions, what each parameter means, and the valid value ranges. This is prompt engineering at heart; for the methodology see Prompt Engineering.
Between the tool implementations and the schema sits the dispatcher — it maps the name and JSON arguments requested by the model to the real functions:
python
def dispatch(name: str, args_json: str) -> str:
"""Dispatch the LLM's tool call request to the real implementation."""
args = json.loads(args_json)
if name == "get_weather":
return get_weather(args.get("city"), args.get("date", "today"))
if name == "add_event":
return add_event(args["date"], args["time"], args["title"])
return f"Error: unknown tool {name}"2. The LLM tool-calling loop: ReAct
Once the tools are defined, everything comes down to that loop. The classic pattern is called ReAct (Reasoning + Acting): the model first "reasons" about what to do next, then "acts" — emitting a tool call request; your program executes the tool and feeds the result back as an "observation"; the model keeps reasoning based on the observation. The cycle repeats until the model decides the task is complete and outputs a final answer directly. For how this concept relates to prompt engineering, see Prompt Engineering.
Loop structure (pseudocode):
while steps < max_steps:
reply = LLM(messages, tools) # 1. let the model decide the next move
if reply has no tool calls: return reply # 2a. done → output the final answer
for each tool call:
result = dispatch(tool_name, args) # 2b. execute the tool
messages.append(tool message) # 3. feed the observation back into the conversationFull implementation (the current OpenAI API):
python
from openai import OpenAI
client = OpenAI() # requires the OPENAI_API_KEY environment variable; for local Ollama see below
def run_agent(user_query: str, max_steps: int = 6, model: str = "gpt-4o-mini") -> str:
messages = [{"role": "user", "content": user_query}]
for step in range(max_steps):
resp = client.chat.completions.create(
model=model,
messages=messages,
tools=TOOLS, # hand the tool schemas to the model
)
msg = resp.choices[0].message
messages.append(msg) # the model's reply (including any tool call requests) joins the conversation
if not msg.tool_calls: # the model decided to answer directly → task complete
return msg.content
# Execute the tool calls one by one; feed results back as role="tool" messages
for call in msg.tool_calls:
result = dispatch(call.function.name, call.function.arguments)
print(f"[step {step}] called {call.function.name}({call.function.arguments}) -> {result}")
messages.append({
"role": "tool",
"tool_call_id": call.id,
"content": result,
})
return "Maximum step limit reached; task not completed."Run it once:
python
print(run_agent("Is today good for a run in Beijing? Check the weather and then give me advice."))A typical output trace (the lines printed by print) looks like this:
[step 0] called get_weather({"city":"Beijing","date":"today"}) -> {"city":"Beijing","date":"today","weather":"Sunny","temp":"27°C"}
It's sunny and 27°C in Beijing today — perfect for a run. Head out in the late afternoon, and remember to stay hydrated.Three easily overlooked details
- Don't drop the message history: the model's reply
msgmust be appended tomessagesverbatim (including itstool_callsfield); otherwise tool results no longer line up with their calls, and the model "loses its memory". - One reply may call multiple tools:
msg.tool_callsis a list — execute each call and append each result one by one. tool_call_idmust be passed back: thetool_call_idin each tool result message corresponds one-to-one with a call; this is the key to keeping multi-tool scenarios error-free.
Local Ollama version: just swap the client and the model name; everything else stays the same, giving you a fully offline agent:
python
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
model = "qwen2.5:7b" # must support function calling; run ollama pull qwen2.5 first3. Adding memory
Everything inside the loop flows through messages — that alone is within-session memory: multi-turn dialogue and each step's tool results all live there. But two kinds of memory need explicit design:
| Memory type | What it stores | Implementation | When it's useful |
|---|---|---|---|
| Within-session memory | Every message in the current task | Just the messages list | Multi-step tasks, multi-turn dialogue |
| Long-term memory | User preferences, facts, past conclusions | Files / databases / vector stores | Cross-session reuse, personalization |
| Working memory | Intermediate computation results | Variables, temporary files | Complex computation tasks |
| Vector memory | Semantically retrievable history | Vector store + similarity search | Recall on demand when history grows long |
Within-session memory needs no extra code (the loop above supports it naturally). The typical vector memory implementation: write the key conclusions produced during a task into a vector store, then retrieve them on demand instead of stuffing the entire history into the context (longer context costs more — see Inference Optimization and Quantization). For the underlying principles and implementation, see Vector Databases and Semantic Search:
python
import numpy as np
import faiss
from sentence_transformers import SentenceTransformer
mem_model = SentenceTransformer("BAAI/bge-small-zh-v1.5")
mem_index = faiss.IndexFlatIP(384) # semantic memory store (in-process)
memory_texts: list[str] = []
def remember(text: str) -> None:
v = mem_model.encode([text], normalize_embeddings=True)
mem_index.add(v.astype("float32"))
memory_texts.append(text)
def recall(query: str, k: int = 3) -> list[str]:
if not memory_texts:
return []
qv = mem_model.encode([query], normalize_embeddings=True)
_, ids = mem_index.search(qv.astype("float32"), k)
return [memory_texts[i] for i in ids[0] if i >= 0]
# Example: store "the user goes for a morning run every Tuesday" in memory so it comes up automatically next time
remember("User habit: goes for a morning run at 7 AM every Tuesday")
print(recall("What are this user's exercise habits?")) # ['User habit: goes for a morning run at 7 AM every Tuesday']Managing memory capacity
All memory eventually becomes tokens in the context. For long sessions you must truncate or compress (keep the most recent N turns plus a summary); keep k small for vector recall. An agent without capacity management will blow up its context after about 20 turns of dialogue — expensive and slow.
4. Adding planning: task decomposition
Complex tasks ("help me plan my weekend") often cannot be finished in a single tool call. Add a planner on top: first have the LLM split the big task into ordered subtasks, then execute them one by one and merge the results. This corresponds to an agent's planning capability:
python
def plan(task: str) -> list[str]:
"""Split the task into 2-4 ordered subtasks, one per line."""
resp = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system",
"content": "Break the user's task into 2-4 ordered subtasks, one per line. Output only the subtasks themselves."},
{"role": "user", "content": task},
],
)
return [ln.strip("- ").strip() for ln in resp.choices[0].message.content.splitlines() if ln.strip()]Then run the execution loop with planning in place: first plan, then call run_agent for each subtask, and finally summarize:
python
def run_with_plan(task: str) -> str:
steps = plan(task)
print("Plan:", steps)
results = [run_agent(f"Subtask: {s}") for s in steps]
# Summarize the subtask results and hand them to the LLM to produce the final reply
resp = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "system", "content": "Based on the subtask results below, give the user a complete, friendly final answer."},
{"role": "user", "content": "\n\n".join(f"### {s}\n{r}" for s, r in zip(steps, results))}],
)
return resp.choices[0].message.contentTwo flavors of planning
There are two ways to plan: explicit planning (split first, then execute, as above — controllable and auditable) and implicit planning (the model thinks and acts on the fly inside the ReAct loop — flexible but hard to predict). At the beginning, go explicit: split, execute, summarize — when something breaks, you can trace it to a specific subtask. For planning boundaries in complex tasks and a deeper discussion of "task decomposition", see Agent Core Concepts.
5. The complete runnable demo: check the weather and add it to the schedule
Put everything above together into agent_demo.py. Target task: "Check tomorrow's weather in Beijing and Shanghai, and schedule an outdoor run at 3 PM tomorrow in whichever city has better weather."
python
"""
agent_demo.py — a minimal runnable agent: weather lookup + calendar write
Usage: python agent_demo.py
Dependencies: pip install openai (for the local Ollama version, see the comment at the end of the code)
"""
import json
import random
from pathlib import Path
from openai import OpenAI
# ── Tool implementations ─────────────────────────────────────────────
def get_weather(city: str, date: str = "today") -> str:
conditions = ["Sunny", "Cloudy", "Light rain", "Windy"]
return json.dumps({"city": city, "date": date,
"weather": random.choice(conditions),
"temp": f"{random.randint(18, 34)}°C"},
ensure_ascii=False)
def add_event(date: str, time: str, title: str) -> str:
path = Path("events.json")
events = json.loads(path.read_text(encoding="utf-8")) if path.exists() else []
events.append({"date": date, "time": time, "title": title})
path.write_text(json.dumps(events, ensure_ascii=False, indent=2), encoding="utf-8")
return f"OK, added '{title}' to {date} {time}"
# ── function calling schema ──────────────────────────────────────────
TOOLS = [
{"type": "function", "function": {
"name": "get_weather",
"description": "Get the weather for a city on a given day. Returns the condition and temperature.",
"parameters": {"type": "object",
"properties": {"city": {"type": "string", "description": "City name"},
"date": {"type": "string", "description": "Date; defaults to today"}},
"required": ["city"]}}},
{"type": "function", "function": {
"name": "add_event",
"description": "Write an event to the calendar.",
"parameters": {"type": "object",
"properties": {"date": {"type": "string", "description": "Date YYYY-MM-DD"},
"time": {"type": "string", "description": "Time HH:MM"},
"title": {"type": "string", "description": "Title"}},
"required": ["date", "time", "title"]}}},
]
def dispatch(name: str, args_json: str) -> str:
args = json.loads(args_json)
if name == "get_weather":
return get_weather(args.get("city"), args.get("date", "today"))
if name == "add_event":
return add_event(args["date"], args["time"], args["title"])
return f"Error: unknown tool {name}"
# ── ReAct loop ───────────────────────────────────────────────────────
client = OpenAI() # requires OPENAI_API_KEY
# Local version: client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
# model = "qwen2.5:7b"
def run_agent(user_query: str, max_steps: int = 8, model: str = "gpt-4o-mini") -> str:
messages = [{"role": "user", "content": user_query}]
for step in range(max_steps):
resp = client.chat.completions.create(model=model, messages=messages, tools=TOOLS)
msg = resp.choices[0].message
messages.append(msg)
if not msg.tool_calls:
return msg.content
for call in msg.tool_calls:
result = dispatch(call.function.name, call.function.arguments)
print(f"[step {step}] {call.function.name}({call.function.arguments}) -> {result}")
messages.append({"role": "tool", "tool_call_id": call.id, "content": result})
return "Maximum step limit reached; task not completed."
if __name__ == "__main__":
q = "Check tomorrow's weather in Beijing and Shanghai, and schedule an outdoor run at 3 PM tomorrow in whichever city has better weather."
print(run_agent(q))
print("— Contents of events.json —")
print(Path("events.json").read_text(encoding="utf-8"))When you run it, you will see the model complete the entire chain on its own — look up two cities → compare → call add_event — and produce a well-reasoned final answer. This 80-line script is the complete skeleton of an agent — everything a framework does is engineering layered on top of this loop.
6. Choosing a framework: when not to roll your own
A loop you write yourself is right for learning and minimal scenarios; in real projects, tool count, state complexity, concurrency, and observability quickly outgrow a hand-written loop's comfort zone. A comparison of common frameworks:
| Framework | Core abstraction | Strengths | Costs | When to use |
|---|---|---|---|---|
| LangGraph | Graph + state machine | Controllable flow, supports branches/loops/parallelism, checkpointing | Many concepts, steep learning curve | Complex multi-step flows needing precise control and resumable runs |
| OpenAI Agents SDK | Agent + handoff | Officially maintained, function calling out of the box, little code | Tied to the OpenAI ecosystem | Fast delivery, mostly single-agent |
| Claude Agent SDK | Agent + tools + computer use | Official Anthropic, great tooling/sandbox experience | Tied to the Claude ecosystem | Deep Claude usage, browser/computer control needed |
| Roll your own (this article) | while loop | Zero dependencies, fully transparent, controllable | Lacks engineering features (state/retries/observability) | Learning, fewer than 5 tools, fixed flow |
The decision rule in one sentence: fewer than 5 tools and a fixed flow → rolling your own is enough; building a product that needs parallelism and auditability → use a framework. Whichever framework you pick, the mental model of "tool schema + loop + memory" is universal — frameworks are just different implementations of it. For model selection and cost reference, see Model and Leaderboard Quick Reference.
7. Engineering essentials: from demo to production
Once the hand-written loop works, these problems surface immediately in real-world scenarios:
1. Truncating tool returns
Tools may return very long text (database queries, file reads). Truncate tool returns before feeding them back, otherwise the context balloons and token costs spike:
python
def safe_tool_result(result: str, max_len: int = 4000) -> str:
if len(result) > max_len:
return result[:max_len] + f"\n... (truncated; the original text was {len(result)} characters)"
return result2. Error retries and fault tolerance
Tools may throw exceptions (API timeouts, invalid arguments). Convert exceptions into text and feed them back to the model so it can correct itself, instead of letting the whole loop crash:
python
try:
result = dispatch(call.function.name, call.function.arguments)
except Exception as e: # for more rigor, use isinstance to distinguish exception types
result = f"Tool execution failed: {type(e).__name__}: {e}"
messages.append({"role": "tool", "tool_call_id": call.id, "content": safe_tool_result(result)})3. Max steps and infinite-loop detection
max_steps is mandatory, and you must detect repeated actions (same tool + same arguments N times in a row) and declare failure immediately. Infinite loops are the most common "incident" pattern for agents.
4. Cost control
Every step of the loop is one LLM call. Ways to keep it under control: use a cheap small model for planning/simple tool calls, add semantic caching (similar questions hit the cache), and compress message history. For these optimizations, see Inference Optimization and Quantization and LLM Deployment and Inference Optimization in Practice.
5. Permissions and security
Tools are the only channel through which an agent affects the outside world — and the most dangerous part:
| Risk | Scenario | Countermeasure |
|---|---|---|
| Prompt injection | Instructions hidden in tool-returned content (e.g. scraped web pages) lure the agent into malicious actions | Treat tool results as "data", not "instructions"; require confirmation for sensitive operations; isolate context |
| Unauthorized tool calls | The model calls a tool it shouldn't (e.g. a delete operation) | Least privilege: authorize each tool separately, require user confirmation for sensitive operations |
| Sandboxed execution | Letting the agent run generated code | Run inside a sandbox/container; forbid access to sensitive paths |
| Parameter validation | Tool arguments come from the model and may be invalid | Validate types and value ranges at the dispatch layer |
For security boundaries and governance frameworks, see AI Safety and Governance.
Two red lines
- Never hand irreversible tools like "delete" or "transfer money" to the agent for automatic invocation — always add a human confirmation step.
- External text mixed into tool output (web pages, emails, documents) is untrusted by default — it may carry attack instructions. Prompt injection is the most realistic attack surface against agents.
8. Evaluating agents and common pitfalls
Evaluation
Agents have to be evaluated on both process and result. The four most practical metrics:
| Metric | What it measures | How to test |
|---|---|---|
| Task success rate | Did the task complete end to end | A batch of real tasks, judged success/failure by humans |
| Tool call correctness | Was the right tool called with the right arguments | Compare the "expected tool call sequence" against the actual one |
| Step efficiency | How many steps it took | Track average steps; outliers usually mean "taking detours" |
| Cost | Tokens per task | Log token usage, correlated with task type |
Recommendation: build an evaluation set of 20-50 tasks first, then change the loop logic, and run the eval after every change — don't rely on "it feels better". For building the pipeline, see Building an LLM Evaluation Pipeline; for the theory behind the metrics, see LLM Evaluation and Benchmarks.
Common pitfalls
| Pitfall | Symptom | Root cause | Fix |
|---|---|---|---|
| Runaway loops | Endless tool calls, tokens burning | No step limit, model retrying endlessly | max_steps + repeated-action detection + timeout |
| Hallucinated tool arguments | Calls with nonexistent cities / wrong argument types | Unclear schema descriptions, model improvising | Strengthen argument descriptions, validate at the dispatch layer, retry when required arguments are missing |
| Context bloat | Increasingly slow and expensive, answers miss the point | Full history stacking up | Truncate/compress/vector memory with on-demand recall |
| Tool result "contamination" | Behavior suddenly goes wrong | Instruction-like text mixed into tool returns | Truncate results, label them as data, isolate context |
| Over-calling tools | Even a simple "hello" triggers a tool call | Schema trigger conditions written too broadly | State clearly in the description "when NOT to call" |
| Silent failure | Task "looks done" but wasn't actually executed | The tool returned an error that was treated as success | Have tools return structured status (ok/error) and check it explicitly in the loop |
For a more complete list of agent antipatterns, see Common Pitfalls and Antipatterns.
9. Advanced: multi-agent collaboration
A single agent has limits on capacity and responsibility; for complex tasks, let multiple agents with distinct roles collaborate. Three mainstream orchestration patterns:
| Pattern | Structure | Best for | Example |
|---|---|---|---|
| Orchestrator-worker | A lead agent splits the task and assigns it to specialized sub-agents | Decomposable tasks with fixed subtask types | Writing a report = retrieval agent + writing agent + formatting agent |
| Pipeline | Each agent's output is the next one's input | Tasks with a natural order | Research → analysis → summary |
| Debate/review | Multiple agents answer independently, then critique each other | High correctness requirements | Code review, fact-checking |
The engineering practices that go with it: sub-agents interact only through "messages/results" (no shared mutable state), each sub-agent gets its own isolated context, and the lead agent handles aggregation and conflict resolution. For a complete teardown of a multi-agent product, see Manus and Agent Applications; what determines the collaboration structure ultimately traces back to the discussion of "task complexity and agent shape" in Agent Core Concepts.
Closing checklist
A "qualified" agent: completes tasks reliably, has step-count and cost guardrails, gets human confirmation for irreversible tool operations, and has an evaluation set that quantifies "nothing got worse". With those four boxes ticked, it is a system you can have an engineering conversation about — not a demo. Look up unfamiliar terms in the Glossary; for the big picture, see What Is AI: Key Concepts.
Further Reading
- Agent Core Concepts — the full theoretical version of this article: definitions, taxonomy, ReAct, memory, planning, and multi-agent
- Manus and Agent Applications — a complete teardown of a production-grade agent product
- Prompt Engineering — methodology for writing tool descriptions, system prompts, and few-shot examples
- Vector Databases and Semantic Search — how vector memory works and how to choose one
- Building a RAG App from Scratch — turning "retrieval" into an agent tool, i.e. Agentic RAG
- LLM Evaluation and Benchmarks — the theoretical basis for agent evaluation metrics
- Building an LLM Evaluation Pipeline — turning Section 8 into a runnable evaluation pipeline
- Inference Optimization and Quantization — the principles behind loop cost control (caching, quantization, streaming)
- AI Safety and Governance — a complete framework for prompt injection, permissions, and alignment
- Common Pitfalls and Antipatterns — the full 25-item version of Section 8
- Model and Leaderboard Quick Reference — model selection data on tool-calling strength
- Anatomy of the Overall Architecture — scaling this article's little agent up to a production system
References
- Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (arXiv 2210.03629) — the original ReAct paper, the source of the agent loop
- OpenAI docs: Function calling — official tutorial and best practices for the tool calling interface
- OpenAI Agents SDK — the official lightweight agent framework
- LangGraph official documentation — graph state-machine agent framework
- Anthropic: Claude Agent SDK and Tool use — guide to agents and tool use in the Claude ecosystem
- Anthropic engineering blog: Building effective agents — engineering judgment on when to (and when not to) use agents
- Ollama website — installation and models for local LLMs (including tool calling)