Skip to content

Build an Agent from Scratch

At a glance Build, from scratch, an agent that can call tools to complete tasks — function definitions and function calling schemas, the ReAct tool loop, memory, and planning — all threaded through a "check the weather and add it to the schedule" example, with framework selection, engineering essentials, evaluation, common pitfalls, and directly runnable code.

This page contains time-sensitive material, accurate as of 2025-06; job listings, leaderboards, and product features may have changed since. Verify against the original source before citing.

Build an Agent from Scratch ​

Reading ten agent tutorials is worth less than getting a "working" agent running with your own hands. This article takes the minimal viable path and builds, from scratch, a little assistant that can call tools to get things done — "check the weather and add it to my schedule" — with runnable code at every step, so you can see for yourself that the core of an agent is nothing more than a loop.

An agent is a software system that uses a large language model (LLM) as its decision-making hub, can call external tools to complete tasks, and can observe tool results to keep acting. It differs from "a plain LLM call" in exactly one way: the model doesn't just answer questions — it decides "what to do next" — look something up, run code, call an API, write a file — then looks at the result and decides what to do after that.

Why write it by hand? Because "agent" is the term most easily wrapped in marketing mystique, yet its core is remarkably plain: a while loop. Frameworks (LangGraph, OpenAI Agents SDK) merely industrialize that loop — adding state management, memory, parallelism, and observability. Once you have implemented the loop yourself, you will instantly understand what any framework is doing under the hood. For the conceptual big picture, see Agent Core Concepts and Manus and Agent Applications.

Here is the overall map first:

User request
   │
   ▼
┌────────────────────────────┐
│ LLM decision layer (ReAct) │ ← loop: think → pick which tool to call
└────────────┬───────────────┘
             │ call tool (function call)
             ▼
┌────────────────────────────┐
│ Tool layer (Tools)         │  get_weather / add_event / ...
└────────────┬───────────────┘
             │ feed the result back to the LLM
             ▼
┌────────────────────────────┐
│ Memory layer (Memory)      │  conversation history / vector memory / files
└────────────┬───────────────┘
             │ result satisfies the user's request
             ▼
       Final answer

Our implementation path has five steps: define the tools → the tool-calling loop → add memory → add planning → the complete demo.

Prerequisites

This article assumes you already know: what an agent is (Agent Core Concepts), the basic capability limits of LLMs (Large Language Models), and how to write instructions for a model (Prompt Engineering). If you want to first see how a ReAct loop reasons in theory, read it alongside Transformers and Attention to understand how "context" carries the entire loop's memory.

1. Define the tools: the function calling schema ​

To make an LLM "know how to use" tools, you give it a machine-readable tool manual (a schema) that spells out each tool's name, what it does, and which parameters it takes. The LLM never executes any code — it only outputs a structured "tool call request", which your program parses and executes.

① Tool implementation (the real logic)

python
import json
from pathlib import Path

def get_weather(city: str, date: str = "today") -> str:
    """Get the weather for a city. Mock data for demo purposes; swap in a weather API for real projects."""
    import random
    conditions = ["Sunny", "Cloudy", "Light rain", "Windy"]
    return json.dumps({
        "city": city, "date": date,
        "weather": random.choice(conditions),
        "temp": f"{random.randint(18, 34)}°C",
    }, ensure_ascii=False)

def add_event(date: str, time: str, title: str) -> str:
    """Write an event to the calendar. Uses a local JSON file for the demo; wire up a calendar API for real projects."""
    path = Path("events.json")
    events = json.loads(path.read_text(encoding="utf-8")) if path.exists() else []
    events.append({"date": date, "time": time, "title": title})
    path.write_text(json.dumps(events, ensure_ascii=False, indent=2), encoding="utf-8")
    return f"OK, added '{title}' to {date} {time}"

② function calling schema (OpenAI style; other providers are similar)

python
TOOLS = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get the weather for a city on a given day. Returns the condition and temperature.",
            "parameters": {
                "type": "object",
                "properties": {
                    "city": {"type": "string", "description": "City name, e.g. Beijing, Shanghai"},
                    "date": {"type": "string", "description": "Date (YYYY-MM-DD); defaults to today"},
                },
                "required": ["city"],
            },
        },
    },
    {
        "type": "function",
        "function": {
            "name": "add_event",
            "description": "Write an event to the calendar.",
            "parameters": {
                "type": "object",
                "properties": {
                    "date": {"type": "string", "description": "Event date (YYYY-MM-DD)"},
                    "time": {"type": "string", "description": "Event time (HH:MM)"},
                    "title": {"type": "string", "description": "Event title"},
                },
                "required": ["date", "time", "title"],
            },
        },
    },
]

Tool descriptions make or break the agent

The most critical part of the schema is not the name but the description and each parameter's description — the LLM relies on descriptions to decide "when to use the tool and how to fill in the arguments." Vague descriptions (such as "handle weather") cause the model to skip the tool when it should use it, or to fill in arguments at random. Descriptions should state the trigger conditions, what each parameter means, and the valid value ranges. This is prompt engineering at heart; for the methodology see Prompt Engineering.

Between the tool implementations and the schema sits the dispatcher — it maps the name and JSON arguments requested by the model to the real functions:

python
def dispatch(name: str, args_json: str) -> str:
    """Dispatch the LLM's tool call request to the real implementation."""
    args = json.loads(args_json)
    if name == "get_weather":
        return get_weather(args.get("city"), args.get("date", "today"))
    if name == "add_event":
        return add_event(args["date"], args["time"], args["title"])
    return f"Error: unknown tool {name}"

2. The LLM tool-calling loop: ReAct ​

Once the tools are defined, everything comes down to that loop. The classic pattern is called ReAct (Reasoning + Acting): the model first "reasons" about what to do next, then "acts" — emitting a tool call request; your program executes the tool and feeds the result back as an "observation"; the model keeps reasoning based on the observation. The cycle repeats until the model decides the task is complete and outputs a final answer directly. For how this concept relates to prompt engineering, see Prompt Engineering.

Loop structure (pseudocode):
while steps < max_steps:
    reply = LLM(messages, tools)          # 1. let the model decide the next move
    if reply has no tool calls: return reply   # 2a. done → output the final answer
    for each tool call:
        result = dispatch(tool_name, args)    # 2b. execute the tool
        messages.append(tool message)         # 3. feed the observation back into the conversation

Full implementation (the current OpenAI API):

python
from openai import OpenAI

client = OpenAI()   # requires the OPENAI_API_KEY environment variable; for local Ollama see below

def run_agent(user_query: str, max_steps: int = 6, model: str = "gpt-4o-mini") -> str:
    messages = [{"role": "user", "content": user_query}]
    for step in range(max_steps):
        resp = client.chat.completions.create(
            model=model,
            messages=messages,
            tools=TOOLS,          # hand the tool schemas to the model
        )
        msg = resp.choices[0].message
        messages.append(msg)      # the model's reply (including any tool call requests) joins the conversation

        if not msg.tool_calls:    # the model decided to answer directly → task complete
            return msg.content

        # Execute the tool calls one by one; feed results back as role="tool" messages
        for call in msg.tool_calls:
            result = dispatch(call.function.name, call.function.arguments)
            print(f"[step {step}] called {call.function.name}({call.function.arguments}) -> {result}")
            messages.append({
                "role": "tool",
                "tool_call_id": call.id,
                "content": result,
            })
    return "Maximum step limit reached; task not completed."

Run it once:

python
print(run_agent("Is today good for a run in Beijing? Check the weather and then give me advice."))

A typical output trace (the lines printed by print) looks like this:

[step 0] called get_weather({"city":"Beijing","date":"today"}) -> {"city":"Beijing","date":"today","weather":"Sunny","temp":"27°C"}
It's sunny and 27°C in Beijing today — perfect for a run. Head out in the late afternoon, and remember to stay hydrated.

Three easily overlooked details

  1. Don't drop the message history: the model's reply msg must be appended to messages verbatim (including its tool_calls field); otherwise tool results no longer line up with their calls, and the model "loses its memory".
  2. One reply may call multiple tools: msg.tool_calls is a list — execute each call and append each result one by one.
  3. tool_call_id must be passed back: the tool_call_id in each tool result message corresponds one-to-one with a call; this is the key to keeping multi-tool scenarios error-free.

Local Ollama version: just swap the client and the model name; everything else stays the same, giving you a fully offline agent:

python
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
model = "qwen2.5:7b"     # must support function calling; run ollama pull qwen2.5 first

3. Adding memory ​

Everything inside the loop flows through messages — that alone is within-session memory: multi-turn dialogue and each step's tool results all live there. But two kinds of memory need explicit design:

Memory typeWhat it storesImplementationWhen it's useful
Within-session memoryEvery message in the current taskJust the messages listMulti-step tasks, multi-turn dialogue
Long-term memoryUser preferences, facts, past conclusionsFiles / databases / vector storesCross-session reuse, personalization
Working memoryIntermediate computation resultsVariables, temporary filesComplex computation tasks
Vector memorySemantically retrievable historyVector store + similarity searchRecall on demand when history grows long

Within-session memory needs no extra code (the loop above supports it naturally). The typical vector memory implementation: write the key conclusions produced during a task into a vector store, then retrieve them on demand instead of stuffing the entire history into the context (longer context costs more — see Inference Optimization and Quantization). For the underlying principles and implementation, see Vector Databases and Semantic Search:

python
import numpy as np
import faiss
from sentence_transformers import SentenceTransformer

mem_model = SentenceTransformer("BAAI/bge-small-zh-v1.5")
mem_index = faiss.IndexFlatIP(384)   # semantic memory store (in-process)
memory_texts: list[str] = []

def remember(text: str) -> None:
    v = mem_model.encode([text], normalize_embeddings=True)
    mem_index.add(v.astype("float32"))
    memory_texts.append(text)

def recall(query: str, k: int = 3) -> list[str]:
    if not memory_texts:
        return []
    qv = mem_model.encode([query], normalize_embeddings=True)
    _, ids = mem_index.search(qv.astype("float32"), k)
    return [memory_texts[i] for i in ids[0] if i >= 0]

# Example: store "the user goes for a morning run every Tuesday" in memory so it comes up automatically next time
remember("User habit: goes for a morning run at 7 AM every Tuesday")
print(recall("What are this user's exercise habits?"))   # ['User habit: goes for a morning run at 7 AM every Tuesday']

Managing memory capacity

All memory eventually becomes tokens in the context. For long sessions you must truncate or compress (keep the most recent N turns plus a summary); keep k small for vector recall. An agent without capacity management will blow up its context after about 20 turns of dialogue — expensive and slow.

4. Adding planning: task decomposition ​

Complex tasks ("help me plan my weekend") often cannot be finished in a single tool call. Add a planner on top: first have the LLM split the big task into ordered subtasks, then execute them one by one and merge the results. This corresponds to an agent's planning capability:

python
def plan(task: str) -> list[str]:
    """Split the task into 2-4 ordered subtasks, one per line."""
    resp = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[
            {"role": "system",
             "content": "Break the user's task into 2-4 ordered subtasks, one per line. Output only the subtasks themselves."},
            {"role": "user", "content": task},
        ],
    )
    return [ln.strip("- ").strip() for ln in resp.choices[0].message.content.splitlines() if ln.strip()]

Then run the execution loop with planning in place: first plan, then call run_agent for each subtask, and finally summarize:

python
def run_with_plan(task: str) -> str:
    steps = plan(task)
    print("Plan:", steps)
    results = [run_agent(f"Subtask: {s}") for s in steps]
    # Summarize the subtask results and hand them to the LLM to produce the final reply
    resp = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "system", "content": "Based on the subtask results below, give the user a complete, friendly final answer."},
                  {"role": "user", "content": "\n\n".join(f"### {s}\n{r}" for s, r in zip(steps, results))}],
    )
    return resp.choices[0].message.content

Two flavors of planning

There are two ways to plan: explicit planning (split first, then execute, as above — controllable and auditable) and implicit planning (the model thinks and acts on the fly inside the ReAct loop — flexible but hard to predict). At the beginning, go explicit: split, execute, summarize — when something breaks, you can trace it to a specific subtask. For planning boundaries in complex tasks and a deeper discussion of "task decomposition", see Agent Core Concepts.

5. The complete runnable demo: check the weather and add it to the schedule ​

Put everything above together into agent_demo.py. Target task: "Check tomorrow's weather in Beijing and Shanghai, and schedule an outdoor run at 3 PM tomorrow in whichever city has better weather."

python
"""
agent_demo.py — a minimal runnable agent: weather lookup + calendar write

Usage: python agent_demo.py
Dependencies: pip install openai   (for the local Ollama version, see the comment at the end of the code)
"""
import json
import random
from pathlib import Path

from openai import OpenAI

# ── Tool implementations ─────────────────────────────────────────────
def get_weather(city: str, date: str = "today") -> str:
    conditions = ["Sunny", "Cloudy", "Light rain", "Windy"]
    return json.dumps({"city": city, "date": date,
                       "weather": random.choice(conditions),
                       "temp": f"{random.randint(18, 34)}°C"},
                      ensure_ascii=False)

def add_event(date: str, time: str, title: str) -> str:
    path = Path("events.json")
    events = json.loads(path.read_text(encoding="utf-8")) if path.exists() else []
    events.append({"date": date, "time": time, "title": title})
    path.write_text(json.dumps(events, ensure_ascii=False, indent=2), encoding="utf-8")
    return f"OK, added '{title}' to {date} {time}"

# ── function calling schema ──────────────────────────────────────────
TOOLS = [
    {"type": "function", "function": {
        "name": "get_weather",
        "description": "Get the weather for a city on a given day. Returns the condition and temperature.",
        "parameters": {"type": "object",
                       "properties": {"city": {"type": "string", "description": "City name"},
                                      "date": {"type": "string", "description": "Date; defaults to today"}},
                       "required": ["city"]}}},
    {"type": "function", "function": {
        "name": "add_event",
        "description": "Write an event to the calendar.",
        "parameters": {"type": "object",
                       "properties": {"date": {"type": "string", "description": "Date YYYY-MM-DD"},
                                      "time": {"type": "string", "description": "Time HH:MM"},
                                      "title": {"type": "string", "description": "Title"}},
                       "required": ["date", "time", "title"]}}},
]

def dispatch(name: str, args_json: str) -> str:
    args = json.loads(args_json)
    if name == "get_weather":
        return get_weather(args.get("city"), args.get("date", "today"))
    if name == "add_event":
        return add_event(args["date"], args["time"], args["title"])
    return f"Error: unknown tool {name}"

# ── ReAct loop ───────────────────────────────────────────────────────
client = OpenAI()                       # requires OPENAI_API_KEY
# Local version: client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
#                model = "qwen2.5:7b"

def run_agent(user_query: str, max_steps: int = 8, model: str = "gpt-4o-mini") -> str:
    messages = [{"role": "user", "content": user_query}]
    for step in range(max_steps):
        resp = client.chat.completions.create(model=model, messages=messages, tools=TOOLS)
        msg = resp.choices[0].message
        messages.append(msg)
        if not msg.tool_calls:
            return msg.content
        for call in msg.tool_calls:
            result = dispatch(call.function.name, call.function.arguments)
            print(f"[step {step}] {call.function.name}({call.function.arguments}) -> {result}")
            messages.append({"role": "tool", "tool_call_id": call.id, "content": result})
    return "Maximum step limit reached; task not completed."

if __name__ == "__main__":
    q = "Check tomorrow's weather in Beijing and Shanghai, and schedule an outdoor run at 3 PM tomorrow in whichever city has better weather."
    print(run_agent(q))
    print("— Contents of events.json —")
    print(Path("events.json").read_text(encoding="utf-8"))

When you run it, you will see the model complete the entire chain on its own — look up two cities → compare → call add_event — and produce a well-reasoned final answer. This 80-line script is the complete skeleton of an agent — everything a framework does is engineering layered on top of this loop.

6. Choosing a framework: when not to roll your own ​

A loop you write yourself is right for learning and minimal scenarios; in real projects, tool count, state complexity, concurrency, and observability quickly outgrow a hand-written loop's comfort zone. A comparison of common frameworks:

FrameworkCore abstractionStrengthsCostsWhen to use
LangGraphGraph + state machineControllable flow, supports branches/loops/parallelism, checkpointingMany concepts, steep learning curveComplex multi-step flows needing precise control and resumable runs
OpenAI Agents SDKAgent + handoffOfficially maintained, function calling out of the box, little codeTied to the OpenAI ecosystemFast delivery, mostly single-agent
Claude Agent SDKAgent + tools + computer useOfficial Anthropic, great tooling/sandbox experienceTied to the Claude ecosystemDeep Claude usage, browser/computer control needed
Roll your own (this article)while loopZero dependencies, fully transparent, controllableLacks engineering features (state/retries/observability)Learning, fewer than 5 tools, fixed flow

The decision rule in one sentence: fewer than 5 tools and a fixed flow → rolling your own is enough; building a product that needs parallelism and auditability → use a framework. Whichever framework you pick, the mental model of "tool schema + loop + memory" is universal — frameworks are just different implementations of it. For model selection and cost reference, see Model and Leaderboard Quick Reference.

7. Engineering essentials: from demo to production ​

Once the hand-written loop works, these problems surface immediately in real-world scenarios:

1. Truncating tool returns ​

Tools may return very long text (database queries, file reads). Truncate tool returns before feeding them back, otherwise the context balloons and token costs spike:

python
def safe_tool_result(result: str, max_len: int = 4000) -> str:
    if len(result) > max_len:
        return result[:max_len] + f"\n... (truncated; the original text was {len(result)} characters)"
    return result

2. Error retries and fault tolerance ​

Tools may throw exceptions (API timeouts, invalid arguments). Convert exceptions into text and feed them back to the model so it can correct itself, instead of letting the whole loop crash:

python
try:
    result = dispatch(call.function.name, call.function.arguments)
except Exception as e:          # for more rigor, use isinstance to distinguish exception types
    result = f"Tool execution failed: {type(e).__name__}: {e}"
messages.append({"role": "tool", "tool_call_id": call.id, "content": safe_tool_result(result)})

3. Max steps and infinite-loop detection ​

max_steps is mandatory, and you must detect repeated actions (same tool + same arguments N times in a row) and declare failure immediately. Infinite loops are the most common "incident" pattern for agents.

4. Cost control ​

Every step of the loop is one LLM call. Ways to keep it under control: use a cheap small model for planning/simple tool calls, add semantic caching (similar questions hit the cache), and compress message history. For these optimizations, see Inference Optimization and Quantization and LLM Deployment and Inference Optimization in Practice.

5. Permissions and security ​

Tools are the only channel through which an agent affects the outside world — and the most dangerous part:

RiskScenarioCountermeasure
Prompt injectionInstructions hidden in tool-returned content (e.g. scraped web pages) lure the agent into malicious actionsTreat tool results as "data", not "instructions"; require confirmation for sensitive operations; isolate context
Unauthorized tool callsThe model calls a tool it shouldn't (e.g. a delete operation)Least privilege: authorize each tool separately, require user confirmation for sensitive operations
Sandboxed executionLetting the agent run generated codeRun inside a sandbox/container; forbid access to sensitive paths
Parameter validationTool arguments come from the model and may be invalidValidate types and value ranges at the dispatch layer

For security boundaries and governance frameworks, see AI Safety and Governance.

Two red lines

  1. Never hand irreversible tools like "delete" or "transfer money" to the agent for automatic invocation — always add a human confirmation step.
  2. External text mixed into tool output (web pages, emails, documents) is untrusted by default — it may carry attack instructions. Prompt injection is the most realistic attack surface against agents.

8. Evaluating agents and common pitfalls ​

Evaluation ​

Agents have to be evaluated on both process and result. The four most practical metrics:

MetricWhat it measuresHow to test
Task success rateDid the task complete end to endA batch of real tasks, judged success/failure by humans
Tool call correctnessWas the right tool called with the right argumentsCompare the "expected tool call sequence" against the actual one
Step efficiencyHow many steps it tookTrack average steps; outliers usually mean "taking detours"
CostTokens per taskLog token usage, correlated with task type

Recommendation: build an evaluation set of 20-50 tasks first, then change the loop logic, and run the eval after every change — don't rely on "it feels better". For building the pipeline, see Building an LLM Evaluation Pipeline; for the theory behind the metrics, see LLM Evaluation and Benchmarks.

Common pitfalls ​

PitfallSymptomRoot causeFix
Runaway loopsEndless tool calls, tokens burningNo step limit, model retrying endlesslymax_steps + repeated-action detection + timeout
Hallucinated tool argumentsCalls with nonexistent cities / wrong argument typesUnclear schema descriptions, model improvisingStrengthen argument descriptions, validate at the dispatch layer, retry when required arguments are missing
Context bloatIncreasingly slow and expensive, answers miss the pointFull history stacking upTruncate/compress/vector memory with on-demand recall
Tool result "contamination"Behavior suddenly goes wrongInstruction-like text mixed into tool returnsTruncate results, label them as data, isolate context
Over-calling toolsEven a simple "hello" triggers a tool callSchema trigger conditions written too broadlyState clearly in the description "when NOT to call"
Silent failureTask "looks done" but wasn't actually executedThe tool returned an error that was treated as successHave tools return structured status (ok/error) and check it explicitly in the loop

For a more complete list of agent antipatterns, see Common Pitfalls and Antipatterns.

9. Advanced: multi-agent collaboration ​

A single agent has limits on capacity and responsibility; for complex tasks, let multiple agents with distinct roles collaborate. Three mainstream orchestration patterns:

PatternStructureBest forExample
Orchestrator-workerA lead agent splits the task and assigns it to specialized sub-agentsDecomposable tasks with fixed subtask typesWriting a report = retrieval agent + writing agent + formatting agent
PipelineEach agent's output is the next one's inputTasks with a natural orderResearch → analysis → summary
Debate/reviewMultiple agents answer independently, then critique each otherHigh correctness requirementsCode review, fact-checking

The engineering practices that go with it: sub-agents interact only through "messages/results" (no shared mutable state), each sub-agent gets its own isolated context, and the lead agent handles aggregation and conflict resolution. For a complete teardown of a multi-agent product, see Manus and Agent Applications; what determines the collaboration structure ultimately traces back to the discussion of "task complexity and agent shape" in Agent Core Concepts.

Closing checklist

A "qualified" agent: completes tasks reliably, has step-count and cost guardrails, gets human confirmation for irreversible tool operations, and has an evaluation set that quantifies "nothing got worse". With those four boxes ticked, it is a system you can have an engineering conversation about — not a demo. Look up unfamiliar terms in the Glossary; for the big picture, see What Is AI: Key Concepts.

Further Reading ​

References ​