Skip to content

Progressive Tutorial: Three Versions, All Running

At a glance Three progressive projects built with the OpenAI Agents SDK: a weather chat Agent you can get running in 30 minutes, a research assistant with a todo list and file-based memory, and a researcher/writer/critic multi-Agent setup with a minimal eval script—plus a comparison table of the three versions' complexity and capability boundaries.

Progressive Tutorial: Three Versions, All Running ​

Build It Yourself walked you through hand-writing a minimal Agent Loop—that was internal training against the raw principles. This page takes the opposite path: build three progressive projects directly on a production-grade framework (the OpenAI Agents SDK), with each version adding one capability you can't avoid in real engineering on top of the last:

  • V1: single Agent + tool calling + multi-turn chat—a chatbot that can look up real weather, with the goal of running in 30 minutes;
  • V2: add planning (a todo checklist) and file-based memory, turning it into a "search → read → write report" research assistant that can carry tasks of 10 minutes or more;
  • V3: split into researcher / writer / critic roles, plus a minimal 5-task eval script, turning "feels good" into "80% pass rate."

All three versions evolve inside one project—no three isolated demos. By the end, you should be able to take any of the three versions and reshape it into your own project skeleton.

Why the OpenAI Agents SDK

It's currently one of the production-grade frameworks with the fewest abstractions: the core is just four primitives—Agent, Runner, tools, and handoff—and by the end of V1 you'll have seen them all. For comparisons and selection rationale, see the Framework Overview and the OpenAI Agents SDK page. The code in this article is based on SDK v0.22.x (released July 2026); treat the official docs as authoritative on API details.

0. Prerequisites ​

bash
mkdir agent-tutorial && cd agent-tutorial
python -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate
pip install -U openai-agents httpx
export OPENAI_API_KEY=sk-...     # Windows PowerShell: $env:OPENAI_API_KEY = "sk-..."

Two things to note:

  • From v0.20.0 onward, openai-agents uses gpt-5.6-luna as its default model (the official release notes state this change explicitly). If you don't set the model parameter, you get the default; the code in this article never specifies a model explicitly—just follow the default.
  • Not specifying model doesn't mean you can only call OpenAI—the SDK supports other compatible endpoints via provider configuration, but to keep the tutorial focused, everything here uses the default configuration.

Directory layout (shared by all three versions):

agent-tutorial/
├── v1_weather.py      # V1: weather chat
├── v2_researcher.py   # V2: research assistant
├── v3_team.py         # V3: multi-Agent team
├── evals.py           # V3: minimal eval script
└── workspace/         # file memory and artifacts for V2/V3 (created automatically by the code)

1. V1: Single Agent + Tools, Running in 30 Minutes ​

Goal ​

A command-line chatbot that answers questions like "What's the weather in Beijing right now?" or "Is Shanghai hotter than Shenzhen?" The skills involved: tool definitions, the Agent Loop (the model decides on its own to look up the city's coordinates first, then the weather), and multi-turn conversation memory.

Architecture ​

User input ──► Runner.run(agent, input, session)
                  │
                  ▼
            ┌── Agent Loop ──────────────────┐
            │  LLM decides to call geocode_city
            │    → tool returns coordinates  │
            │  LLM decides to call get_current_weather
            │    → tool returns weather JSON │
            │  LLM produces a natural-language answer
            └────────────────────────────────┘
                  │
                  ▼
            final_output (also written to SQLiteSession)

Weather data comes from Open-Meteo: free, no API key, no registration—ideal for teaching. It splits into two endpoints—geocoding (city name → coordinates) and forecast (coordinates → weather)—which conveniently demonstrates "the model chaining multiple tools together on its own."

Full code ​

python
# v1_weather.py
import json

import httpx
from agents import Agent, Runner, SQLiteSession
from agents.decorators import tool  # equivalent to the older `from agents import function_tool`


@tool
def geocode_city(city: str) -> str:
    """Resolve a city name to latitude/longitude. Prefer English names or pinyin, e.g. Beijing, Shanghai."""
    resp = httpx.get(
        "https://geocoding-api.open-meteo.com/v1/search",
        params={"name": city, "count": 1, "language": "zh"},
        timeout=10,
    )
    results = resp.json().get("results") or []
    if not results:
        return f"City {city} not found; try a different spelling"
    r = results[0]
    return json.dumps(
        {"name": r["name"], "latitude": r["latitude"], "longitude": r["longitude"]},
        ensure_ascii=False,
    )


@tool
def get_current_weather(latitude: float, longitude: float) -> str:
    """Look up the current weather at the given coordinates; returns temperature (°C) and a weather condition code."""
    resp = httpx.get(
        "https://api.open-meteo.com/v1/forecast",
        params={
            "latitude": latitude,
            "longitude": longitude,
            "current": "temperature_2m,weather_code,relative_humidity_2m",
        },
        timeout=10,
    )
    return json.dumps(resp.json()["current"], ensure_ascii=False)


agent = Agent(
    name="Weather Assistant",
    instructions=(
        "You are a weather assistant. For any weather question, you must first call geocode_city to resolve the city's coordinates, "
        "then call get_current_weather for live data—never invent temperatures from memory. "
        "weather_code is a WMO code (0 clear, 1-3 partly cloudy, 45/48 fog, 51-67 rain, 71-77 snow, 95+ thunderstorms); "
        "translate it into natural language in your answer. When comparing cities, query them one by one."
    ),
    tools=[geocode_city, get_current_weather],
)

session = SQLiteSession("weather_chat", "workspace/chat.db")  # file-persisted session memory

if __name__ == "__main__":
    print("Weather assistant online. Type quit to exit.")
    while True:
        question = input("\nYou: ").strip()
        if question.lower() in {"quit", "exit"}:
            break
        result = Runner.run_sync(agent, question, session=session)
        print(f"Assistant: {result.final_output}")

What it looks like running ​

You: What's the weather in Beijing right now?
Assistant: Beijing is currently 32.4°C with light drizzle—take an umbrella if you're heading out.

You: What about Shanghai? Hotter than Beijing?      ← the session remembers the earlier city
Assistant: Shanghai is at 34.1°C, cloudy. About 1.7 degrees hotter than Beijing (32.4°C).

Three details worth stopping to look at:

  1. A tool's docstring is interface documentation aimed at the model. Writing "prefer English names or pinyin" in geocode_city's docstring teaches the model how to construct arguments—it's part of the prompt, not a comment. Style details in Prompt Engineering.
  2. You didn't write a single line of the loop. Runner.run internally is the full Agent Loop: call the model → if there's a tool call, execute it and feed the result back into the context → call the model again → until the model outputs plain text. Compare against the hand-written version to see what the framework encapsulates for you.
  3. SQLiteSession solves multi-turn memory in one line. The Runner reads history automatically before each run and writes new entries after; the database lands in workspace/chat.db, and the memory survives a process restart. More memory strategies in Memory Systems.

Return JSON strings from tools, not Python objects

Tool return values get serialized into the conversation history. Returning a json.dumps(...) string is the most reliable: the model can read it, and the session can store it. Returning custom objects or dicts triggers serialization warnings in some versions.

What you learned in this version ​

  • Agent = name + instructions + tool list; there's no more magic than that;
  • The @tool decorator generates JSON Schema automatically from the function signature and docstring;
  • Runner.run / run_sync encapsulate the complete tool-calling loop;
  • SQLiteSession provides out-of-the-box multi-turn memory.

V1's capability boundary is also clear: no planning. Ask it to "survey the current state of Agent frameworks and write a report," and it flails through a couple of searches and turns in something hasty—because there's no mechanism for it to break a big task into steps or remember intermediate artifacts. That's what V2 adds.

2. V2: A Research Assistant with Planning and Memory ​

Goal ​

Hand it a research topic (e.g. "compare multi-Agent framework choices in 2026") and the Agent autonomously: breaks the topic into sub-questions → searches the web one by one → reads the sources → saves notes to disk → assembles a markdown report. The whole task may run a dozen-plus tool-call rounds, and if it crashes midway, a restart picks up where it left off.

This version introduces three new mechanisms:

  • todo planning: the Agent must first write its plan as a checklist file, then update statuses as it completes each step;
  • file memory: search results and reading notes are written to workspace/ instead of piling up in the context window—the key to long tasks not "forgetting things";
  • web tools: the SDK's built-in WebSearchTool (OpenAI-hosted web search), plus a self-written read_url to fetch page content.

Architecture ​

Research topic
   │
   ▼
┌─ Research Agent (max_turns=25) ──────────────────────┐
│ 1. update_todo: split the topic into 4-6 sub-questions → plan.md
│ 2. Loop: web_search → read_url → save_note           │
│         (notes land in workspace/notes/*.md)         │
│ 3. Once all sub-questions are done, read back the notes → write report.md
└───────────────────────────────────────────────────────┘
   │
   ▼
workspace/report.md (final deliverable)

The context window is not memory

In V1 everything lives in the conversation history, which is fine within 10 turns; but research tasks routinely involve dozens of tool calls and tens of thousands of tokens of page text, and stuffing it all into context is both expensive and a trigger for lost-in-the-middle. V2's principle: raw material goes into files; the context only holds pointers (file paths + summaries). That's the core idea of Context Engineering, and the shared practice of long-task products like Claude Code and Manus.

Key code skeleton ​

The tool layer (files are the memory; deliberately no vector store):

python
# v2_researcher.py (excerpt; tool section shown in full)
import json
import re
from pathlib import Path

import httpx
from agents import Agent, Runner, WebSearchTool
from agents.decorators import tool

WORKSPACE = Path("workspace")
NOTES = WORKSPACE / "notes"


@tool
def update_todo(plan_markdown: str) -> str:
    """Write or update the task plan. plan_markdown is the complete plan,
    using - [ ] / - [x] for pending and completed items. Call this every time a subtask starts or finishes."""
    WORKSPACE.mkdir(exist_ok=True)
    (WORKSPACE / "plan.md").write_text(plan_markdown, encoding="utf-8")
    return "Plan updated at workspace/plan.md"


@tool
def read_url(url: str) -> str:
    """Fetch a web page's main text (crude extraction, truncated to 6000 chars) for reading search results in depth."""
    try:
        resp = httpx.get(url, timeout=15, follow_redirects=True)
        resp.raise_for_status()
    except Exception as e:
        return f"Fetch failed: {e}"
    text = re.sub(r"<(script|style)[^>]*>.*?</\1>", "", resp.text, flags=re.S)
    text = re.sub(r"<[^>]+>", " ", text)          # strip HTML tags
    text = re.sub(r"\s+", " ", text).strip()
    return text[:6000]                             # truncate rather than flood the context


@tool
def save_note(filename: str, content: str) -> str:
    """Save a research note to workspace/notes/. filename excludes path and extension."""
    NOTES.mkdir(parents=True, exist_ok=True)
    (NOTES / f"{filename}.md").write_text(content, encoding="utf-8")
    return f"Note saved: notes/{filename}.md"


@tool
def read_all_notes() -> str:
    """Read every note in the notes directory (call once before writing the report)."""
    if not NOTES.exists():
        return "No notes yet"
    parts = [f"===== {p.name} =====\n{p.read_text(encoding='utf-8')}"
             for p in sorted(NOTES.glob("*.md"))]
    return "\n\n".join(parts)


@tool
def write_report(content: str) -> str:
    """Write the final research report to workspace/report.md."""
    WORKSPACE.mkdir(exist_ok=True)
    (WORKSPACE / "report.md").write_text(content, encoding="utf-8")
    return "Report written to workspace/report.md"

The Agent definition—notice how the instructions hard-code "workflow discipline," which is the main lever for keeping the Agent on the rails:

python
# v2_researcher.py (Agent and entry-point section)
agent = Agent(
    name="Research Assistant",
    instructions=(
        "You are a rigorous research assistant. On receiving a topic, follow this workflow exactly:\n"
        "1. First call update_todo: break the topic into 4-6 independently searchable sub-questions, written as a checklist;\n"
        "2. Research the sub-questions one by one: web_search to find sources → read_url on the 1-2 most valuable links "
        "→ save_note to store the key facts (with source URLs) as notes → update_todo to mark it done;\n"
        "3. When all sub-questions are done, call read_all_notes to consolidate, then use write_report to produce a structured report: "
        "summary, thematic sections (with sources cited), points of contention and uncertainty, and a reference list;\n"
        "4. Your final reply only needs to summarize the conclusion and the report path.\n"
        "Principles: every concrete fact must come from a searched source; never invent numbers or dates from memory; "
        "every fact in a note must carry its source URL."
    ),
    tools=[WebSearchTool(), update_todo, read_url, save_note, read_all_notes, write_report],
)

if __name__ == "__main__":
    import asyncio
    topic = input("Research topic: ").strip()
    result = asyncio.run(Runner.run(agent, topic, max_turns=25))  # long tasks need a higher turn cap
    print(result.final_output)

A few design decisions, explained:

  • max_turns=25 is a seatbelt, not a target. One turn = one model call. A dozen-plus turns is normal for research, but there must be a cap to keep a runaway loop from burning tokens. V1 doesn't set one because chat tasks wrap up in 3-5 turns.
  • read_url is the teaching edition—its text extraction is regex hacking and comes back empty on JS-rendered pages. In production, swap in trafilatura, jina reader, or a headless browser; the function signature doesn't need to change.
  • The todo is a file, not an in-memory variable, which buys a cheap but useful capability: after an interruption and rerun, the Agent reads plan.md and notes/ and knows exactly where it left off. For systematic planning patterns (ReAct, Plan-and-Execute, etc.) see Planning & Reasoning.

What it looks like running ​

Research topic: Compare the leading 2026 Agent frameworks and give selection advice

(Agent internals: update_todo splits out 5 sub-questions → searches and reads 8 pages
  → saves 6 notes → generates the report)

Research complete. Core conclusion: LangGraph suits complex orchestration that needs fine-grained state control;
OpenAI Agents SDK suits fast delivery......full report at workspace/report.md.

workspace/ ends up containing plan.md (a checklist with checked states), notes/ (6 sourced notes), and report.md (the final report)—every process artifact is auditable, so when something goes wrong you can trace exactly which search went sideways.

What you learned in this version ​

  • Long task = planning (todo) + external memory (files) + a turn cap—all three are indispensable;
  • Instructions are the carrier of workflow discipline; the more it reads like an SOP, the steadier the Agent;
  • Pointers in context, originals on disk: the survival rule for long tasks;
  • But "one Agent does the whole pipeline" is starting to strain: search, writing, and review share a single context, prompts keep growing, and quality drags on quality. That's the motivation for V3's role split.

3. V3: Multi-Agent and a Minimal Eval ​

Goal ​

Split V2's single Agent into three roles, and answer the question we've been dodging all along: how reliable is it, actually?

  • Researcher: only searches and reads sources; outputs a fact list with sources attached;
  • Writer: only organizes the fact list into a report; not allowed to invent facts of its own;
  • Critic: only reviews—checks whether facts are backed by sources and the structure is complete, then sends it back or lets it through.

Alongside, write an evals.py: 5 fixed test tasks, each with a predefined pass criterion, printing a pass rate at the end. This is the watershed between "toy" and "engineering."

Architecture ​

The SDK offers two multi-Agent orchestration styles: handoff (a triage Agent hands the whole conversation over to a specialist) and agents as tools (a manager Agent calls specialists as tools and always owns the final answer). A research report needs someone accountable for the final artifact, so we pick the latter:

                User topic
                   │
                   ▼
        ┌──── Orchestrator ────┐
        │  tools:              │
        │   research(...)  ────┼──► Researcher ──► WebSearchTool / read_url
        │   write_report(...) ─┼──► Writer      (no tools, writes only)
        │   critique(...)  ────┼──► Critic     (no tools, reviews only)
        └──────────────────────┘
                   │   On FAIL, sends the critique back to the Writer for a rewrite (≤2 rounds)
                   ▼
              workspace/report.md

Why not let the three Agents hand off freely? Because research → writing → review is a directed flow, not a routing problem. Letting the LLM decide the order (agents as tools) is already enough; if you want more determinism, V3's code can easily be rewritten as three sequential Runner.run calls—see the trade-off discussion below.

Key code skeleton ​

The three specialist Agents (each with very short, single-purpose instructions—this is the core payoff of splitting roles):

python
# v3_team.py (excerpt: role definitions)
from agents import Agent, Runner, WebSearchTool
from agents.decorators import tool
# read_url and write_report are reused from v2: from v2_researcher import read_url, write_report

researcher = Agent(
    name="Researcher",
    instructions=(
        "You only do research. Search and read primary sources around the topic and output a fact list: "
        "one sentence per fact + the source URL. No commentary, no essay."
    ),
    tools=[WebSearchTool(), read_url],
)

writer = Agent(
    name="Writer",
    instructions=(
        "You only write. Organize the given fact list into a structured research report: "
        "summary, thematic sections, points of contention, references. "
        "You must not add any concrete fact, number, or date beyond the fact list; "
        "when you receive revision feedback, address each point and output a complete new version."
    ),
    tools=[write_report],
)

critic = Agent(
    name="Critic",
    instructions=(
        "You are a strict reviewer. Check the report: "
        "1) does every concrete fact have a source behind it; 2) is the structure complete; 3) any exaggeration or fabrication. "
        "Output VERDICT: PASS or VERDICT: FAIL; on FAIL you must give specific, actionable revision points."
    ),
)

The orchestrator calls the three roles as tools:

python
# v3_team.py (excerpt: orchestration and entry point)
orchestrator = Agent(
    name="Orchestrator",
    instructions=(
        "You coordinate a research-and-writing team. Workflow:\n"
        "1. Call research to get the fact list;\n"
        "2. Call write_report (Writer) to draft the report from the fact list and save it to disk;\n"
        "3. Call critique to review the draft;\n"
        "4. On FAIL, send the critique along with the fact list back to write_report for a rewrite, at most 2 rewrites;\n"
        "5. Final reply: report path + review verdict + number of rewrites."
    ),
    tools=[
        researcher.as_tool(tool_name="research",
                           tool_description="Research the topic; return a sourced fact list"),
        writer.as_tool(tool_name="write_report",
                       tool_description="Write a report from the fact list (accepts revision feedback for rewrites)"),
        critic.as_tool(tool_name="critique",
                       tool_description="Review report quality; returns PASS/FAIL with comments"),
    ],
)

if __name__ == "__main__":
    import asyncio
    topic = input("Research topic: ").strip()
    result = asyncio.run(Runner.run(orchestrator, topic, max_turns=20))
    print(result.final_output)

agent.as_tool() wraps a sub-Agent into an ordinary tool: what the Orchestrator sees are three "functions," while underneath, each call is an independent sub-Agent run with its own clean context. Context isolation is exactly the engineering value of splitting roles—the tens of thousands of words of web text the Researcher read never pollute the Writer's window.

Multi-Agent is over-engineering most of the time

V3 splits into three roles because "research-writing-review" genuinely interferes across contexts and prompts. If your task passes evals reliably with a single Agent plus a good prompt, don't split—every extra Agent means an extra LLM call, another prompt to maintain, another point of failure. For the full discussion of when to split, see Multi-Agent Architectures.

evals.py: the minimal eval script ​

No framework, 30 lines: 5 tasks, keyword criteria, pass-rate stats.

python
# evals.py
import asyncio

from agents import Runner
from v3_team import orchestrator

# (task, pass criteria—keyword groups the report should contain; all must hit to pass)
TASKS = [
    ("Research the design goals and transport modes of Model Context Protocol",
     ["MCP", "stdio"]),
    ("Compare the applicable scenarios of LangGraph and the OpenAI Agents SDK",
     ["LangGraph", "OpenAI Agents SDK"]),
    ("Research mainstream chunking strategies in RAG systems",
     ["chunk", " embedding"]),
    ("Summarize the core idea of the ReAct paper",
     ["ReAct", "reasoning"]),
    ("Survey the mainstream AI Agent benchmarks of 2026",
     ["benchmark"]),
]


async def run_eval():
    passed = 0
    for i, (task, keywords) in enumerate(TASKS, 1):
        result = await Runner.run(orchestrator, task, max_turns=20)
        report = open("workspace/report.md", encoding="utf-8").read()
        missing = [k for k in keywords if k.lower() not in report.lower()]
        ok = not missing
        passed += ok
        print(f"[{i}/5] {'PASS' if ok else 'FAIL (missing: ' + ','.join(missing) + ')'} | {task[:20]}...")
    print(f"\nPass rate: {passed}/{len(TASKS)} = {passed / len(TASKS):.0%}")


if __name__ == "__main__":
    asyncio.run(run_eval())

Typical output:

[1/5] PASS | Research Model Context Proto...
[2/5] PASS | Compare LangGraph and the Ope...
[3/5] FAIL (missing: embedding) | Research chunking in RAG s...
[4/5] PASS | Summarize the core idea of th...
[5/5] PASS | Survey 2026 AI Agent benchm...

Pass rate: 4/5 = 80%

The naivety of this script is deliberate—keyword matching misfires (a report that says "vector" but never "embedding" fails). But its value isn't precision; it's that for the first time you have a baseline you can compare against before and after a change: tweaked the Orchestrator's prompt? Run evals.py—60% → 80% is a real improvement; a drop to 40% means you broke it. Without it, every one of your "optimizations" is astrology. The natural next steps: swap keywords for LLM-as-judge, grow the task set to 30+, wire in a platform like Langfuse—see Evaluation Systems and Evals in Practice.

Eval task sets overfit too

With only 5 tasks, it's easy to unconsciously tune your prompts to "just barely pass these 5." Rule: only ever add tasks, never modify existing ones, and top up regularly; if the pass rate rises after a prompt change but your gut says nothing improved, suspect the task set first.

What you learned in this version ​

  • Agents as tools: the manager keeps control of the whole while each specialist works in a clean context;
  • Splitting roles is about context isolation and prompt focus, not "more hands make light work";
  • A minimal eval = a fixed task set + machine-checkable pass criteria + a pass rate; 30 lines is enough to start;
  • Critic-rejects-and-rewrites is the cheapest self-improvement loop.

4. Three Versions Compared: Complexity, Code Size, and Capability Boundaries ​

DimensionV1 weather chatV2 research assistantV3 multi-Agent team
Code size (approx.)60 lines120 lines150 lines + 30-line eval
Number of Agents114 (orchestrator + 3 specialists)
Number of tools2 (self-written)5 self-written + WebSearchTool2 self-written + WebSearchTool + 3 agent tools
Planningnonetodo file checklistorchestration flow baked into the instructions
MemorySQLiteSession (chat history)file system (notes/plan/report)file system + sub-Agent context isolation
Turns per task3-5 turns10-25 turns15-40 turns (including sub-Agents)
Typical durationseconds3-10 minutes5-15 minutes
Quality assurancenonesource citations, human spot checksCritic review + eval pass rate
Capability boundarysingle-fact Q&Asingle-topic research, limited depthdivisible long tasks, but coordination cost rises
Main failure modewrong tool argumentsno good sources found, note gapsorchestrator skips review, sub-Agent output format drift

One thread runs through all three versions: every increment of complexity maps to a specific failure mode. V2 adds file memory because long tasks don't fit in context; V3 splits roles because a single Agent's prompts interfere with each other; the eval exists because "feels good" can't be trusted. If your scenario never hits that failure mode, stop at the corresponding version—that's engineering judgment, not laziness.

5. Where to Go Next ​

Once all three versions run, pick a direction based on your goal:

  • Keep building projects: turn V3 into your own work—switch to a different vertical (investment research, competitor monitoring, literature reviews), add human confirmation (Human-in-the-Loop) and cost tracking (Cost Optimization). More ideas in Portfolio Projects.
  • Go deeper on frameworks: read LangGraph (explicit state graphs, good for complex orchestration) and Claude Agent SDK side by side, to feel the spectrum of "how much state the framework manages for you."
  • Fill in the theory: return to the core components series and digest the concepts this article used (Agent Loop, context engineering, planning, memory) one by one; then read core papers like ReAct.
  • Job hunting: turn the V3 + eval experience into a quantified resume line ("built a multi-Agent research system, authored a 5-task eval set, iterated prompts to lift the pass rate from 60% to 80%"); for how to write it, see Resume Analysis.

Whichever path you take, keep the mindset from evals.py: define "what counts as good" before you start optimizing.

References ​