Skip to content

OpenAI Agents SDK

At a glance OpenAI's official lightweight agent framework — the production-grade successor to Swarm. Six primitives (Agent / Handoff / Guardrail / Session / Tracing) cover most agent applications; it defaults to the Responses API but is not locked to OpenAI models.

This page contains time-sensitive content; data is current as of 2026-08. Job listings, pricing, and product features may have changed — verify against the original sources before citing.

OpenAI Agents SDK ​

In the framework selection overview, the OpenAI Agents SDK's positioning can be stated in one sentence: the one with the least abstraction. It has only a handful of primitives, the orchestration logic is ordinary Python code, and it introduces no graphs and no DSL. The price is giving up fine-grained control over the execution flow — the last section of this page compares it head-to-head with LangGraph to judge whether that trade is worth it.

1. Positioning and History: From Swarm to a Production-Grade SDK ​

Timeline (all verified):

  • 2024: OpenAI released the experimental Swarm project, introducing the two concepts of Agent and handoff, with the explicit disclaimer "educational exploration only, not for production."
  • 2025-03-11: OpenAI shipped the Responses API, three built-in tools, and the Agents SDK; the Agents SDK was positioned as the production-ready upgrade of Swarm. The Traces observability dashboard launched the same day. The Swarm repo was archived shortly after and redirected to the Agents SDK.
  • 2025-03-27: The SDK officially added MCP support — you can attach local or remote MCP servers.
  • 2025-2026: Rapid iteration. As of 2026-08, the latest Python package openai-agents is v0.22.0 (still 0.x; the official stance is that minor version bumps signal "substantive runtime changes"). Along the way it gained Sessions (multiple backends), Realtime/Voice agents, a Sandbox agent, Programmatic Tool Calling, and more. The default models have evolved to gpt-5.6-luna (low-cost default) / gpt-5.6-sol (high capability).

The relationship to the Responses API ​

This is the key layer for understanding the SDK. The Responses API is OpenAI's new API primitive from March 2025, merging the simplicity of Chat Completions with the tooling capability of the Assistants API (which has been announced for deprecation, with a target sunset in mid-2026). The Agents SDK is not the API — it is a runtime on top of the API:

Your code
   │
   ▼
Agents SDK (Agent / Runner / Guardrail / Session ...)   ← runs the agent loop, tool dispatch, guardrails, handoffs
   │
   ▼
Responses API (built-in tools: web_search / file_search / computer ...) ← OpenAI's server side
   │
   ▼
Models (the gpt-5.6 family, etc.)

The official docs' selection criteria are blunt:

  • Want to control the loop, tool dispatch, and state yourself → use the Responses API directly;
  • Want the runtime to manage multi-turn loops, tool execution, guardrails, handoffs, and sessions for you → use the Agents SDK.

The two aren't mutually exclusive, and many production systems are hybrid: the main flow runs on the SDK while a few latency-sensitive paths call the Responses API directly.

What a 0.x version number means

The SDK is still 0.x. Its release cadence is extremely fast (nearly every 1-2 weeks within 2026), and minor versions occasionally carry behavior changes (for example, v0.22.0 tightened the configuration contract for OpenAIProvider with an explicit client). For production use, pin the version and watch the Releases page.

2. Core Primitives: Six Concepts Carry the Whole Framework ​

The official design principle is "enough features to be worth using, few enough primitives to actually learn." There are exactly six:

PrimitiveIn one sentenceCorresponding code
AgentAn LLM with instructions and toolsAgent(name=..., instructions=..., tools=[...])
RunnerThe runtime that executes the agent loopRunner.run() / run_sync() / run_streamed()
HandoffTransfers control between agentshandoffs=[other_agent]
GuardrailValidation and circuit-breaking for inputs / outputs / tool calls@input_guardrail / @output_guardrail
SessionConversation memory across runsSQLiteSession("conv_123")
TracingBuilt-in observabilityOn by default, zero code

The Runner's internal loop looks roughly like this:

        ┌────────────────────────── Runner loop ──────────────────────────┐
        │                                                                 │
 input ─┼─▶ [input guardrails] ─▶ LLM call ─▶ tool call? ───yes───▶ run tool ─┐
        │        (parallel/blocking)    │              │                    │
        │                               │             no            (feed back)│
        │                               ▼              │                    │
        │                          handoff? ───yes───▶ switch agent ─────────┤
        │                               │              │                    │
        │                              no              ▼                    │
        │                               │        [output guardrails]         │
        │                               │              │                    │
        └───────────────────────────────┴──────────────┴────────────────────┘
                                                             ▼
                                                     final_output

Agent ​

python
from agents import Agent, ModelSettings

agent = Agent(
    name="Support assistant",
    instructions="You are an after-sales support agent. Only answer questions about orders and refunds.",
    model="gpt-5.6-sol",                      # omit to use the default model (currently gpt-5.6-luna)
    model_settings=ModelSettings(temperature=0.2),
    tools=[...],                              # function tools / hosted tools / MCP tools
    handoffs=[...],                           # downstream agents you can hand off to
)

instructions accepts a static string or a function that receives context (dynamically injecting tenant info, user profiles, and so on — a standard context engineering move). Agent is a generic class, Agent[TContext], which pairs with RunContextWrapper for typed dependency injection.

Runner ​

Runner.run(agent, input, ...) is the async entry point; run_sync is a synchronous wrapper; run_streamed returns a streaming result you consume event by event via stream_events(). Important parameters: max_turns (guards against infinite loops, raises MaxTurnsExceeded when exceeded), session, context, run_config. The returned RunResult carries final_output, new_items (every item produced this run), last_agent, and more.

Handoff: making "transfer" a first-class citizen ​

Handoff is the SDK's most distinctive design. Under the hood it is modeled as a tool the model can call: when the model calls transfer_to_xxx, the Runner hands control of the conversation (including the full history) to the target agent. This differs from the "manager pattern" (the main agent calls a sub-agent and takes control back — implemented in the SDK via agent.as_tool()). For when to use which, see Multi-Agent Architecture.

python
refund_agent = Agent(name="Refund specialist", instructions="Handle refunds. Always verify the order number first.")
triage_agent = Agent(
    name="Triage",
    instructions="Hand refund questions to the refund specialist; answer everything else yourself.",
    handoffs=[refund_agent],
)

For customization (renaming the tool, passing structured input, filtering history), use handoff(agent, on_handoff=..., input_type=..., input_filter=...).

Guardrails: guardrails, not decoration ​

Three kinds of guardrail, each with a different execution point — the most common beginner trap:

  • Input guardrails run only when the first agent in the chain receives user input. By default they run in parallel with the agent (run_in_parallel=True), which gives the lowest latency; but when a tripwire fires, the main model may already have burned some tokens. To save cost or prevent tool side effects, set run_in_parallel=False to make it blocking.
  • Output guardrails run only after the last agent produces the final result, and never in parallel.
  • Tool guardrails wrap every function tool call (before/after execution), suitable for rules like "never pass secrets to an external API." Note they do not apply to handoffs or hosted tools.

A guardrail returns GuardrailFunctionOutput(tripwire_triggered=...); when it trips, the Runner raises InputGuardrailTripwireTriggered / OutputGuardrailTripwireTriggered, which you catch at the application layer and degrade gracefully. The typical pattern is a cheap small model acting as the guardrail agent alongside an expensive main model. For more patterns, see Security & Guardrails.

Sessions: conversation memory, official edition ​

Sessions solve multi-turn conversation history management — no more manually feeding result.to_input_list() back in. After Runner.run(..., session=session), the Runner automatically fetches the history before each run and stores new items after it.

The built-in backends are already quite complete: SQLiteSession (default, file or in-memory), AsyncSQLiteSession, RedisSession, SQLAlchemySession, MongoDBSession, DaprSession, OpenAIConversationsSession (history stored on OpenAI's servers), OpenAIResponsesCompactionSession (auto-compaction for very long conversations), plus custom backends implementing the Session protocol (the four methods get_items / add_items / pop_item / clear_session). The difference from long-term memory: a session manages "working context," while cross-session user profiles and accumulated knowledge still need your own infrastructure.

Tracing ​

On by default, zero code. Every run automatically produces a trace, organized internally as spans: agent spans, generation spans (LLM calls), function spans (tools), guardrail spans, handoff spans. By default it batch-uploads to the OpenAI Traces dashboard. See Section 5 for details.

3. Built-in Tools: Hosted Tools Are the Biggest Differentiator ​

The SDK's tools come in five classes, and the truly unique one is hosted tools — the tool executes on OpenAI's servers, so your code only declares it, never implements it:

ToolWhat it doesNotes
WebSearchToolWeb search with citationsSupports user_location, search_context_size
FileSearchToolRetrieval over OpenAI Vector StoresManaged RAG with metadata filters and reranking
ComputerToolGUI/browser operationExecutes locally: the model emits actions, you implement the Computer interface (screenshot, click, etc.) and feed results back
CodeInterpreterToolRuns code in a sandboxServer-side container
HostedMCPToolConnects directly to a remote MCP serverServer-side execution
ImageGenerationToolText-to-imageServer-side execution
python
from agents import Agent, FileSearchTool, Runner, WebSearchTool

agent = Agent(
    name="Research assistant",
    tools=[
        WebSearchTool(),                       # web search
        FileSearchTool(                        # managed RAG, replaces a self-built retrieval pipeline
            max_num_results=3,
            vector_store_ids=["vs_xxx"],
        ),
    ],
)
result = Runner.run_sync(agent, "Compare plan A in our product docs with the latest industry practice today")

A few judgments:

  • FileSearchTool fits the "quickly wire up internal docs" scenario and eliminates an entire RAG pipeline; but the retrieval logic is a black box and the data goes to OpenAI, so fine-grained retrieval needs should still be self-built.
  • ComputerTool nominally belongs to the hosted tool family, but the actual actions execute on your machine (the official example uses a Playwright harness). The state of things in 2026: computer use went GA starting with gpt-5.5, with the tool payload migrating from the preview computer_use_preview (single action) to computer (batch actions); the SDK picks the wire format automatically based on the model actually requested. Before GA, CUA models scored only 38.1% on OSWorld — for desktop scenarios, definitely add human confirmation.
  • Custom tools use the @function_tool decorator (or the shorter @tool alias since 0.19) on any Python function; the schema is generated automatically from type annotations plus the docstring (based on Pydantic + griffe). This covers 90% of day-to-day development.
  • MCP tools: local servers use MCPServerStdio, remote ones use MCPServerStreamableHttp / MCPServerSse, attached to the agent's mcp_servers.

4. Complete Example: Customer-Support Triage with Guardrails, Handoffs, and a Session ​

The example below wires together all the primitives above; the syntax matches the official docs of the 0.2x line (pip install openai-agents):

python
import asyncio
from pydantic import BaseModel
from agents import (
    Agent, Runner, SQLiteSession, WebSearchTool,
    GuardrailFunctionOutput, InputGuardrailTripwireTriggered,
    RunContextWrapper, TResponseInputItem,
    function_tool, input_guardrail,
)

# ---------- 1. Custom function tool ----------
@function_tool
def lookup_order(order_id: str) -> str:
    """Look up order status by order number.

    Args:
        order_id: the order number, e.g. OD-12345
    """
    # In production this would call your order system
    return f"Order {order_id}: shipped, expected within 3 days"

# ---------- 2. Input guardrail: a small model intercepts off-topic requests ----------
class RelevanceCheck(BaseModel):
    is_off_topic: bool
    reasoning: str

guardrail_agent = Agent(
    name="Topic check",
    instructions="Decide whether the user is asking something unrelated to e-commerce support (e.g. coding, math homework).",
    model="gpt-5-mini",                 # guardrails use a cheap model
    output_type=RelevanceCheck,
)

@input_guardrail
async def relevance_guardrail(
    ctx: RunContextWrapper[None],
    agent: Agent,
    input: str | list[TResponseInputItem],
) -> GuardrailFunctionOutput:
    result = await Runner.run(guardrail_agent, input, context=ctx.context)
    return GuardrailFunctionOutput(
        output_info=result.final_output,
        tripwire_triggered=result.final_output.is_off_topic,
    )

# ---------- 3. Specialist agents + triage agent (handoffs) ----------
order_agent = Agent(
    name="Order specialist",
    instructions="You only handle order queries. Call lookup_order before answering.",
    tools=[lookup_order],
)

research_agent = Agent(
    name="Product researcher",
    instructions="You answer product comparisons and buying advice; search the web when needed.",
    tools=[WebSearchTool()],
)

triage_agent = Agent(
    name="Support triage",
    instructions=(
        "You are the entry point for e-commerce support. Hand order queries to the order "
        "specialist and product questions to the product researcher; answer everything "
        "else yourself, briefly."
    ),
    handoffs=[order_agent, research_agent],
    input_guardrails=[relevance_guardrail],
)

# ---------- 4. Multi-turn runs with a session ----------
async def main():
    session = SQLiteSession("user_42", "conversations.db")  # history persisted to disk

    try:
        r1 = await Runner.run(triage_agent, "Where is my order OD-12345?",
                              session=session, max_turns=10)
        print("Support:", r1.final_output)

        # The second turn automatically carries the previous history
        r2 = await Runner.run(triage_agent, "While you're at it, any cheaper alternatives to recommend?",
                              session=session, max_turns=10)
        print("Support:", r2.final_output)

    except InputGuardrailTripwireTriggered:
        print("Support: That question is outside my scope, sorry!")

if __name__ == "__main__":
    asyncio.run(main())

Before running, export OPENAI_API_KEY=sk-.... This skeleton works as a production starting point as-is: guardrails handle boundaries, handoffs handle division of labor, sessions handle memory, and tracing is recording everything by default. To turn it into a course project, see the breakdown in the hands-on tutorial.

5. Tracing and Evaluation Integration ​

Observability is one of the SDK's hidden advantages over other frameworks, because it's there by default:

  • Every run is wrapped in a trace (named Agent workflow by default; use with trace("name") to customize and merge multiple runs into one trace).
  • Spans cover: LLM generations, function tool calls, guardrails, handoffs, speech transcription/synthesis, and more.
  • Free viewing: the Traces dashboard on the OpenAI platform. Don't want traces going to OpenAI? Three switches: the environment variable OPENAI_AGENTS_DISABLE_TRACING=1, set_tracing_disabled(True), or RunConfig(tracing_disabled=True). Tracing is unavailable for organizations under ZDR (zero data retention).
  • Custom backends: add_trace_processor() appends, set_trace_processors() replaces. The official docs list a long roster of third-party integrations — Langfuse, LangSmith, Arize Phoenix, Braintrust, Pydantic Logfire, AgentOps, MLflow, Datadog, Langtrace, and twenty-odd more — basically covering the mainstream agent observability stack.

On the evaluation side: trace data can be distilled into datasets and fed to OpenAI's Evals product for offline evaluation, and the real LLM call records inside traces can even be used for distillation/fine-tuning. If your organization already has an evaluation pipeline, spans exported via third-party processors can feed your in-house evals too — see the methodology in Evaluation Systems and Evals in Practice.

Flush traces for long-running tasks

The default BatchTraceProcessor uploads in batches every few seconds and flushes on process exit. In long-lived workers like Celery or FastAPI BackgroundTasks, if you need a trace visible in the dashboard the moment a task finishes, call flush_traces() explicitly after the trace() context exits.

6. Model Neutrality vs Vendor Lock-in: Half Neutral, Half Bound ​

"Can the Agents SDK work without OpenAI models?" Yes — but know the boundary.

The neutral part:

  • Any provider offering an OpenAI-compatible endpoint, wired up in three lines:
python
from agents import Agent, AsyncOpenAI, OpenAIChatCompletionsModel, set_tracing_disabled

set_tracing_disabled(True)  # turn off when you have no OpenAI key, or configure tracing with its own key

client = AsyncOpenAI(base_url="https://your-provider/v1", api_key="...")
agent = Agent(
    name="Assistant",
    model=OpenAIChatCompletionsModel(model="provider-model-name", openai_client=client),
)
  • Broader coverage comes via third-party adapters (beta): LitellmModel from openai-agents[litellm] and AnyLLMModel from openai-agents[any-llm], claiming 100+ models. You can also mix models per agent — a small model for triage and large models for specialists is a common cost optimization move.
  • Tracing is decoupled from the model: even when models come from elsewhere, traces can still go to the OpenAI dashboard with a separate OpenAI key (set_tracing_export_api_key) — free-riding on its observability.

The bound part:

  • All hosted tools (web search / file search / computer / code interpreter) work only on the OpenAI Responses path. Switch providers and they're all gone — you'd have to rebuild search and retrieval yourself.
  • A batch of advanced features is Responses-only: tool search, ProgrammaticToolCallingTool, server-side compaction, previous_response_id chaining, and so on. On the Chat Completions adapter these fields get silently dropped (set strict_feature_validation=True to turn that into errors).
  • Structured output and multimodal input support is patchy on non-OpenAI providers; OpenAI itself warns you to "check the feature matrix before switching providers."

Conclusion: the SDK's runtime abstractions (loop, handoff, guardrail, session) are model-neutral, but its fully-loaded form is bound to the OpenAI platform. If your strategy is all-in on OpenAI, that's fine; if you need multi-cloud / multi-model, list the hosted tools as "dependencies that would need replacing" in your architecture diagram.

7. Pros, Cons, and Where It Fits ​

DimensionAssessment
Learning curveExtremely low. Six primitives + plain Python; an afternoon to get productive
Code intrusivenessLow. Agents are dataclasses, orchestration is function calls — easy to test and refactor
ObservabilityDefault tracing + dashboard + rich third-party integrations; first tier in its class
GuardrailsInput/output/tool guardrails are built-in first-class citizens; most frameworks bolt them on
Ecosystem bindingHosted tools, the Traces dashboard, and Evals all live inside the OpenAI platform — deep users benefit but are also constrained
ControlThe Runner's black-box loop resists customization: complex state machines, branch rollback, and fine-grained checkpoints are awkward
Maturity0.x, fast iteration, occasional behavior changes; by 2026 the core API (Agent/Runner/handoff/guardrail) is quite stable, with changes mostly at the edges
LanguagesPython-first; an official TypeScript version (@openai/agents) exists with roughly aligned features

Good fit: new businesses on an OpenAI stack (support, research assistants, content pipelines); teams that need to ship fast and don't want to maintain an orchestration framework; triage/routing-style multi-agent systems where handoff is the core interaction pattern.

Poor fit: complex deterministic workflow orchestration (state machines, approval flows, precise retry semantics) — that's LangGraph's home turf; strongly regulated / private deployments that can't touch OpenAI; research projects that need deep customization of loop behavior (in that case, better to hand-write the loop).

8. The Trade-off Against LangGraph ​

This is the most frequently asked pairing in selection discussions. It isn't "which is better" — they are two different worldviews:

OpenAI Agents SDKLangGraph
Orchestration paradigmImplicit: the Runner runs the loop as a black box; handoffs decentralize transfersExplicit: you draw the state graph; nodes/edges/conditions are all yours to define
State modelSessions manage conversation history; business state goes into your own contextFirst-class State schema + checkpointer; persistable, replayable, forkable
Control granularityCoarse. The loop's internals aren't customizable — only hooks and guardrailsFine. Every step, every interrupt, every checkpoint is under your control
Time to learnHoursDays (graph thinking has a learning cost)
Debugging experienceThe Traces dashboard works out of the boxLangSmith is equally strong, but the graph structure itself is more explainable
Model bindingNeutral runtime, swappable models; fully-loaded form binds to OpenAICompletely model-agnostic
Typical scenariosConversational, triage/routing, quick launchesLong-horizon tasks, approval flows, multi-step deterministic processes

Rules of thumb:

  1. First ask whether you need a graph at all. If your "workflow" is really just "triage + a few specialists + guardrails," the Agents SDK ships in a day and LangGraph is over-engineering.
  2. Switch to LangGraph once these signals appear: you need precise interrupt/resume semantics, you need intermediate state in your own database for audit, the flow has genuinely complex branching and merging (not just "who takes over"), or you can't depend on OpenAI at all.
  3. The two can coexist. Plenty of teams use the Agents SDK for the user-facing conversational layer and delegate one heavy process node to a LangGraph subgraph — frameworks are not a religion.

Don't be held hostage by the word "official"

The Agents SDK is OpenAI's official product, but "official" doesn't mean "automatically right." Its design is deeply tied to the capability boundaries of the Responses API, and the conveniences you enjoy (hosted tools, the tracing dashboard) are the very lock-in points. Write down the question "if I had to drop OpenAI tomorrow, what would I rewrite?" and answer it before going to production.

References ​