Appearance
OpenAI Agents SDK
In the framework selection overview, the OpenAI Agents SDK's positioning can be stated in one sentence: the one with the least abstraction. It has only a handful of primitives, the orchestration logic is ordinary Python code, and it introduces no graphs and no DSL. The price is giving up fine-grained control over the execution flow — the last section of this page compares it head-to-head with LangGraph to judge whether that trade is worth it.
1. Positioning and History: From Swarm to a Production-Grade SDK
Timeline (all verified):
- 2024: OpenAI released the experimental Swarm project, introducing the two concepts of
Agentandhandoff, with the explicit disclaimer "educational exploration only, not for production." - 2025-03-11: OpenAI shipped the Responses API, three built-in tools, and the Agents SDK; the Agents SDK was positioned as the production-ready upgrade of Swarm. The Traces observability dashboard launched the same day. The Swarm repo was archived shortly after and redirected to the Agents SDK.
- 2025-03-27: The SDK officially added MCP support — you can attach local or remote MCP servers.
- 2025-2026: Rapid iteration. As of 2026-08, the latest Python package
openai-agentsis v0.22.0 (still 0.x; the official stance is that minor version bumps signal "substantive runtime changes"). Along the way it gained Sessions (multiple backends), Realtime/Voice agents, a Sandbox agent, Programmatic Tool Calling, and more. The default models have evolved togpt-5.6-luna(low-cost default) /gpt-5.6-sol(high capability).
The relationship to the Responses API
This is the key layer for understanding the SDK. The Responses API is OpenAI's new API primitive from March 2025, merging the simplicity of Chat Completions with the tooling capability of the Assistants API (which has been announced for deprecation, with a target sunset in mid-2026). The Agents SDK is not the API — it is a runtime on top of the API:
Your code
│
▼
Agents SDK (Agent / Runner / Guardrail / Session ...) ← runs the agent loop, tool dispatch, guardrails, handoffs
│
▼
Responses API (built-in tools: web_search / file_search / computer ...) ← OpenAI's server side
│
▼
Models (the gpt-5.6 family, etc.)The official docs' selection criteria are blunt:
- Want to control the loop, tool dispatch, and state yourself → use the Responses API directly;
- Want the runtime to manage multi-turn loops, tool execution, guardrails, handoffs, and sessions for you → use the Agents SDK.
The two aren't mutually exclusive, and many production systems are hybrid: the main flow runs on the SDK while a few latency-sensitive paths call the Responses API directly.
What a 0.x version number means
The SDK is still 0.x. Its release cadence is extremely fast (nearly every 1-2 weeks within 2026), and minor versions occasionally carry behavior changes (for example, v0.22.0 tightened the configuration contract for OpenAIProvider with an explicit client). For production use, pin the version and watch the Releases page.
2. Core Primitives: Six Concepts Carry the Whole Framework
The official design principle is "enough features to be worth using, few enough primitives to actually learn." There are exactly six:
| Primitive | In one sentence | Corresponding code |
|---|---|---|
| Agent | An LLM with instructions and tools | Agent(name=..., instructions=..., tools=[...]) |
| Runner | The runtime that executes the agent loop | Runner.run() / run_sync() / run_streamed() |
| Handoff | Transfers control between agents | handoffs=[other_agent] |
| Guardrail | Validation and circuit-breaking for inputs / outputs / tool calls | @input_guardrail / @output_guardrail |
| Session | Conversation memory across runs | SQLiteSession("conv_123") |
| Tracing | Built-in observability | On by default, zero code |
The Runner's internal loop looks roughly like this:
┌────────────────────────── Runner loop ──────────────────────────┐
│ │
input ─┼─▶ [input guardrails] ─▶ LLM call ─▶ tool call? ───yes───▶ run tool ─┐
│ (parallel/blocking) │ │ │
│ │ no (feed back)│
│ ▼ │ │
│ handoff? ───yes───▶ switch agent ─────────┤
│ │ │ │
│ no ▼ │
│ │ [output guardrails] │
│ │ │ │
└───────────────────────────────┴──────────────┴────────────────────┘
▼
final_outputAgent
python
from agents import Agent, ModelSettings
agent = Agent(
name="Support assistant",
instructions="You are an after-sales support agent. Only answer questions about orders and refunds.",
model="gpt-5.6-sol", # omit to use the default model (currently gpt-5.6-luna)
model_settings=ModelSettings(temperature=0.2),
tools=[...], # function tools / hosted tools / MCP tools
handoffs=[...], # downstream agents you can hand off to
)instructions accepts a static string or a function that receives context (dynamically injecting tenant info, user profiles, and so on — a standard context engineering move). Agent is a generic class, Agent[TContext], which pairs with RunContextWrapper for typed dependency injection.
Runner
Runner.run(agent, input, ...) is the async entry point; run_sync is a synchronous wrapper; run_streamed returns a streaming result you consume event by event via stream_events(). Important parameters: max_turns (guards against infinite loops, raises MaxTurnsExceeded when exceeded), session, context, run_config. The returned RunResult carries final_output, new_items (every item produced this run), last_agent, and more.
Handoff: making "transfer" a first-class citizen
Handoff is the SDK's most distinctive design. Under the hood it is modeled as a tool the model can call: when the model calls transfer_to_xxx, the Runner hands control of the conversation (including the full history) to the target agent. This differs from the "manager pattern" (the main agent calls a sub-agent and takes control back — implemented in the SDK via agent.as_tool()). For when to use which, see Multi-Agent Architecture.
python
refund_agent = Agent(name="Refund specialist", instructions="Handle refunds. Always verify the order number first.")
triage_agent = Agent(
name="Triage",
instructions="Hand refund questions to the refund specialist; answer everything else yourself.",
handoffs=[refund_agent],
)For customization (renaming the tool, passing structured input, filtering history), use handoff(agent, on_handoff=..., input_type=..., input_filter=...).
Guardrails: guardrails, not decoration
Three kinds of guardrail, each with a different execution point — the most common beginner trap:
- Input guardrails run only when the first agent in the chain receives user input. By default they run in parallel with the agent (
run_in_parallel=True), which gives the lowest latency; but when a tripwire fires, the main model may already have burned some tokens. To save cost or prevent tool side effects, setrun_in_parallel=Falseto make it blocking. - Output guardrails run only after the last agent produces the final result, and never in parallel.
- Tool guardrails wrap every function tool call (before/after execution), suitable for rules like "never pass secrets to an external API." Note they do not apply to handoffs or hosted tools.
A guardrail returns GuardrailFunctionOutput(tripwire_triggered=...); when it trips, the Runner raises InputGuardrailTripwireTriggered / OutputGuardrailTripwireTriggered, which you catch at the application layer and degrade gracefully. The typical pattern is a cheap small model acting as the guardrail agent alongside an expensive main model. For more patterns, see Security & Guardrails.
Sessions: conversation memory, official edition
Sessions solve multi-turn conversation history management — no more manually feeding result.to_input_list() back in. After Runner.run(..., session=session), the Runner automatically fetches the history before each run and stores new items after it.
The built-in backends are already quite complete: SQLiteSession (default, file or in-memory), AsyncSQLiteSession, RedisSession, SQLAlchemySession, MongoDBSession, DaprSession, OpenAIConversationsSession (history stored on OpenAI's servers), OpenAIResponsesCompactionSession (auto-compaction for very long conversations), plus custom backends implementing the Session protocol (the four methods get_items / add_items / pop_item / clear_session). The difference from long-term memory: a session manages "working context," while cross-session user profiles and accumulated knowledge still need your own infrastructure.
Tracing
On by default, zero code. Every run automatically produces a trace, organized internally as spans: agent spans, generation spans (LLM calls), function spans (tools), guardrail spans, handoff spans. By default it batch-uploads to the OpenAI Traces dashboard. See Section 5 for details.
3. Built-in Tools: Hosted Tools Are the Biggest Differentiator
The SDK's tools come in five classes, and the truly unique one is hosted tools — the tool executes on OpenAI's servers, so your code only declares it, never implements it:
| Tool | What it does | Notes |
|---|---|---|
WebSearchTool | Web search with citations | Supports user_location, search_context_size |
FileSearchTool | Retrieval over OpenAI Vector Stores | Managed RAG with metadata filters and reranking |
ComputerTool | GUI/browser operation | Executes locally: the model emits actions, you implement the Computer interface (screenshot, click, etc.) and feed results back |
CodeInterpreterTool | Runs code in a sandbox | Server-side container |
HostedMCPTool | Connects directly to a remote MCP server | Server-side execution |
ImageGenerationTool | Text-to-image | Server-side execution |
python
from agents import Agent, FileSearchTool, Runner, WebSearchTool
agent = Agent(
name="Research assistant",
tools=[
WebSearchTool(), # web search
FileSearchTool( # managed RAG, replaces a self-built retrieval pipeline
max_num_results=3,
vector_store_ids=["vs_xxx"],
),
],
)
result = Runner.run_sync(agent, "Compare plan A in our product docs with the latest industry practice today")A few judgments:
FileSearchToolfits the "quickly wire up internal docs" scenario and eliminates an entire RAG pipeline; but the retrieval logic is a black box and the data goes to OpenAI, so fine-grained retrieval needs should still be self-built.ComputerToolnominally belongs to the hosted tool family, but the actual actions execute on your machine (the official example uses a Playwright harness). The state of things in 2026: computer use went GA starting withgpt-5.5, with the tool payload migrating from the previewcomputer_use_preview(single action) tocomputer(batch actions); the SDK picks the wire format automatically based on the model actually requested. Before GA, CUA models scored only 38.1% on OSWorld — for desktop scenarios, definitely add human confirmation.- Custom tools use the
@function_tooldecorator (or the shorter@toolalias since 0.19) on any Python function; the schema is generated automatically from type annotations plus the docstring (based on Pydantic + griffe). This covers 90% of day-to-day development. - MCP tools: local servers use
MCPServerStdio, remote ones useMCPServerStreamableHttp/MCPServerSse, attached to the agent'smcp_servers.
4. Complete Example: Customer-Support Triage with Guardrails, Handoffs, and a Session
The example below wires together all the primitives above; the syntax matches the official docs of the 0.2x line (pip install openai-agents):
python
import asyncio
from pydantic import BaseModel
from agents import (
Agent, Runner, SQLiteSession, WebSearchTool,
GuardrailFunctionOutput, InputGuardrailTripwireTriggered,
RunContextWrapper, TResponseInputItem,
function_tool, input_guardrail,
)
# ---------- 1. Custom function tool ----------
@function_tool
def lookup_order(order_id: str) -> str:
"""Look up order status by order number.
Args:
order_id: the order number, e.g. OD-12345
"""
# In production this would call your order system
return f"Order {order_id}: shipped, expected within 3 days"
# ---------- 2. Input guardrail: a small model intercepts off-topic requests ----------
class RelevanceCheck(BaseModel):
is_off_topic: bool
reasoning: str
guardrail_agent = Agent(
name="Topic check",
instructions="Decide whether the user is asking something unrelated to e-commerce support (e.g. coding, math homework).",
model="gpt-5-mini", # guardrails use a cheap model
output_type=RelevanceCheck,
)
@input_guardrail
async def relevance_guardrail(
ctx: RunContextWrapper[None],
agent: Agent,
input: str | list[TResponseInputItem],
) -> GuardrailFunctionOutput:
result = await Runner.run(guardrail_agent, input, context=ctx.context)
return GuardrailFunctionOutput(
output_info=result.final_output,
tripwire_triggered=result.final_output.is_off_topic,
)
# ---------- 3. Specialist agents + triage agent (handoffs) ----------
order_agent = Agent(
name="Order specialist",
instructions="You only handle order queries. Call lookup_order before answering.",
tools=[lookup_order],
)
research_agent = Agent(
name="Product researcher",
instructions="You answer product comparisons and buying advice; search the web when needed.",
tools=[WebSearchTool()],
)
triage_agent = Agent(
name="Support triage",
instructions=(
"You are the entry point for e-commerce support. Hand order queries to the order "
"specialist and product questions to the product researcher; answer everything "
"else yourself, briefly."
),
handoffs=[order_agent, research_agent],
input_guardrails=[relevance_guardrail],
)
# ---------- 4. Multi-turn runs with a session ----------
async def main():
session = SQLiteSession("user_42", "conversations.db") # history persisted to disk
try:
r1 = await Runner.run(triage_agent, "Where is my order OD-12345?",
session=session, max_turns=10)
print("Support:", r1.final_output)
# The second turn automatically carries the previous history
r2 = await Runner.run(triage_agent, "While you're at it, any cheaper alternatives to recommend?",
session=session, max_turns=10)
print("Support:", r2.final_output)
except InputGuardrailTripwireTriggered:
print("Support: That question is outside my scope, sorry!")
if __name__ == "__main__":
asyncio.run(main())Before running, export OPENAI_API_KEY=sk-.... This skeleton works as a production starting point as-is: guardrails handle boundaries, handoffs handle division of labor, sessions handle memory, and tracing is recording everything by default. To turn it into a course project, see the breakdown in the hands-on tutorial.
5. Tracing and Evaluation Integration
Observability is one of the SDK's hidden advantages over other frameworks, because it's there by default:
- Every run is wrapped in a trace (named
Agent workflowby default; usewith trace("name")to customize and merge multiple runs into one trace). - Spans cover: LLM generations, function tool calls, guardrails, handoffs, speech transcription/synthesis, and more.
- Free viewing: the Traces dashboard on the OpenAI platform. Don't want traces going to OpenAI? Three switches: the environment variable
OPENAI_AGENTS_DISABLE_TRACING=1,set_tracing_disabled(True), orRunConfig(tracing_disabled=True). Tracing is unavailable for organizations under ZDR (zero data retention). - Custom backends:
add_trace_processor()appends,set_trace_processors()replaces. The official docs list a long roster of third-party integrations — Langfuse, LangSmith, Arize Phoenix, Braintrust, Pydantic Logfire, AgentOps, MLflow, Datadog, Langtrace, and twenty-odd more — basically covering the mainstream agent observability stack.
On the evaluation side: trace data can be distilled into datasets and fed to OpenAI's Evals product for offline evaluation, and the real LLM call records inside traces can even be used for distillation/fine-tuning. If your organization already has an evaluation pipeline, spans exported via third-party processors can feed your in-house evals too — see the methodology in Evaluation Systems and Evals in Practice.
Flush traces for long-running tasks
The default BatchTraceProcessor uploads in batches every few seconds and flushes on process exit. In long-lived workers like Celery or FastAPI BackgroundTasks, if you need a trace visible in the dashboard the moment a task finishes, call flush_traces() explicitly after the trace() context exits.
6. Model Neutrality vs Vendor Lock-in: Half Neutral, Half Bound
"Can the Agents SDK work without OpenAI models?" Yes — but know the boundary.
The neutral part:
- Any provider offering an OpenAI-compatible endpoint, wired up in three lines:
python
from agents import Agent, AsyncOpenAI, OpenAIChatCompletionsModel, set_tracing_disabled
set_tracing_disabled(True) # turn off when you have no OpenAI key, or configure tracing with its own key
client = AsyncOpenAI(base_url="https://your-provider/v1", api_key="...")
agent = Agent(
name="Assistant",
model=OpenAIChatCompletionsModel(model="provider-model-name", openai_client=client),
)- Broader coverage comes via third-party adapters (beta):
LitellmModelfromopenai-agents[litellm]andAnyLLMModelfromopenai-agents[any-llm], claiming 100+ models. You can also mix models per agent — a small model for triage and large models for specialists is a common cost optimization move. - Tracing is decoupled from the model: even when models come from elsewhere, traces can still go to the OpenAI dashboard with a separate OpenAI key (
set_tracing_export_api_key) — free-riding on its observability.
The bound part:
- All hosted tools (web search / file search / computer / code interpreter) work only on the OpenAI Responses path. Switch providers and they're all gone — you'd have to rebuild search and retrieval yourself.
- A batch of advanced features is Responses-only: tool search,
ProgrammaticToolCallingTool, server-side compaction,previous_response_idchaining, and so on. On the Chat Completions adapter these fields get silently dropped (setstrict_feature_validation=Trueto turn that into errors). - Structured output and multimodal input support is patchy on non-OpenAI providers; OpenAI itself warns you to "check the feature matrix before switching providers."
Conclusion: the SDK's runtime abstractions (loop, handoff, guardrail, session) are model-neutral, but its fully-loaded form is bound to the OpenAI platform. If your strategy is all-in on OpenAI, that's fine; if you need multi-cloud / multi-model, list the hosted tools as "dependencies that would need replacing" in your architecture diagram.
7. Pros, Cons, and Where It Fits
| Dimension | Assessment |
|---|---|
| Learning curve | Extremely low. Six primitives + plain Python; an afternoon to get productive |
| Code intrusiveness | Low. Agents are dataclasses, orchestration is function calls — easy to test and refactor |
| Observability | Default tracing + dashboard + rich third-party integrations; first tier in its class |
| Guardrails | Input/output/tool guardrails are built-in first-class citizens; most frameworks bolt them on |
| Ecosystem binding | Hosted tools, the Traces dashboard, and Evals all live inside the OpenAI platform — deep users benefit but are also constrained |
| Control | The Runner's black-box loop resists customization: complex state machines, branch rollback, and fine-grained checkpoints are awkward |
| Maturity | 0.x, fast iteration, occasional behavior changes; by 2026 the core API (Agent/Runner/handoff/guardrail) is quite stable, with changes mostly at the edges |
| Languages | Python-first; an official TypeScript version (@openai/agents) exists with roughly aligned features |
Good fit: new businesses on an OpenAI stack (support, research assistants, content pipelines); teams that need to ship fast and don't want to maintain an orchestration framework; triage/routing-style multi-agent systems where handoff is the core interaction pattern.
Poor fit: complex deterministic workflow orchestration (state machines, approval flows, precise retry semantics) — that's LangGraph's home turf; strongly regulated / private deployments that can't touch OpenAI; research projects that need deep customization of loop behavior (in that case, better to hand-write the loop).
8. The Trade-off Against LangGraph
This is the most frequently asked pairing in selection discussions. It isn't "which is better" — they are two different worldviews:
| OpenAI Agents SDK | LangGraph | |
|---|---|---|
| Orchestration paradigm | Implicit: the Runner runs the loop as a black box; handoffs decentralize transfers | Explicit: you draw the state graph; nodes/edges/conditions are all yours to define |
| State model | Sessions manage conversation history; business state goes into your own context | First-class State schema + checkpointer; persistable, replayable, forkable |
| Control granularity | Coarse. The loop's internals aren't customizable — only hooks and guardrails | Fine. Every step, every interrupt, every checkpoint is under your control |
| Time to learn | Hours | Days (graph thinking has a learning cost) |
| Debugging experience | The Traces dashboard works out of the box | LangSmith is equally strong, but the graph structure itself is more explainable |
| Model binding | Neutral runtime, swappable models; fully-loaded form binds to OpenAI | Completely model-agnostic |
| Typical scenarios | Conversational, triage/routing, quick launches | Long-horizon tasks, approval flows, multi-step deterministic processes |
Rules of thumb:
- First ask whether you need a graph at all. If your "workflow" is really just "triage + a few specialists + guardrails," the Agents SDK ships in a day and LangGraph is over-engineering.
- Switch to LangGraph once these signals appear: you need precise interrupt/resume semantics, you need intermediate state in your own database for audit, the flow has genuinely complex branching and merging (not just "who takes over"), or you can't depend on OpenAI at all.
- The two can coexist. Plenty of teams use the Agents SDK for the user-facing conversational layer and delegate one heavy process node to a LangGraph subgraph — frameworks are not a religion.
Don't be held hostage by the word "official"
The Agents SDK is OpenAI's official product, but "official" doesn't mean "automatically right." Its design is deeply tied to the capability boundaries of the Responses API, and the conveniences you enjoy (hosted tools, the tracing dashboard) are the very lock-in points. Write down the question "if I had to drop OpenAI tomorrow, what would I rewrite?" and answer it before going to production.
References
- OpenAI: New tools for building agents (2025-03-11 announcement) — the original release note for the Responses API, built-in tools, and the Agents SDK
- OpenAI Agents SDK official docs — the authoritative reference for primitives, tools, Sessions, and Tracing
- openai/openai-agents-python (GitHub) — source code and the examples directory, the best material for learning the API
- Releases · openai-agents-python — version changelog, a must-read before production upgrades
- Models docs: non-OpenAI model integration — the Chat Completions adapter, LiteLLM/Any-LLM adapters, and feature differences
- Guardrails docs — the execution semantics of the input/output/tool guardrail layers
- Tracing docs — the trace/span model and the list of twenty-plus third-party observability integrations