Skip to content

Claude Agent SDK

At a glance The agent harness that powers Claude Code, opened up as an official SDK — file/Bash tools out of the box, with subagents, Skills, hooks, MCP, and a permission system built in. Build production-grade long-running agents in a few dozen lines of Python or TypeScript.

This page contains time-sensitive content; data is current as of 2026-08. Job listings, pricing, and product features may have changed — verify against the original sources before citing.

Claude Agent SDK ​

The Claude Agent SDK is Anthropic taking the agent harness that powers Claude Code (the agent loop, built-in tools, context management, permission system) and extracting it into a programmable library, available in both Python and TypeScript. You no longer have to write the tool-use loop yourself — a single query() call lets Claude read files, run commands, edit code, and search the web inside your application's process until the task is done.

1. Positioning: Claude Code's Harness, Opened Up as an SDK ​

The SDK's origin determines its character. Anthropic was explicit in its official September 2025 blog post: inside the company, Claude Code had long outgrown being just a coding tool — they used the same harness for deep research, video creation, and note-taking, and it "drives almost all of our major agent loops." So they renamed the library from Claude Code SDK to Claude Agent SDK to reflect that it goes far beyond coding. The rename happened in September 2025 and completed within that month (first released as the Claude Code SDK, renamed at the end of the month): the Python package moved from claude-code-sdk to claude-agent-sdk, the TypeScript package from @anthropic-ai/claude-code-sdk to @anthropic-ai/claude-agent-sdk, and the core options class was renamed from ClaudeCodeOptions to ClaudeAgentOptions.

The best way to understand its positioning is to compare it with Anthropic's lower-level Messages API:

DimensionMessages API (Client SDK)Claude Agent SDK
Abstraction levelSingle request/responseA complete agent loop
Tool executionYou write handlers and assemble results yourselfThe SDK executes tools; file reads/Bash work out of the box
Context managementYou maintain the messages array yourselfAutomatic, with compaction as you approach the limit
Permission controlNo concept of it — you intercept everything yourselfBuilt-in permission modes / hooks / tool allowlists
Session persistenceStore it in your own databaseSessions persist as JSONL files and can be resumed across processes
Best forChat, structured extraction, custom loopsLong-running, multi-step agents that operate on a real environment

A common analogy: the Messages API is the engine, the Agent SDK is the whole car. Most people who want to build agents want the car. Its architectural position looks roughly like this:

┌─────────────────────────────────────────────────┐
│       Your application (Python / TS process)    │
│   query() / ClaudeSDKClient                     │
├─────────────────────────────────────────────────┤
│       Claude Agent SDK (agent harness)          │
│  agent loop │ permissions │ hooks │ compaction  │
│  subagent orchestration │ sessions │ MCP client │
├─────────────────────────────────────────────────┤
│       Claude Code CLI (bundled by the SDK)      │
│  Read/Write/Edit/Bash/Glob/Grep/WebSearch...    │
├─────────────────────────────────────────────────┤
│       Claude models (Anthropic API /            │
│       Bedrock / Vertex / Foundry)               │
└─────────────────────────────────────────────────┘

The relationship to Claude Code

The SDK's Python package automatically bundles the Claude Code CLI for your platform — no separate installation needed; the TS package likewise ships the native binary. So "using the SDK" essentially means programmatically driving Claude Code's core engine inside your own process, rather than interacting with it as a human in a terminal. For the product form itself, see the Claude Code case study.

2. Core Philosophy: Give the Model a Computer ​

The SDK's design philosophy is condensed in the official blog into one sentence: give your agents a computer. Don't build a pile of narrow API wrappers for the model — hand it the same tools programmers use every day, the filesystem and the terminal, and let it work the way a person does: find the files, write code, run it, read the error, fix it, iterate until it succeeds.

Three corollaries follow directly:

  1. The filesystem is the context infrastructure. Conversation history, documents, and logs all live in directories; the agent pulls information into context on demand with grep, tail, and Glob instead of stuffing everything in up front. Folder structure is itself a form of context engineering.
  2. Agentic search beats semantic search — by default. The official guidance: try agentic search first (let the agent grep and browse files itself), because it is more transparent and easier to maintain; only add vector retrieval when you need faster responses or more recall diversity. This echoes the trade-offs discussed in the RAG chapter.
  3. Code generation is a general-purpose capability, not just coding. Code is precise, composable, and reusable. Does the agent need an Excel report? Write a Python script to generate it — more reliable than any dedicated tool. That's exactly how Claude.ai's file-creation feature works.

The SDK's built-in agent loop follows the loop Claude Code has proven in production: gather context → take action → verify work → repeat. The built-in toolset comes straight from production rather than toy demos, including:

  • File operations: Read, Write, Edit
  • Command execution: Bash (running scripts, git, builds, tests)
  • Code navigation: Glob (find files by pattern), Grep (regex content search)
  • Web access: WebSearch, WebFetch
  • Orchestration and human interaction: Agent (spawn subagents), AskUserQuestion (ask the user)

The key point: you never write a single handler for these tools. Compare that with a hand-written loop, where you implement execution logic, serialize results, and handle errors for every tool — this is the SDK's most tangible productivity gain. For general tool-design principles, see Tools & MCP.

3. Core Capabilities ​

3.1 Subagents ​

The SDK supports subagents natively: the main agent delegates focused subtasks to specialized subagents via the Agent tool, and each subagent gets its own independent context window, toolset, and system prompt, reporting back only its conclusions when done. This solves two problems at once — parallelism (multiple subagents working on different tasks simultaneously) and context isolation (a subagent can wade through hundreds of emails and return only the relevant excerpts, without polluting the main context). This is the built-in implementation of the orchestrator-worker pattern from Multi-Agent Architecture.

Subagents can be defined two ways: programmatically via AgentDefinition in code, or — like Claude Code — as Markdown files with YAML frontmatter in the .claude/agents/ directory.

3.2 Skills ​

Agent Skills package domain expertise into directories of "instructions + scripts + resources" (including a SKILL.md with YAML frontmatter), and the agent decides when to invoke them based on the description field. The SDK loads them via the setting_sources configuration:

  • setting_sources=["project"]: loads project-level .claude/skills/, shareable via git with the whole team
  • Including "user": loads personal Skills from ~/.claude/skills/
  • Skills bundled with installed Claude Code plugins are also available

Once setting_sources is configured, the Skill tool is automatically allowed — no need to add it to allowed_tools manually. This means the large ecosystem of ready-made Skills accumulated around Claude Code is directly reusable by your SDK agents.

3.3 Hooks ​

A hook is a deterministic function invoked by the harness (not the model) at specific event points in the agent loop. Available event points include PreToolUse, PostToolUse, Stop, SessionStart, SessionEnd, UserPromptSubmit, and more. Typical uses:

  • Safety interception: in PreToolUse, inspect Bash commands and return permissionDecision: "deny" on a blacklist hit
  • Audit logging: in PostToolUse, append every file modification to an audit file
  • Result post-processing: transform or redact tool results before they reach the model

The value of hooks lies in their determinism — model behavior is probabilistic, but compliance and audit requirements are hard constraints, and that kind of logic should never be left to the model's good judgment. For more security design, see Agent Security.

3.4 MCP Integration ​

The SDK ships with an MCP client supporting two kinds of server:

  • External MCP servers: stdio subprocesses or remote HTTP/SSE — hundreds of ready-made servers for Playwright, databases, third-party APIs, and so on
  • In-process SDK MCP servers (create_sdk_mcp_server): turn Python/TS functions directly into agent-callable tools — no subprocess, no IPC overhead, type-safe — and the two kinds can be mixed

The in-process MCP server is the SDK's sweet perk: developing and debugging custom tools feels exactly like writing ordinary functions, and deployment is a single process.

3.5 The Permission System ​

The SDK ports Claude Code's permission model wholesale. The decision chain is roughly: the allowed_tools allowlist → permission_mode → the can_use_tool callback → hooks. Key points:

  • allowed_tools is an auto-approval allowlist, not an availability switch; to disable tools use disallowed_tools
  • Common permission_mode values: "default" (ask for everything), "acceptEdits" (auto-accept file edits), "bypassPermissions" (fully automatic — CI/CD territory), "plan" (produce a plan first, then execute)
  • The can_use_tool callback lets you implement arbitrarily complex approval logic in code (e.g., "write operations are only allowed under src/")

Don't run production bare

Many tutorials set permission_mode="bypassPermissions" with all tools enabled just to make the demo work. Before putting this on a server, at minimum: route Bash commands through a hook for review, restrict writes to specific directories, and gate critical operations through can_use_tool for human confirmation. An agent with computer access means the attack surface for prompt injection is the entire machine.

3.6 Session Management ​

SDK sessions persist as JSONL files on disk, and three primitives cover most scenarios:

  • resume: continue a previous session by session ID — across processes, across days
  • fork_session: branch from an existing session for "same starting point, different exploration" experiments
  • Or start fresh each time and inject the previous run's summary into the system prompt as a lightweight relay

In long pipelines, the common pattern: phase one finishes analysis and yields a session_id, and each subsequent phase calls resume to pick up where it left off, with the full chain of reasoning preserved throughout.

3.7 Context Compaction ​

When a long task approaches the context limit, the SDK automatically summarizes and compacts the message history (inherited from Claude Code's /compact capability), so the agent doesn't grind to a halt because the context blew up. This is one of the underlying guarantees behind "tasks that run for hours" — with the caveat that summarization is lossy, so key conclusions are best written to files along the way, which echoes the "filesystem as memory" philosophy.

4. Code Examples ​

Installation (Python 3.10+; Node 18+):

bash
pip install claude-agent-sdk            # Python; bundles the Claude Code CLI automatically
npm install @anthropic-ai/claude-agent-sdk   # TypeScript
export ANTHROPIC_API_KEY=your-api-key   # or use Bedrock / Vertex / Foundry

4.1 Minimal working agent (Python) ​

python
import asyncio
from claude_agent_sdk import query, ClaudeAgentOptions

async def main():
    async for message in query(
        prompt="Find every TODO comment in this codebase and summarize them",
        options=ClaudeAgentOptions(
            allowed_tools=["Read", "Glob", "Grep"],  # read-only agent: auto-approve these three tools
        ),
    ):
        if hasattr(message, "result"):
            print(message.result)

asyncio.run(main())

query() returns an async message iterator. Message types include AssistantMessage (containing TextBlock/ToolUseBlock), SystemMessage (carrying session_id when subtype == "init"), and ResultMessage (the final result, with timing and token usage). Consuming this stream gives you the full execution history, easy to pipe into logs or an observability system.

4.2 TypeScript version ​

typescript
import { query } from "@anthropic-ai/claude-agent-sdk";

for await (const message of query({
  prompt: "Review the codebase for common security issues (SQL injection, XSS, hardcoded secrets) and write security-report.md",
  options: {
    allowedTools: ["Read", "Write", "Glob", "Grep", "Bash"],
    permissionMode: "acceptEdits",  // auto-accept file edits
  },
})) {
  if ("result" in message) console.log(message.result);
}

4.3 Custom tools (in-process MCP server) ​

python
from claude_agent_sdk import (
    tool, create_sdk_mcp_server, ClaudeAgentOptions, ClaudeSDKClient,
)

# The @tool decorator turns an ordinary function into an agent-callable tool
@tool("greet", "Greet a user", {"name": str})
async def greet_user(args):
    return {"content": [{"type": "text", "text": f"Hello, {args['name']}!"}]}

server = create_sdk_mcp_server(name="my-tools", version="1.0.0", tools=[greet_user])

options = ClaudeAgentOptions(
    mcp_servers={"tools": server},
    allowed_tools=["mcp__tools__greet"],  # custom tool naming: mcp__<server>__<tool>
)

# Custom tools and hooks require ClaudeSDKClient, which supports bidirectional interaction
async with ClaudeSDKClient(options=options) as client:
    await client.query("Say hello to Alice")
    async for msg in client.receive_response():
        print(msg)

4.4 Safety interception with hooks ​

python
from claude_agent_sdk import ClaudeAgentOptions, ClaudeSDKClient, HookMatcher

async def block_dangerous_bash(input_data, tool_use_id, context):
    """PreToolUse hook: reject blacklisted Bash commands outright"""
    if input_data["tool_name"] != "Bash":
        return {}
    command = input_data["tool_input"].get("command", "")
    if "rm -rf" in command:
        return {
            "hookSpecificOutput": {
                "hookEventName": "PreToolUse",
                "permissionDecision": "deny",
                "permissionDecisionReason": "rm -rf is forbidden",
            }
        }
    return {}

options = ClaudeAgentOptions(
    allowed_tools=["Bash"],
    hooks={"PreToolUse": [HookMatcher(matcher="Bash", hooks=[block_dangerous_bash])]},
)

4.5 Subagents and session resumption ​

python
from claude_agent_sdk import query, ClaudeAgentOptions, AgentDefinition, SystemMessage

options = ClaudeAgentOptions(
    allowed_tools=["Read", "Glob", "Grep", "Agent"],  # the Agent tool must be allowed to spawn subagents
    agents={
        "code-reviewer": AgentDefinition(
            description="Senior code reviewer focused on quality and security",
            prompt="Analyze code quality, flag potential issues, and report only your conclusions.",
            tools=["Read", "Glob", "Grep"],  # the subagent's own toolset
        )
    },
)

session_id = None
async for message in query(prompt="Have the code-reviewer review this codebase", options=options):
    if isinstance(message, SystemMessage) and message.subtype == "init":
        session_id = message.data["session_id"]

# Continue on the same session for the follow-up task, keeping full context
async for message in query(
    prompt="Fix the three most severe issues from the review findings",
    options=ClaudeAgentOptions(resume=session_id, allowed_tools=["Read", "Edit", "Bash"]),
):
    pass

An easy trap to fall into

By default the SDK inherits Claude Code's identity — unless you override it, your agent will introduce itself as "Claude Code." Two fixes: override the identity with system_prompt, or go fully isolated with setting_sources=[], which loads no user/project configuration at all (no .claude/, no CLAUDE.md, no Skills). For production deployments, set both explicitly to avoid behavior drifting between local debugging and production.

5. The Relationship to Claude Code: One Harness, Two Forms ​

Here's how to think about the division of labor:

          The same agent harness (loop / tools / permissions / hooks)
          ┌──────────────────┴──────────────────┐
   Claude Code (product form)          Claude Agent SDK (library form)
   Interactive coding agent in the     A library embedded in your own app
   terminal                            Your code sits in the driver's seat
   You sit in the driver's seat        Configured via query() parameters
   CLAUDE.md / .claude/ config

Practical corollaries:

  • Shared ecosystem. .claude/skills/, .claude/agents/, CLAUDE.md, and MCP configurations work on both sides. A Skill tuned inside Claude Code works directly in an SDK agent with setting_sources=["project"].
  • A debugging path. When an SDK agent behaves strangely, hand the same task to Claude Code in the terminal to reproduce it and rule out harness-level issues.
  • Very fast iteration. The SDK releases on Claude Code's cadence; the TypeScript package's npm latest reached 0.3.x (0.3.231) in August 2026 while Python sits at 0.1.x — the API surface can still shift in small breaking ways, so check the CHANGELOG before upgrading and pin your versions.

6. Where It Fits ​

The official blog names typical directions: financial analysis agents (read portfolios, call external APIs, run calculations), personal assistants (book travel, manage calendars, track context across apps), support agents (handle highly ambiguous tickets, escalate to humans when needed), and deep research agents (search across large document sets, cross-reference, generate reports). Combined with community practice, its sweet spot comes down to three statements:

  1. Long-running tasks. Auto-compaction + session persistence + subagent context isolation let it run tasks measured in hours — something most hand-written loop frameworks struggle to achieve.
  2. Coding and engineering automation. It inherits everything from Claude Code: code review, refactoring, bug fixing, running tests, git operations, auto-remediation agents in CI.
  3. Research/ops-style "digital work" agents. Any task reducible to "manipulate files, run commands, call APIs" — report generation, log analysis, SRE inspection bots, content pipelines — fits this "give the model a computer" playbook.

Conversely, it is not the best choice for: pure chat / structured extraction (the Messages API is lighter); graph-style orchestration where you need fine control over every state transition (see LangGraph); or multi-model setups and non-Claude models (see next section).

7. Pros, Cons, and Vendor Lock-in ​

ProsCons
The harness is battle-tested by Claude Code at production scale, not a paper designClaude models only — the model layer is not swappable
Built-in tools work out of the box, saving large amounts of glue codeFast iteration; the 0.x stage still has occasional small API breaks
Production concerns (permissions, hooks, compaction) are built inHarness behavior is a black box; fine-grained control trails a hand-written loop
Seamless reuse of the Claude Code ecosystem (Skills, MCP, plugins)Bundles a CLI binary — deployment size and platform compatibility need consideration
Both Python and TS; sessions persist as auditable, replayable JSONLLimited room for deep customization (swapping the loop or context strategy)

On vendor lock-in, one thing deserves to be said precisely: what you're actually locked into is not "the API interface" (that layer is not hard to swap) but the engineering assets accumulated around this harness — CLAUDE.md files, Skills, subagent definitions, hooks, permission policies. These assets are deeply tied to Claude Code's conventions and would largely need rewriting to move to another framework. So the selection question is: do you buy the "gather context → take action → verify" loop and the "give the model a computer" philosophy? If yes, the lock-in is the dividend of deep integration. If not, choose a hand-written loop or a more neutral orchestration framework from day one (see the framework selection overview and the OpenAI Agents SDK comparison).

The selection call

If your task is "get an agent to finish real work in a real environment" rather than "study agent architecture itself," the Claude Agent SDK is one of the default answers in 2026. Only two reasons remain for hand-writing the loop: you need model neutrality, or you need a control granularity the harness can't provide. To build loop intuition with your own hands, see Build an Agent from Scratch.

References ​