Skip to content

Framework Selection Overview

At a glance The 2026 agent framework landscape — from bare SDKs and the OpenAI Agents SDK to the Claude Agent SDK, LangGraph, CrewAI, Microsoft Agent Framework, and low-code platforms — with a comparison table, hidden-cost analysis, and a decision tree you can actually follow.

This page contains time-sensitive content; data is current as of 2026-08. Job listings, pricing, and product features may have changed — verify against the original sources before citing.

Framework Selection Overview ​

"Which framework should I use to build agents?" is the most frequently asked — and the worst answered — question in this field. The problem with most answers is that they hand you a name without first asking three questions: is your task a fixed workflow or open-ended exploration, what stack does your team know, and how much abstraction can you afford? This page doesn't offer "the standard answer." Instead, it gives you a spectrum map, a comparison table verified as of mid-2026, an honest discussion of whether you need a framework at all, and a decision tree you can actually walk through. Deep dives on each framework live in the sub-pages; low-code platforms are covered under the Product Case Studies section.

One judgment runs through the whole page: frameworks solve orchestration problems, not intelligence problems. Model capability, prompts, tools, and context determine whether your agent is smart (see Agent Loop and Context Engineering); a framework only determines how those calls are organized, how state is stored, and how failures recover. Don't expect a framework to compensate for a weak model.

1. The Framework Spectrum: From Bare SDKs to Low-Code ​

Agent development tooling in 2026 forms a continuous spectrum: moving left to right, abstraction increases, control decreases, and time-to-first-agent shrinks:

High control ◄──────────────────────────────────────────────► Fast to start
Less abstraction                                                More abstraction

┌──────────┐  ┌────────────┐  ┌────────────┐  ┌────────────┐  ┌────────────┐
│ Bare SDK │  │ Lightweight│  │ Graph/FSM  │  │ Role-based │  │ Low-code   │
│          │  │ SDKs       │  │ frameworks │  │ frameworks │  │ platforms  │
│ OpenAI / │→ │ OpenAI     │→ │ LangGraph  │→ │ CrewAI     │→ │ Dify       │
│ Anthropic│  │ Agents SDK │  │ Google ADK │  │ Microsoft  │  │ Coze       │
│ official │  │ Claude     │  │ MAF        │  │ Agent      │  │ n8n        │
│ libs     │  │ Agent SDK  │  │            │  │ Framework  │  │            │
└──────────┘  └────────────┘  └────────────┘  └────────────┘  └────────────┘
  ~100 lines    A few            Explicit state,  "Role +          Drag-and-drop
  hand-rolled   primitives:      branching,       process"         canvas, no
  loop          handoffs/hooks   resumable runs   metaphor         code

This spectrum is not an evolution ladder — the right side is not more "advanced" than the left. These tools serve different scenarios and different people:

  • Bare SDK (no framework): use the official OpenAI / Anthropic libraries directly and hand-write a while loop plus tool calling. A working agent loop core comes in under 100 lines of code. Good for learning, good for simple tasks, and — surprisingly — good for many production scenarios (more on that below).
  • Lightweight SDKs: represented by the OpenAI Agents SDK (released March 2025) and the Claude Agent SDK (renamed from the Claude Code SDK in 2025). They give you a small set of production-proven primitives — handoff, guardrail, hook, session — without taking over your control flow. In essence, vendors packaging the loops they run agents with internally into a library.
  • Graph / state-machine frameworks: LangGraph (1.0 released October 2025) is the flagship; Google ADK and Microsoft Agent Framework sit in the same tier. Agent execution is modeled as an explicit state graph: nodes are steps, edges are transition conditions, and state flows between nodes and can be persisted. What you get in return is durable execution (resume from breakpoints), human-in-the-loop interrupts, and auditable execution traces.
  • Role-play orchestration frameworks: CrewAI and (now maintenance-mode) AutoGen. Multi-agent systems organized around the metaphor "give every agent a role and let them collaborate." Fastest to get started, but the price of the metaphor is that during debugging you don't know where the conversation will wander.
  • Low-code platforms: Dify, Coze, n8n. Drag-and-drop canvases replace code, aimed at business users or rapid validation. They sit at the far end of the "framework" spectrum; strictly speaking, they are platforms rather than frameworks.

An easy-to-miss fact

Lightweight SDKs and graph frameworks are not two mutually exclusive rungs on a ladder. On top of LangGraph, the LangChain team built DeepAgents (a batteries-included harness with planning, sub-agents, and filesystem memory), and Anthropic's Claude Agent SDK is itself the Claude Code runtime packaged as a library. The 2026 trend is "vendors hand you a proven harness," not "assemble everything from graph primitives."

2. The Big Comparison of Mainstream Frameworks ​

The table below was verified as of mid-2026 (primary sources are listed in the references at the end). Version status changes fast; re-check the official repositories before citing any specific number.

DimensionBare SDKOpenAI Agents SDKClaude Agent SDKLangGraphCrewAIMicrosoft Agent Framework
PositioningHand-rolled loopLightweight multi-agent primitivesClaude Code runtime as a libraryGraph orchestration + state managementRole-based multi-agentAutoGen/SK successor
Abstraction levelNoneLowLow-midMid-highMidMid-high
LanguagesPython or TSPython + TS (separate package @openai/agents)Python 3.10+ / TSPython / JS-TSPythonPython / .NET (C#)
Core primitivesmessages + toolsAgent, handoff, guardrail, sessionhooks, in-process MCP, permission allowlist, sub-agentsNodes/edges/state, interrupt, checkpointerCrew (role collaboration) + Flow (event-driven pipeline)Graph workflows, middleware, declarative YAML definitions
State managementDIYSessions manage conversation historySession + filesystem stateCheckpointer persistence, durable executionIn-memory/external storage + checkpointSession state + telemetry
Human-in-the-loopDIYCompose it yourselfPermission allowlist + AskUser hooksFirst-class: interrupt() pause and resumeApproval nodes can be plugged into FlowsApprovals and interrupts supported
Model bindingBound to whichever API you callNominal neutrality; 100+ models via LiteLLMTightly bound to ClaudeModel-agnosticModel-agnosticModel-agnostic, deep Azure integration
EcosystemNoneOpenAI ecosystem, built-in tracingAnthropic ecosystem + MCPDeepest LangChain integration catalogStandalone ecosystem (not built on LangChain)Azure / Foundry
Version status (mid-2026)—Still 0.x (~0.18 as of July 2026), iterating fastStable iteration, evolving with Claude Code1.0 (GA 2025-10, API-stability pledge until 2.0)1.0 GA (2025-10)1.0 GA (2026-04)
Production fitAmple for simple tasksFine for quick launches; watch API churn long termThe pragmatic choice for Claude-family agentsA de facto standard for complex long-running tasksPrototypes to small/mid-scale productionThe default for Microsoft/.NET teams

Beyond the table, a few facts worth spelling out:

LangGraph's position. After 1.0, it doubled down on production concerns: durable execution (exact recovery from failure points), human-in-the-loop interrupts, and short- plus long-term memory. Official materials name Uber, LinkedIn, Klarna, and JP Morgan as production users (October 2025 1.0 announcement). The companion LangGraph Platform went GA in May 2025. It is the most-cited default for "complex, stateful, long-horizon" workflows — note the qualifier; it is not the default for "all agents."

The OpenAI Agents SDK's 0.x status. Its primitive design (agent / handoff / guardrail / session) is clean and elegant, but as of July 2026 it is still 0.x with frequent releases, which means the API can change. Despite the OpenAI name, it supports hundreds of models through LiteLLM and similar integrations. OpenAI has announced a phased retirement of the Assistants API starting mid-2026 in favor of the Responses API; the Agents SDK is the official vehicle on that road.

What makes the Claude Agent SDK unique. It is not "yet another orchestration framework" — it exposes the runtime that powers Claude Code (file read/write, Bash, sub-agents, permission control, hooks, in-process MCP servers) directly as a library. If your agent fundamentally needs to "look around and get work done inside the filesystem and shell," this harness has already been polished by enormous real-world Claude Code usage, and rewriting it yourself will almost certainly turn out worse.

AutoGen is done — but know exactly what that means. AutoGen entered maintenance mode in late 2025: its README now states plainly that it no longer accepts new features and is community-maintained, with only bug fixes and security patches. Microsoft merged AutoGen with Semantic Kernel into the Microsoft Agent Framework, which went GA in April 2026 and is the only first-tier option in the .NET ecosystem. Existing AutoGen code keeps running; new projects should not pick it. This site keeps an AutoGen page because its "conversation as computation" idea and multi-agent patterns are still worth studying — as history and design material, not as a selection recommendation.

Second-tier names worth knowing: Google ADK (Gemini ecosystem, now at 2.x), Pydantic AI (type-safe Python with FastAPI ergonomics), Mastra and the Vercel AI SDK (the top two in the TypeScript camp), AWS's Strands Agents, and HuggingFace's smolagents (a minimal CodeAgent with a ~1,000-line core). When shortlisting, match them by language and cloud-vendor affinity.

The three most common head-to-head matchups ​

In practice, selection is rarely an open audition; it's usually a short-list showdown between two candidates. The three matchups engineering teams debated most in 2026:

  • LangGraph vs OpenAI Agents SDK: control versus simplicity. The Agents SDK's four primitives (agent / handoff / guardrail / session) can be picked up in half a day and are great for getting a multi-agent prototype running quickly; LangGraph asks you to think through a state graph first, but in exchange you get auditable, replayable, interruptible execution. "Prototype with the former, go to production with the latter" is a common path — but be aware the codebases are incompatible, so that "migrate later" is a real rewrite, not a refactor.
  • LangGraph vs Pydantic AI: an intra-Python-team matchup. Pydantic AI uses type annotations to describe agent inputs, tool signatures, and outputs; dependency injection, result validation, and OpenTelemetry instrumentation are handled by the framework, and the code footprint is an order of magnitude smaller. It fits single agents and small systems embedded in Python services. LangGraph wins on explicit orchestration: multiple roles, long horizons, checkpoints, approval interrupts. The one-line rule of thumb: if your complexity lives in "the output quality of a single agent," pick Pydantic AI; if it lives in "how multiple execution units coordinate and recover," pick LangGraph.
  • Mastra vs Vercel AI SDK (the TypeScript camp): Mastra is the all-in-one — agents, graph workflows, RAG, memory, and evals bundled into a TS-first package. The Vercel AI SDK is "grow agent capabilities on your existing Next.js/Node stack"; AI SDK 7 provides ToolLoopAgent, HarnessAgent, and WorkflowAgent abstractions. If your entire product lives in the JS/TS world, the former is more cohesive; if you already have AI SDK code in production, the latter integrates more smoothly.

The shared lesson from all three matchups: cut by language and cloud-vendor affinity first, then benchmark the 2-3 surviving candidates with real tests, rather than reading ten comparison posts.

3. Do You Even Need a Framework? An Honest Discussion ​

Anthropic's advice: write it bare first, then consider a framework ​

Anthropic's "Building Effective Agents," published in December 2024, remains the most-cited piece of engineering advice in this field. Its two core claims:

  1. After working with dozens of teams across industries, they found that the most successful implementations used no complex frameworks — just simple, composable patterns.
  2. Strictly distinguish workflows (orchestration where the code path is fixed in advance) from agents (systems where the model decides the path), and advise: find the simplest solution first; introduce autonomy only when fixed steps can't solve the problem.

This is not a "frameworks are worthless" argument — it's a question of sequencing. Durable execution, multi-agent delegation, human-in-the-loop, tracing: each one only pays for itself once "you actually hit that problem." Before that, it's just baggage you carry before ever stepping on the mine.

A minimal bare agent loop looks like this (Anthropic SDK, Python):

python
import anthropic

client = anthropic.Anthropic()

def run_agent(user_task: str, tools: list, max_turns: int = 20):
    """Minimal agent loop: the model calls tools autonomously until it produces a final answer."""
    messages = [{"role": "user", "content": user_task}]

    for _ in range(max_turns):
        resp = client.messages.create(
            model="claude-sonnet-4-5",
            max_tokens=4096,
            tools=tools,                # [{name, description, input_schema}, ...]
            messages=messages,
        )
        messages.append({"role": "assistant", "content": resp.content})

        # No tool call = the model considers the task done; exit the loop
        if resp.stop_reason != "tool_use":
            return resp

        # Execute the requested tools and feed the results back as tool_result
        results = []
        for block in resp.content:
            if block.type == "tool_use":
                output = execute_tool(block.name, block.input)  # your tool dispatcher
                results.append({
                    "type": "tool_result",
                    "tool_use_id": block.id,
                    "content": str(output),
                })
        messages.append({"role": "user", "content": results})

    raise RuntimeError("Max turns exceeded; agent failed to converge")

That is the entire core of the "bare SDK" tier. Read it once and you'll see: a loop, tool dispatch, message appending — no black magic. Jumping to LangGraph without understanding this code is building the roof before the foundation, which is why this site puts Build Your Own Agent as the first stop in the practice track.

The hidden costs of frameworks ​

Four bills the framework marketing page won't show you:

  • Learning cost up front. LangGraph's concept system — nodes, edges, reducers, checkpointers — takes weeks to learn properly, and your first agent probably won't use half of it. Lightweight libraries like the OpenAI Agents SDK have a gentle learning curve, but during 0.x you have to keep tracking breaking changes.
  • Debugging cost later. The heavier the framework, the longer the distance between "why did the model decide that" and "which framework layer changed the behavior." A trace you'd see with a single print(messages) in a bare loop may require digging through three layers of callbacks in a heavy framework. Observability tooling can recover some of it (see Observability), but not as well as never losing it in the first place.
  • Version coupling cost. Agent frameworks are among the most violently iterating software categories of 2025-2026. The LangChain ecosystem's early reputation — "every minor release breaks something" — wasn't unearned; AutoGen went from viral to maintenance mode in two years. Choosing a framework means joining a team you'll be following through upgrades for the next year.
  • Talent and migration cost. Framework-specific idioms (LangGraph's StateGraph, CrewAI's YAML-style role definitions) are not transferable skills. When your team switches frameworks, that experience resets to zero.

Leaky abstractions: the original sin of agent frameworks ​

Joel Spolsky's Law of Leaky Abstractions plays out in agent frameworks almost verbatim: frameworks try to abstract "an LLM call" into nodes, agents, and Crews, but LLM non-determinism leaks through every abstraction straight to your door. The model emits schema-invalid JSON inside a node, half the context disappears after a handoff, two roles in a Crew get stuck in a mutual-flattery loop — none of these problems can even be described in the framework's vocabulary, and you end up back at the level of prompts, message lists, and tool results to fix them.

When a framework really is over-engineering

When all of the following hold, a heavy framework is over-engineering in most cases: single agent, fewer than 10 tools, tasks that finish within minutes, cheap failure-and-rerun, no mid-run human intervention needed. A while loop with retry is all you need. The counter-case is just as clear: tasks that run for hours, cross-session recovery, approval gates, multi-agent division of labor — those are exactly where durable execution and state machines earn their keep, and rolling your own is more expensive.

So what does a framework actually buy you? ​

Four things, ranked by value:

  1. State and recovery: checkpoints, durable execution, resuming long tasks from a breakpoint. Building this yourself is far more painful than it sounds — it's the hardest-selling point of the LangGraph tier.
  2. Human-in-the-loop infrastructure: interrupts, approvals, resuming after a pause. Agents that touch money, production, or outbound messaging can't skip it (see Human-in-the-Loop).
  3. Multi-agent primitives: handoffs, sub-agents with isolated context, group coordination. When you genuinely need multi-agent (first check Multi-Agent Architecture to confirm you're not over-engineering), primitives are steadier than hand-rolling.
  4. Tracing and evaluation hooks: a production necessity — but note that standalone observability tools like Langfuse and LangSmith can be attached without any orchestration framework, so they don't constitute a lock-in reason to pick one.

4. The Decision Tree ​

Compress "team background × task complexity × autonomy needs" into a single tree. Start at the root and follow the branches; whichever leaf you land on is the path to try first:

Start: I want to build an Agent
│
├─ Q1: Is the task a fixed process (steps enumerable in advance)?
│   │
│   ├─ Yes ──► Don't build an Agent — build a workflow
│   │         ├─ Can you code? ──► Bare SDK + plain function orchestration (prompt chaining / routing)
│   │         └─ No code? ─► Low-code platform: Dify / Coze / n8n
│   │
│   └─ No (the path must be decided by the model) ──► Q2
│
├─ Q2: Is this your first agent prototype / a learning project?
│   │
│   └─ Yes ──► Bare SDK, hand-write the agent loop (~100 lines)
│             The goal isn't saving money — it's building intuition for the loop
│             └─ When you hit a real bottleneck, come back to Q3
│
├─ Q3: What are the task's runtime characteristics?
│   │
│   ├─ Short tasks, cheap to re-run, single agent
│   │   ├─ Mostly OpenAI models ─► OpenAI Agents SDK
│   │   ├─ Mostly Claude and needs file/shell work ─► Claude Agent SDK
│   │   └─ Want vendor neutrality ─► Bare SDK or Pydantic AI
│   │
│   ├─ Long tasks / needs resumability / needs approvals / complex branching
│   │   ├─ Python/TS team ─► LangGraph
│   │   ├─ .NET / Azure team ─► Microsoft Agent Framework
│   │   └─ Gemini / GCP team ─► Google ADK
│   │
│   └─ Clearly multi-role collaboration (research → writing → review style)
│       ├─ Quick idea validation ─► CrewAI (Crews + Flows)
│       └─ Production with complex flows ─► Model it explicitly in LangGraph;
│                                          roles stay at the prompt layer
│
└─ Q4: TypeScript full-stack team?
    └─ Yes ──► Mastra or Vercel AI SDK first, LangGraph.js as fallback

Two usage notes:

  • This tree gives you the "first candidate," not a lifetime commitment. The reliable approach (suggested by Langfuse): implement the same minimal task with your top two candidates, pipe both sides' traces into the same observability platform, compare cost, latency, and failure modes, then pick the framework whose traces you'd rather be debugging a year from now.
  • There is no "switching cost" branch in the tree, because the way to lower switching cost isn't in the tree: keep prompts, tool definitions, and eval sets separate from the orchestration code. Those three assets are what actually travel between frameworks.
  • One closing meta-tip: framework hype has a half-life measured in months, but the variables this tree is built on — task structure, autonomy needs, recovery cost, team stack — have a half-life measured in years. Pull selection discussions back to these four variables and you won't be whipsawed by every quarterly framework launch.

5. Guide to the Deep-Dive Pages ​

This page handles selection only; installation, core APIs, complete examples, and gotchas for each framework live on their dedicated pages:

  • LangGraph deep dive: the StateGraph mental model, checkpointers and durable execution, human-in-the-loop via interrupt, subgraphs and multi-agent.
  • OpenAI Agents SDK deep dive: the four core primitives — agent / handoff / guardrail / session — tracing integration, and its position in the Responses API era.
  • Claude Agent SDK deep dive: reusing the Claude Code runtime, hooks that intercept the agent loop, in-process MCP servers, permission allowlists.
  • CrewAI deep dive: the Crew role-collaboration model, event-driven Flows after 1.0, when CrewAI is enough — and when to switch.
  • AutoGen deep dive: its maintenance-mode status, the migration path to Microsoft Agent Framework, and a retrospective on the "conversation as computation" design philosophy.

If you haven't read the core components section yet, start with Tools & MCP — the tool integration layers of all these frameworks have almost entirely converged on MCP, which affects your long-term architecture more than which framework you pick.

6. Low-Code Platforms: The Other End of the Spectrum ​

Strictly speaking, Dify, Coze, and n8n are not "frameworks" but platforms: what you get is not an abstraction inside your codebase but an entire environment with UI, hosting, permissions, and app distribution. The rough landscape as of mid-2026:

PlatformPositioningOpen sourceWho it's for
DifyLLM app development platform (Workflow + RAG + Agent in one)Yes; ~142k GitHub stars (2026-05)Teams with engineering resources that want to self-host
CozeNo-code bot/agent builder, by ByteDanceNo (mostly SaaS)Ops and business users starting from zero code
n8nAutomation/integration platform that grew AI capabilities; 400+ connectorsYes (fair-code); ~188k GitHub starsCases where an agent's value is mostly connecting external systems

The differences between the three boil down to their starting points: Dify started from LLM apps and added workflows, n8n started from automation connectors and added AI nodes, and Coze started from "let people who can't code build bots" and built ecosystem distribution around it.

A common pragmatic combination is a Dify + n8n dual stack: the former owns LLM application logic, the latter owns system integration. But be clear-eyed about the ceiling: when agent behavior needs fine tuning (custom loops, fine-grained context control, unusual tool protocols), you will hit a wall — and there is no code behind that wall to fix. Both have dedicated pages under Product Case Studies: Coze case study and Dify case study.

One line for job seekers

"Knows LangGraph" appears in job descriptions far more often than any other framework, but the question that actually separates candidates in interviews is "how would you do it without a framework, and what does the framework buy you?" Write it bare first, master one graph framework, and know the rest at "read the docs, know the trade-offs" depth — that's the highest-ROI combination. See Job Landscape and Knowledge Map for role details.

References ​