Skip to content

LangGraph

At a glance LangGraph is the agent orchestration layer of the LangChain ecosystem — a graph state machine that gives you explicit control over every step an agent takes. Based on the v1.x docs as of August 2026, this page covers State/Node/Edge, the Checkpointer, human-in-the-loop with interrupt, deployment with LangSmith Deployment, how the ecosystem divides responsibilities — and an honest take on the complexity controversy.

This page contains time-sensitive content; data is current as of 2026-08. Job listings, pricing, and product features may have changed — verify against the original sources before citing.

LangGraph ​

LangGraph is currently the de facto standard agent orchestration framework in the Python ecosystem. It is not "yet another agent library" but a runtime that models agent execution as a graph state machine: you break the flow into nodes and edges, state flows between nodes, and the framework handles the "dirty work" — persistence, resuming from checkpoints, streaming output, and human-in-the-loop.

This page is based on the official LangGraph v1.x documentation as of August 2026. LangChain and LangGraph both shipped v1.0 in October 2025 (official release notes), and the core graph API has been frozen stable ever since. A large share of the 2024-era tutorials online still use AgentExecutor and create_react_agent patterns that are now obsolete — treat them with suspicion when you see them.

1. Positioning: the Agent Orchestration Layer of the LangChain Ecosystem ​

Start with the version timeline to understand why it looks the way it does today:

TimeEvent
2024-01LangGraph first appeared in the LangChain v0.1 announcement, positioned as "building language agents with graphs"
2024-06v0.1 released, alongside LangGraph Cloud
2024-08v0.2: checkpointer split into its own library; stronger session memory, checkpoint recovery, human-in-the-loop
2025-02v0.3 series: control-flow primitives like Command and Send introduced
2025-05LangGraph Platform GA (the commercial deployment product)
2025-10v1.0 GA; LangGraph Platform renamed LangSmith Deployment; LangChain v1's create_agent rebuilt on top of LangGraph
2026v1.x keeps polishing stability and type safety, with no breaking changes to the core API

After v1.0, the official division of labor among the three products is very clear:

┌─────────────────────────────────────────────────────┐
│ LangChain  (v1)                                     │
│  = Agent API layer: create_agent + middleware       │
│  The standard agent loop for quick starts           │
├─────────────────────────────────────────────────────┤
│ LangGraph                                           │
│  = Agent Runtime: StateGraph / Node / Edge          │
│  Persistence, streaming, HITL, custom orchestration │
├─────────────────────────────────────────────────────┤
│ LangSmith                                           │
│  = Engineering platform: tracing / eval / Deployment│
└─────────────────────────────────────────────────────┘

In one sentence: LangChain is the integration and abstraction layer, LangGraph is the execution and orchestration layer, and LangSmith is the observability and deployment layer. create_agent runs on the LangGraph runtime under the hood — you start fast with LangChain and drop down to LangGraph when you need fine-grained control; the two are not competitors. For how this layering maps onto the general anatomy of an agent, see the Anatomy and Agent Loop pages.

2. Core Concepts ​

LangGraph's entire API surface collapses into four primitives: State, Node, and Edge, plus the Checkpointer that manages state.

2.1 State and Reducers ​

State is a TypedDict (or Pydantic model, or dataclass) representing the full snapshot of the graph at any moment. The key concept is the reducer: each state field can declare a merge function that decides how a node's partial update is written into the global state.

python
from typing import Annotated
from typing_extensions import TypedDict
from langgraph.graph import MessagesState          # built-in messages field
from langgraph.graph.message import add_messages   # merges by message ID, deduplicates

class State(MessagesState):
    # messages: Annotated[list, add_messages] is provided by MessagesState
    documents: list[str]          # no reducer → default is full-value overwrite
    retry_count: int
  • Without a declared reducer, a node's return value overwrites the field directly;
  • add_messages is the standard reducer for message fields: new messages are appended, same-ID messages are overwritten, and dicts are automatically deserialized into Message objects;
  • The graph executes in Pregel-style super-steps: multiple nodes run in parallel within one super-step, and the next step begins only after all of them finish.

2.2 Nodes, Edges, and Conditional Routing ​

A node is just a plain Python function: it receives state, does work, and returns a partial update. An edge decides "where to go next." There are three ways to route:

python
builder = StateGraph(State)
builder.add_node("agent", call_model)
builder.add_node("tools", ToolNode(tools))   # prebuilt tool-execution node

builder.add_edge(START, "agent")              # fixed entry
builder.add_conditional_edges("agent", route) # dynamic routing
builder.add_edge("tools", "agent")            # fixed edge
  • add_conditional_edges("agent", route_fn): route_fn returns the name of the next node (or a list of nodes — returning multiple triggers a parallel fan-out).
  • Command: when a node wants to update state and decide the destination at the same time, just return Command(update={...}, goto="next_node") instead of splitting into a node plus a routing function. Note that you must annotate the return type as Command[Literal["next_node"]], otherwise the graph won't render.
  • Send: for map-reduce scenarios — dynamically fan out any number of parallel branches from a routing function, each carrying its own sub-state.
  • Subgraphs: a compiled graph can be used directly as a node in another graph, with shared fields flowing through automatically; inside a subgraph you can use Command(goto=..., graph=Command.PARENT) to jump back to a parent-graph node — this is the official pattern for multi-agent handoffs.

2.3 Checkpointer Persistence ​

This is one of LangGraph's core advantages over "hand-writing a while loop." Attach a checkpointer at compile time, and every completed super-step is automatically persisted:

python
from langgraph.checkpoint.memory import InMemorySaver          # for development
# for production: langgraph-checkpoint-sqlite / langgraph-checkpoint-postgres
checkpointer = InMemorySaver()
graph = builder.compile(checkpointer=checkpointer)

config = {"configurable": {"thread_id": "user-42"}}
graph.invoke({"messages": [("user", "Hello")]}, config)

The thread_id is your persistence cursor: invoking again with the same ID resumes from the last checkpoint, surviving process crashes and machine restarts. Around it you also get time travel (branch off from any historical checkpoint) and breakpoint debugging (interrupt_before / interrupt_after).

2.4 interrupt: Human-in-the-Loop ​

interrupt() is the recommended human-in-the-loop primitive in v1 (the older static breakpoints still exist, but dynamic interrupts are far more flexible):

python
from langgraph.types import interrupt, Command

def human_review(state: State):
    # Pause the graph and throw the payload to the caller until someone resumes
    answer = interrupt({"question": "Approve this transfer?", "amount": state["amount"]})
    return {"approved": answer == "yes"}

# First invocation: returns {"__interrupt__": [...]} and the graph suspends
result = graph.invoke(input, config)
# After a human decides, resume: the return value of interrupt() is "yes"
graph.invoke(Command(resume="yes"), config)

Three interrupt gotchas (the official "rules of interrupts")

  • Never wrap interrupt in try/except: it suspends by raising a special exception; if you catch it, the interrupt silently stops working.
  • The node re-runs from the top on resume: code before the interrupt executes again, so side effects (database writes, API calls) must either be idempotent or be moved after the interrupt into a separate node.
  • Multiple interrupts in one node must have a deterministic order: resume values are matched by index; conditionally skipping an interrupt causes a mismatch.

Combined with the checkpointer, this mechanism turns scenarios like "an approval flow suspended for three days until a human clicks confirm" into a few lines of code. See the pattern summary on the Human-in-the-Loop page.

3. A Complete v1 Code Example ​

Below is a runnable customer-support agent: tool calling + conditional routing + persistence + human approval for large refunds. All APIs were verified against the official docs of 2026-08 (langgraph v1.x / langchain v1.x).

python
# pip install -U langgraph langchain langchain-anthropic
from typing import Annotated
from typing_extensions import TypedDict

from langchain.chat_models import init_chat_model
from langchain_core.tools import tool
from langgraph.graph import StateGraph, MessagesState, START, END
from langgraph.graph.message import add_messages
from langgraph.prebuilt import ToolNode
from langgraph.checkpoint.memory import InMemorySaver
from langgraph.types import interrupt, Command

# ---------- 1. Tools ----------
@tool
def query_order(order_id: str) -> str:
    """Look up order status."""
    return f"Order {order_id}: shipped, expected delivery tomorrow"

@tool
def refund(order_id: str, amount: float) -> str:
    """Initiate a refund. Large refunds trigger human approval first."""
    if amount > 500:
        # Interrupt inside the tool: approval logic travels with the tool,
        # so it works in any graph that reuses it
        decision = interrupt({
            "action": "refund",
            "order_id": order_id,
            "amount": amount,
        })
        if decision != "approve":
            return f"Refund rejected (human review decision: {decision})"
    return f"Refunded ¥{amount} for order {order_id}"

tools = [query_order, refund]

# ---------- 2. State ----------
class State(MessagesState):
    pass  # this example only needs the built-in messages field (add_messages reducer)

# ---------- 3. Nodes ----------
llm = init_chat_model("anthropic:claude-sonnet-4-6").bind_tools(tools)

def call_model(state: State):
    response = llm.invoke(state["messages"])
    return {"messages": [response]}

def should_continue(state: State) -> str:
    """Conditional routing: go to tools if the model wants another call, else end."""
    last = state["messages"][-1]
    return "tools" if last.tool_calls else END

# ---------- 4. Build the graph ----------
builder = StateGraph(State)
builder.add_node("agent", call_model)
builder.add_node("tools", ToolNode(tools))
builder.add_edge(START, "agent")
builder.add_conditional_edges("agent", should_continue)
builder.add_edge("tools", "agent")          # tool results go back to the model, closing the agent loop

graph = builder.compile(checkpointer=InMemorySaver())

# ---------- 5. Run (including human approval and resume) ----------
config = {"configurable": {"thread_id": "thread-001"}}

result = graph.invoke(
    {"messages": [("user", "Please refund order A123 for 800 yuan")]},
    config,
)

# The large-refund interrupt fired: result carries the __interrupt__ key
if "__interrupt__" in result:
    print("Waiting for approval:", result["__interrupt__"][0].value)
    # Simulate an approver clicking "Approve," then resume execution
    final = graph.invoke(Command(resume="approve"), config)
    print(final["messages"][-1].content)

When you don't need to build the graph yourself

If all you want is the standard "model + tools + loop" ReAct agent, use LangChain's create_agent after v1 — it runs on LangGraph underneath, but you write zero graph code:

python
from langchain.agents import create_agent
agent = create_agent(
    model="anthropic:claude-sonnet-4-6",
    tools=[query_order, refund],
    system_prompt="You are a customer support assistant. Treat large refunds with caution.",
)
agent.invoke({"messages": [{"role": "user", "content": "Look up order A123"}]})

It also ships a middleware system (PII scrubbing, conversation compaction, HumanInTheLoopMiddleware, and more) that covers most "standard agent plus a bit of customization" needs. Only when routing logic, state structure, or multi-agent collaboration outgrows the template is it worth dropping down to a hand-written StateGraph. Note that langgraph.prebuilt.create_react_agent is deprecated — stop copying it from old tutorials.

The hand-written graph version is essentially the Agent Loop made explicit: every step's inputs and outputs, state transitions, and interruption points are all visible and persistable — which is exactly what you get in exchange for giving up the black box and choosing LangGraph.

4. LangSmith Deployment (formerly LangGraph Platform) and Studio ​

Getting the graph running locally is only step one; deploying a long-running, stateful agent as a service is where the real engineering effort lives. LangChain's commercial deployment product evolved through LangGraph Cloud → LangGraph Platform → LangSmith Deployment (renamed October 2025), with the core components unchanged:

  • LangGraph Server: a service runtime purpose-built for agents — horizontally scalable task queues, long-running support (this is not an ordinary web framework designed for second-scale HTTP requests), cross-session persistence, streaming/background runs, concurrency control (how to handle a user firing off multiple messages), plus cron and webhooks.
  • Studio: a visual debugger that shows the graph structure, executes step by step, and lets you modify mid-run state and continue — extremely valuable during development.
  • CLI + Python/JS SDK: for running the service locally and managing deployments.

For local development, a langgraph.json declares dependencies and graph entry points:

json
{
  "dependencies": ["."],
  "graphs": {
    "support_agent": "./agent.py:graph"
  },
  "env": ".env"
}

Then langgraph dev starts a local server with Studio built in (connect to the local port through the LangSmith web UI), with breakpoints, state inspection, and streaming output all available.

Deployment comes in four tiers (official announcement):

PlanDescriptionFits
Self-Hosted LiteFree self-hosting, capped at 1M node executionsSmall teams validating, internal tools
Cloud SaaSFully managed, bundled with LangSmith plansMost teams
BYOCDeployed in your VPC (AWS only for now), operated by themData-residency requirements
Self-Hosted EnterpriseFull control over infrastructureLarge enterprises

On pricing (mid-2026 figures): LangSmith's Developer tier is free (1 seat, 5k base traces per month), Plus is $39/seat/month (10k base traces included; overage billed per trace at roughly $2.5 per 1k). On the Deployment side, billing moved to a per-run model in early 2026 (about $0.005/run). These numbers change frequently; check langchain.com/pricing before committing.

You can deploy without the platform

The langgraph library itself is MIT-licensed, and a compiled graph is just a plain Python object — wrap it in FastAPI and serve it with no restrictions. What LangSmith Deployment sells is "you don't have to build the task queue, persistence layer, and ops scaffolding yourself," not a license to run. With modest traffic and no long-running tasks, self-hosting is entirely sufficient.

5. Ecosystem Roles: LangChain / LangGraph / LangSmith ​

One table to pin down the post-v1 division of labor, so you never confuse them again:

RoleWhen you'll touch it
LangChain v1Agent API layer: create_agent, middleware, init_chat_model as a unified model interface, hundreds of third-party integrationsEvery project — at minimum the model and tool integration layer
LangGraphAgent runtime: StateGraph orchestration, checkpointer, interrupt, subgraphsWhen the standard ReAct template isn't enough and you need custom control flow
LangSmithObservability and engineering platform: tracing, eval, prompt management, DeploymentDebugging production behavior, regression evaluation, managed deployment (see Observability)

A note on historical baggage: the langchain package slimmed down dramatically in v1 — LLMChain, RetrievalQAChain, legacy retrievers, the hub, and more all moved into langchain-classic. When maintaining old code, fix the imports per the migration guide; don't use these in new code.

6. Pros and Cons: An Honest Assessment ​

Strengths:

  • Control granularity that nothing else offers. State structure, routing logic, and interrupt points are all declared explicitly, so you can answer "which step is the agent on right now, what's the state, why did it take this edge" — something black-box agent SDKs cannot give you.
  • Persistence and fault tolerance are among the most mature in the industry. Super-step-level checkpoints, crash recovery, time travel, idempotent semantics; add a Postgres checkpointer and you have a durable execution engine with almost no peers among similar frameworks.
  • Ecosystem gravity. The broadest model/tool integrations, the easiest hiring, the most Stack Overflow answers and community material (Chinese included), and LangSmith's tracing experience is currently the benchmark.

Weaknesses and real community criticism:

  • A steep learning curve. The concept density of state/reducer/super-step/Command/Send/interrupt is high; to write a correct human-in-the-loop flow you must first internalize counterintuitive semantics like "the node re-runs from the top on resume." Official docs churned heavily in the v0.x era, and the mix of stale and current tutorials online is the biggest source of beginner noise (improved after v1 froze the API).
  • Overkill for simple scenarios. Hacker News has hosted recurring high-heat threads (e.g., the March 2025 "We chose LangGraph to build our coding agent" post) with two camps: one argues the graph abstraction is a useful mental model, the other says outright "most scenarios need no framework — calling the model API in a hand-written loop is clearer." The criticism is fully valid for linear, stateless flows with no human intervention.
  • The debugging tax of the abstraction layer. Graph behavior is scheduled by the framework, so when something breaks you need to understand both your own code and LangGraph's execution model. Version mismatches between langgraph-prebuilt and the langgraph core package causing import errors are not rare in the issue tracker (e.g., issue #7404 in April 2026). Fortunately, LangSmith traces provide a backstop; debugging LangGraph with print statements alone is painful.
  • Commercial gravitational pull. The library itself is open source, but the best-experience path (Studio, Deployment, tracing) all leads into the paid LangSmith ecosystem; teams that mind this should think it through up front.

7. When to Use It, and Alternatives ​

Signals that LangGraph is the right pick:

  1. The flow has branches, loops, parallel fan-out — it isn't a straight line;
  2. Cross-session memory, crash recovery, or approval-gated pauses are needed — i.e., "stateful + long-running";
  3. A multi-agent system where handoffs must be controllable and observable;
  4. The team is already inside the Python/LangChain ecosystem.

Alternatives at a glance:

ScenarioBetter-fitting option
Standard ReAct agent, no orchestration details wantedLangChain create_agent, or the OpenAI Agents SDK (lighter, OpenAI ecosystem)
Coding / terminal agentsClaude Agent SDK (the same harness as Claude Code)
Role-play style multi-agent collaboration, low-code expressionCrewAI
Research / conversational multi-agentAutoGen (AG2 / absorbed into the Microsoft Agent Framework ecosystem)
Simple flow, self-managed stateHand-write the agent loop — a while loop + SQLite often does it in ~200 lines

The one-line conclusion: LangGraph's value is positively correlated with flow complexity. Using it for linear tasks is a pure self-imposed tax; but as soon as the business shows requirements like "pause for approval for three days, then resume," "crash and continue from the breakpoint," or "five agents collaborating as a state machine," it is currently the choice that saves the most engineering effort. A suggested path: prototype with create_agent (see Build an Agent from Scratch), drop down to StateGraph when you hit the orchestration ceiling, and read through Common Pitfalls along the way.

References ​