Skip to content

Case Study: LangGraph

At a glance A deep dive into LangGraph's harness design — the "framework camp" approach that models agent control flow explicitly as a state graph, how its philosophy compares with the Claude Code-style free-running loop, and when to reach for a framework versus writing your own loop.

Case Study: LangGraph ​

An agent orchestration framework released by the LangChain team in January 2024. It represents the "framework camp" of harness design: rather than trusting a bare while loop, it models agent control flow explicitly as a graph — states, nodes, edges, breakpoints, persistence, all of it declared by the developer.

What It Is: A Graph, Not a Loop ​

In the Claude Code case study you'll meet the other extreme: the harness is, at its core, a single main loop. The model decides for itself when to call tools and when to stop; the entire "cognitive architecture" hides in the prompt and tool design, and the control flow is nearly invisible to the developer.

LangGraph stands on the opposite side. Its thesis: once an agent is headed for production, a free-running loop isn't controllable enough. LangChain was blunt in the launch announcement — in practice, they found, companies pushing agents into production kept needing firmer control: forcing the agent to call a specific tool as its first step, using different prompts in different states, constraining precisely how tools get called. Internally, they called this more disciplined shape a "state machine":

A state machine keeps the power of the loop — it can handle inputs far more ambiguous than a simple chain — but human guidance remains in how the loop is constructed.

LangGraph is the way of expressing that state machine as a graph. The core interface the whole framework exposes is astonishingly narrow: a single StateGraph class. You define the state, add nodes, connect edges, compile, run — that's it.

In one sentence

LangGraph officially positions itself as a "low-level agent orchestration framework": no hidden prompts, no built-in cognitive architecture — what you get is durable execution and fine-grained control. The problem it aims to solve is not "helping you stand up a demo quickly" but "keeping complex agentic systems running reliably in production."

Why You Need It: From DAG to Loop, From Loop to Controlled Loop ​

To understand LangGraph's motivation, you first have to look at the lessons of LangChain's own evolution:

  1. Stage one: chains (Chain/DAG). LangChain's early core abstraction was the chain — a pipeline that goes one step after another, at heart a directed acyclic graph (DAG). The canonical example is RAG: retrieve → generate, one straight road to the end. The problem is that a DAG cannot express loops like "the retrieval results are bad, rewrite the query and try again."
  2. Stage two: the free loop (AgentExecutor). Introduce a loop and let the LLM reason about what to do next inside it — "essentially putting the LLM into a for loop." This is the AutoGPT school, and a close cousin of Claude Code's main loop. It is the simplest design, and also the most radical: almost all decision-making authority goes to the model.
  3. Stage three: the controlled loop (LangGraph). The price of freedom is unpredictability. What production needs is: loops can exist, but the loop's topology is defined by humans, and the model only chooses within the bounds the structure allows.
Evolution path:

  Chain (DAG)               AgentExecutor                 LangGraph
 ┌────────┐   ┌────────┐    ┌────────────┐              ┌────────────┐
 │retrieve│ → │generate│    │    LLM     │              │    LLM     │←──┐
 └────────┘   └────────┘    │  ↓     ↑   │              │     ↓      │   │
                            │ tool call  │              │conditional │   │
                            └────────────┘              │     ↓      │   │
  No loop; controllable     └─ model decides ─┘         │   tools    │───┘
  but rigid                                             └────────────┘
                                                        Has a loop, but the topology is drawn by humans

Note a subtle fact: the free loop is a proper subset of LangGraph. Recreating a ReAct loop in LangGraph takes just two nodes (a model node plus a tool node) and one conditional edge. So this "philosophical dispute" was never about what's possible — it's about who owns the defaults and the control.

Core Mechanics, Dissected ​

State: Shared State With Merge Rules ​

Every node in the graph shares one state object. You pass the state definition when you create the StateGraph, and each field declares two things: what it stores and how it merges (the reducer):

  • Override: the node returns a new value that directly replaces the old one. Suited to fields like "current plan" or "next action."
  • Append (add): the node returns a delta that is automatically accumulated onto the old value. The classic application is the message list — each node returns only the messages it just produced, and the framework stitches the history together.
python
from typing import TypedDict, Annotated
import operator
from langgraph.graph import StateGraph

class State(TypedDict):
    input: str
    # Annotated[..., operator.add] declares: this field merges by "add"
    all_actions: Annotated[list[str], operator.add]

graph = StateGraph(State)

This reducer design is LangGraph's most underrated decision. It keeps nodes pure: a node doesn't care about the history of the global state; it only declares "here is the delta I produced this step." State management gets stripped out of business code and sinks down into the framework layer. Compare that with the messages list in a hand-written loop — the one that keeps growing while everyone reads and writes it directly. This is framework-camp fastidiousness at its finest: one more layer of abstraction, traded for one more rule of discipline.

Node: State In, Delta Out ​

Nodes are plain functions (or LangChain Runnables), with the signature "read the state dict → return a dict of the fields to update":

python
def call_model(state: State) -> dict:
    # Return only the delta; never touch the full state
    return {"all_actions": ["model_called"]}

graph.add_node("model", call_model)
graph.add_node("tools", tool_executor)

There is also a special node, END, marking the graph's terminus. The announcement specifically warns: your loop must be able to eventually reach END — the graph's expressive power includes infinite loops, and the responsibility for "when to stop" is handed back to the graph's designer.

Edges: The Ones Humans Draw and the Ones the Model Chooses ​

LangGraph has three kinds of edges, and the three map neatly onto three levels of control authority:

Edge typeHow it's writtenWhat it meansWho decides
Entry edgeset_entry_point("model")The graph starts hereHuman
Normal edgeadd_edge("tools", "model")After A, always go to BHuman
Conditional edgeadd_conditional_edges(...)A function (usually LLM-driven) decides where to goHuman draws the options, model picks

The conditional edge is the key to understanding LangGraph's philosophy. It takes three things: the upstream node, a routing function, and a mapping from routing results to node names. Note the existence of that mapping — the model can only choose among predefined options; it cannot go anywhere the graph doesn't contain. This is where "controlled" actually lands: the model's freedom is fenced, precisely, within the set of edges.

python
graph.add_conditional_edges(
    "model",
    should_continue,          # routing function: decides where to go based on the model's output
    {
        "end": END,           # model says "done" → finish
        "continue": "tools",  # model says "call a tool" → tool node
    },
)

app = graph.compile()  # compile into a runnable Runnable

Checkpointer: Making "Where Are We?" a First-Class Citizen ​

Up to this point, LangGraph is still just a "flowchart engine with loops." What truly sets it apart from ordinary orchestration libraries is the persistence layer.

The checkpointer writes a full snapshot of the state to storage (in-memory, SQLite, Postgres, and so on) at every super-step of graph execution, and uses a thread_id to identify one thread of execution:

python
from langgraph.checkpoint.memory import MemorySaver

checkpointer = MemorySaver()
app = graph.compile(checkpointer=checkpointer)

config = {"configurable": {"thread_id": "user-42"}}
app.invoke({"input": "I want a refund"}, config)   # auto-archived after every node runs

With it, a whole set of capabilities you'd have to hack together by hand in a free loop become infrastructure:

  • Crash recovery: the process dies, the machine reboots — execution resumes from the last checkpoint, no progress lost.
  • Time travel: replay the state at any step in history and re-run from a fork at some checkpoint in the middle — the killer debugging technique for agents.
  • Multi-session isolation: different thread_ids are independent state streams, naturally supporting multi-tenancy and multiple concurrent conversations.

The fundamental difference from a hand-written loop

In a hand-written loop, the "current state" is just a few variables in process memory; when the process dies, it all dies. LangGraph turns agent execution into something like a database transaction: state lives outside the process, every step is archived, everything is resumable and replayable. This is the framework camp's hardest engineering value, and the most practical support for long-running tasks (jobs that run for hours, approvals that span days of human review).

Interrupt: Modeling the Human as a Kind of Graph Event ​

Human-in-the-loop in a free loop usually means breaking the loop, popping a confirmation dialog, and stuffing the result back in — special-case code everywhere. LangGraph's approach is to make pause/resume a first-class runtime mechanism.

Call interrupt() anywhere inside a node and the graph stops exactly there: the state is archived, the payload is thrown to the caller, and the graph waits indefinitely. You then re-invoke the graph with Command(resume=...); the resume value becomes the return value of the interrupt() call, and the node keeps running:

python
from langgraph.types import interrupt, Command

def refund_node(state: State) -> dict:
    # When execution reaches this line, the graph pauses and hands the refund details out for human review
    decision = interrupt({"action": "refund", "amount": state["amount"]})
    if decision["approved"]:
        do_refund(state["amount"])
    return {"all_actions": ["refund_processed"]}

# Resumed from outside (hours later, in another process):
app.invoke(Command(resume={"approved": True}), config)

Besides the dynamic interrupt(), there are also static breakpoints: set interrupt_before=["tools"] at compile time or run time, and the graph pauses before executing a given node — a direct expression of the most common safety pattern, "human approval before a tool call."

An easy trap to fall into

interrupt()'s resume semantics are "re-run the whole node from the top," not "pick up from the line after interrupt()." The official docs therefore lay down a few rules: don't put non-idempotent side effects before interrupt (write to the database first and ask a human second, and the write happens twice on resume); don't wrap interrupt in a bare try/except (pausing is implemented by throwing an exception, which you would swallow); and if a node has multiple interrupts, their order must stay consistent across every execution. These constraints are the price of the "replay-based resume" design choice.

State Machine vs. Free Loop: What the Debate Is Really About ​

Putting LangGraph and Claude Code side by side is the best controlled experiment for understanding the entire harness design space. Both sides run models of the same caliber, yet they reached opposite architectural conclusions:

DimensionLangGraph (framework camp)Claude Code (free-loop camp)
Control flowExplicit: the graph's topology is hard-codedImplicit: the model improvises inside the loop
Core abstractionState / Node / Edge / CheckpointerOne main loop + a set of tools
Model's freedomFenced in by conditional edges, can only pick existing optionsNearly full authority, until it says "done" itself
Where the cognitive architecture livesIn the graph structure — visible and auditable as codeIn the system prompt — visible as text
State persistenceFramework-level checkpointer, for freeMostly process memory; session storage is DIY
Human in the loopFirst-class mechanism (interrupt / breakpoints)Permission layer intercepts tool calls
DebuggingTime travel, state snapshots, graph visualizationRead the trace logs, replay the transcript
Where it hurts mostThe graph design itself becomes software that must be maintainedBehavior is never fully predictable; boundaries are held by prompts

The substance of the debate is how much you trust the model. The free-loop camp's bet: models are already smart enough, so the best harness hands them clean tools and ample context, then gets out of the way; any flowchart a human draws is a prior prejudice that will turn into a shackle once models get stronger. The framework camp's bet: enterprise settings cannot tolerate "right most of the time" — predictability, auditability, and recoverability at every step are hard requirements, and none of them can grow out of a tangle of improvised loops.

Notably, even Anthropic itself concedes this tension. Its "Building effective agents" essay divides systems into two kinds: workflows (the LLM is orchestrated along predefined code paths) and agents (the LLM dynamically directs its own process) — which is nearly a definition of LangGraph versus Claude Code. And the essay's advice: look for the simplest solution first, since many tasks don't need an agentic system at all; frameworks let you start fast, but they also make it easy to pile up hard-to-debug layers of abstraction.

A subtler observation: both sides are converging toward each other

The official narrative of the LangGraph 1.0 era has already softened: it ships prebuilt high-level interfaces like create_agent, and underneath they are exactly the free loop of "model node + tool node + conditional edge" — implementing a loop with a graph, conceding that the loop is the right default starting point. Meanwhile, products like Claude Code have been growing framework-camp fixtures: permission rules, hooks, subagents — mechanisms that are all, at bottom, explicit structure added around a free loop. Pure positions no longer exist; what remains is a choice of where to sit on a slider.

Where It Fits and Where It Doesn't ​

Where LangGraph shines:

  • Businesses with highly deterministic flows: support ticket handling, claims approval, document review pipelines — the steps are mostly fixed, and the LLM handles the understanding and generation within them. Draw the flow as a graph and readability, auditability, and testability all come as gains.
  • Heavy-compliance / heavy-human-review settings: finance, healthcare, government. The interrupt + checkpointer combination makes "who approved each step, and what the state was at the time" fully traceable, and approvals can span hours or even days.
  • Long-running tasks and unreliable environments: batch agents that run for hours, where process crashes are the norm rather than the exception — durable execution is a must.
  • Multi-agent orchestration: complex topologies such as supervisor patterns and hierarchical teams are far clearer expressed as a graph than as nested loops.
  • Situations that need fine-grained control over model behavior: force a tool call as the first step, switch prompts per state, restrict the set of allowed actions — conditional edges were born for this.

Where LangGraph gets in the way:

  • Open-ended exploration tasks: "fix this bug in the repo," "research this topic" — the solution space cannot be drawn as a graph in advance. Force it, and all you get is a two-node graph of "model node → tool node → model node"; the framework's value degenerates into providing a checkpointer, while you still pay the complexity in full.
  • Rapid prototypes / one-off scripts: paying the concept tax of state schemas, reducers, compilation, and checkpointers for a weekend project isn't worth it.
  • Periods of fast model improvement: models get a notch stronger every few months, and the control flow you so carefully drew can quickly become a ceiling that caps the model. A free loop collects the gains automatically; a graph has to be redrawn by hand.
  • Teams with no budget for learning a framework: LangGraph's concepts are few, but they interlock (state merging, super-steps, replay semantics, Command), and the hidden cost of using them wrong is not small — interrupt's replay semantics have burned plenty of people.

A Decision Framework: Framework or Hand-Written Loop? ​

Compress the decision into four questions, and ask them in order:

Q1. Can the task's control flow be drawn out in advance?
    ├─ Yes, and the steps are mostly fixed ──────→ High payoff from a graph/framework (Q2)
    └─ No, the path is wide open ────────────────→ Lean toward a hand-written loop (Q4)

Q2. Do you need cross-process persistence, resume-from-crash, or human-approval breakpoints?
    ├─ Yes ──────────────────────────→ LangGraph is nearly a ready-made answer
    └─ No ───────────────────────────→ A hand-written DAG/pipeline may be simpler

Q3. What failure cost is acceptable?
    ├─ One wrong step = money lost / compliance incident → Controlled loop, every branch drawn by hand
    └─ Rerun and it's fine ──────────────────────────────→ Free loop + good tools + evals

Q4. Are you betting on models getting better?
    ├─ Yes, want to ride the gains automatically ─────→ Thin harness, minimal hard-coded flow
    └─ No, need to lock in today's behavior ──────────→ Explicit state machine, behavior frozen in the graph

In practice, many teams land on a hybrid architecture: the outer layer is a deterministic graph (ticket in → classify → route → archive), and inside one of its nodes runs a free-loop agent ("handle this category of ticket"). The graph owns the governable part, the loop owns the unpredictable part — the same layering idea this site's architecture anatomy keeps emphasizing.

And remember one anti-pattern: don't stuff everything into a graph just because LangGraph exists. If a while loop and twenty lines of code can express your logic, reaching for a graph abstraction is not engineering rigor — it's ceremony.

Further Reading ​

References ​