Lesson 12.6 · 28 min
LangGraph: Graph-Based Agent Orchestration Explained
How do we build an agent that can loop, branch, remember a customer across sessions and pause for a manager's approval, without it turning into a tangle of if-statements?
In short: LangGraph is an open-source library from the LangChain team for building agents as graphs. We define a typed state, nodes that update it, and edges (fixed or conditional) that decide what runs next; cycles are allowed, so agent loops are natural. A checkpointer saves the state after every step per conversation thread, which gives memory, crash recovery, time travel and human-in-the-loop pauses. It is graph engineering, packaged as a library.
What is LangGraph?
LangGraph is an open-source library (Python and JavaScript) from the company behind LangChain, first released in early 2024. It lets us describe an LLM application as a graph: nodes are steps, edges connect them, and a shared state flows through. Unlike a simple chain, the graph may contain cycles, which is exactly what an agent needs: think, act, observe, think again.
LangGraph is the practical form of the graph engineering ideas from earlier in this module. It is a lower-level tool than LangChain's chains: it does not decide how our agent should behave; it gives us precise control over the flow and handles the hard infrastructure parts (state merging, persistence, streaming, interrupts). LangChain 1.0's create_agent is itself built on LangGraph.
Think of a board game The board (graph) has squares (nodes) and arrows (edges), some of which say “roll: if 6, take the shortcut” (conditional edges). Each player carries a card with their position and items (state). You can save the game, leave, and come back tomorrow exactly where you were (checkpointer), and a referee can pause play to approve a move (human-in-the-loop).
Running example: a shop support agent that looks up orders with a tool, remembers each customer's conversation, and asks a human before issuing refunds.
Why do we need LangGraph?
Chains are great for straight pipelines: retrieve → prompt → model → parse. Agents are different. They need to:
- Loop an unknown number of times (call tools until the answer is ready).
- Branch on decisions (tool call or final answer? refund or question?).
- Keep state across steps and across separate user messages.
- Survive interruptions: resume after a crash, or wait hours for a human approval.
- Stay controllable: limit steps, force certain checks, show progress while running.
We could hand-write all of this with while-loops and dictionaries, and every team used to. LangGraph provides it as a tested runtime so we write only the business logic: the nodes and the routing.
What is a graph and what is state in LangGraph?
A LangGraph graph starts with StateGraph(State), where State is a schema (usually a Python TypedDict or a Pydantic model) listing the fields every node can read and update. For a chat agent the state is often just a list of messages.
Nodes do not overwrite the state; they return updates, and LangGraph merges them. How a field is merged is set by a reducer. With no reducer, a new value replaces the old one. With Annotated[list, add_messages], new messages are appended to the list (and a message with an existing id replaces that one). Reducers are what make parallel branches safe: two nodes can both add messages without one erasing the other.
Pause and think: A state field notes: list has no reducer. Node A returns {"notes": ["x"]}, then node B returns {"notes": ["y"]}. What is notes after both?
["y"]. Without a reducer the default is replace, so B's value overwrites A's. With Annotated[list, operator.add] it would be ["x", "y"].
Nodes, edges and conditional edges
A node is a plain function node(state) -> update, added with builder.add_node("name", fn). It can call an LLM, run a tool, query a database, or do pure logic.
A normal edge, builder.add_edge("a", "b"), always goes from a to b. Special names START and END mark where a run enters and leaves. A conditional edge, builder.add_conditional_edges("a", router), calls router(state) after node a; the router returns the name of the next node (or END). This is where decisions live, for example: “did the model ask for a tool? then go to tools, else finish”.
Calling builder.compile() checks the graph (for example, that edges point at real nodes) and returns a runnable app with invoke, stream and async versions. A recursion_limit setting caps the number of steps per run, so a cycle cannot spin forever.
A complete example
Below: the real LangGraph code for our support agent, and a runnable mini version that implements the same mechanics (reducer, conditional edge, cycle, checkpointer per thread) in plain Python.
Tools and who calls them
A frequent confusion: the model never runs a tool itself. It only outputs a structured request: “call get_order_status with order_id='42'”. Something else must execute that request. In LangGraph that something is a node, usually the prebuilt ToolNode, which reads the tool calls from the last AI message, runs the matching Python functions, and appends ToolMessage results to state.
Life of one tool call
- Bind:
llm.bind_tools([...])sends the tool names, descriptions and argument schemas with every model request. - Request: The model replies with an AI message whose
tool_callslist names the tool and arguments, instead of (or as well as) text. - Route: The conditional edge (
tools_condition) sees the tool calls and routes to the tools node. - Execute:
ToolNoderuns each function with the given arguments. Errors can be caught and returned as messages so the model can try again. - Return: Results are appended as ToolMessages, the edge goes back to the agent node, and the model reads them.
Why this split is good Because our code executes tools, we control them: validate arguments, enforce permissions, add timeouts, log every call, or pause before risky ones. The model only proposes.
Memory and persistence
When we compile with a checkpointer, LangGraph saves a snapshot of the state after every step (technically, every “super-step”), keyed by the thread_id we pass in the config. A thread is one conversation or one job. This single feature gives several capabilities:
- Short-term memory: calling again with the same thread id continues the conversation, as in our example.
- Fault tolerance: if a run crashes mid-way, it can resume from the last saved step.
- Time travel:
get_state_historylists past snapshots; we can inspect or re-run from any of them to debug. - Human-in-the-loop: a run can stop and wait indefinitely, because its state is safely stored.
InMemorySaver is for development; production uses database-backed checkpointers (for example SQLite or Postgres versions). For long-term memory shared across threads (a customer's preferences remembered in every new chat), LangGraph offers a separate store interface for key-value documents, which nodes can read and write.
Thread ids are a privacy boundary Everything in a thread is replayed to the model. Reusing one thread id for different users leaks one customer's conversation to another. Derive thread ids from the authenticated user and conversation, never from user-supplied text.
Human-in-the-loop
Because state is checkpointed, LangGraph can interrupt a run. Inside a node we call interrupt(payload); the graph saves its state and returns the payload to our application (for example, “approve refund of $250 for order 42?”). Later we resume with graph.invoke(Command(resume=answer), config) on the same thread, and interrupt returns the human's answer inside the node, which continues from there. Graphs can also be compiled with static breakpoints such as interrupt_before=["tools"] to pause before every tool call.
approval step (real LangGraph, not run here)
from langgraph.types import interrupt, Command
def refund(state):
amount = state["refund_amount"]
if amount > 200: # policy: big refunds need a human
decision = interrupt({"question": f"Approve refund of ${amount}?"})
if decision != "approve":
return {"messages": [("ai", "A manager declined this refund.")]}
return {"messages": [("ai", f"Refund of ${amount} issued.")]}
# first call pauses at interrupt(); later, after the manager clicks approve:
# graph.invoke(Command(resume="approve"), {"configurable": {"thread_id": "customer-1"}})Pause and think: Why can't human-in-the-loop work reliably without a checkpointer?
Because the run must stop and later continue exactly where it was, possibly in a different process hours later. Without saved state there is nothing to resume from; the whole run would have to start over.
When to use LangGraph
Common mistakes Forgetting a reducer and silently overwriting a list; using one shared thread id; putting giant blobs (whole PDFs) in state so every checkpoint is huge; and building a 30-node graph for a job a single prompt could do.
Real-world use Teams use LangGraph for customer-support agents with escalation, research agents with plan-search-write loops, document-processing pipelines with validation and review, and supervisor setups where one node routes work to specialist sub-agents.
Worked example, step by step
The merge rule decides everything a node can see, so let us trace it by hand on the mini version from this lesson. The state has one field, messages, with an append reducer. The user asks “where is order 42”.
| Step | Who ran | Update returned | Messages after merge |
|---|---|---|---|
| 0 | input | user: where is order 42 | 1 |
| 1 | agent | tool_call:get_order_status:42 | 2 |
| 2 | tools | tool:shipped | 3 |
| 3 | agent | ai: Your order is shipped. | 4 |
Each node returned one message, not the whole list. The reducer did the appending. After step 3 the conditional edge finds no tool call and routes to END, and 4 messages are saved under t1. The second turn loads those 4 and adds 4 more, which is the 8 in the output.
The same run with the default rule, replace
- Input arrives:
messagesis replaced by the new input: 1 message. No harm yet. - agent runs: It returns the tool call.
messagesnow holds only the tool call. The user's question is gone. - tools runs: It reads the last message, runs the tool and returns the result. Only
tool:shippedis left. - agent runs again: Our fake model looks only at the last message, so it still answers. A real model would see a tool result with no question in front of it.
- Second turn: Loading the thread gives 1 message, not 4. The conversation memory is one line long.
Nothing crashed in that replay. This is why a missing reducer is hard to notice: the graph runs, and the damage is a model that quietly lacks context. When a list field in a saved state looks too short, check its reducer first.
Practice: try it yourself
We will build two ideas in plain Python: the merge rule with a reducer, and a pause that waits for a human. A refund over 200 stops the run before it is issued. The state is saved under a thread id, and a later call that carries the human's answer picks up where the run stopped. This is a sketch of the mechanism, not LangGraph's real API.
practice_pause_resume.py
import operator # a plain-Python sketch of the ideas, not the real LangGraph API
REDUCERS = {"notes": operator.add} # no entry = replace (the default)
def merge(state, update):
# new_state[field] = reducer(old_state[field], update[field])
new = dict(state)
for field, value in update.items():
new[field] = REDUCERS[field](state[field], value) if field in REDUCERS else value
return new
def lookup(state): return {"notes": ["order 42 costs 250"], "amount": 250}
def refund(state):
if state["amount"] > 200 and state["approval"] is None:
return "PAUSE" # like an interrupt: wait for a human
ok = state["amount"] <= 200 or state["approval"] == "approve"
return {"notes": ["refund issued" if ok else "refund declined"], "status": "done"}
NODES = [("lookup", lookup), ("refund", refund)]
saved = {} # thread id -> (next node index, state)
EMPTY = {"notes": [], "amount": 0, "approval": None, "status": "open"}
def invoke(thread_id, resume=None):
index, state = saved.get(thread_id, (0, EMPTY))
if resume is not None:
state = merge(state, {"approval": resume}) # the human's answer is an update
while index < len(NODES):
name, fn = NODES[index]
update = fn(state)
if update == "PAUSE":
saved[thread_id] = (index, state) # checkpoint, then stop
return f"paused at {name}: approve refund of {state['amount']}?"
state, index = merge(state, update), index + 1
saved[thread_id] = (index, state) # checkpoint after every step
return f"{state['status']}: {state['notes']}"
print(invoke("customer-1"))
print(invoke("customer-1", resume="approve"))
print(invoke("customer-2"))
print(invoke("customer-2", resume="no"))Output:
paused at refund: approve refund of 250? done: ['order 42 costs 250', 'refund issued'] paused at refund: approve refund of 250? done: ['order 42 costs 250', 'refund declined']
Now change it:
- Remove
"notes"fromREDUCERS. Predict the finalnoteslist for customer-1. - Change the amount in
lookupto 150. Predict what the first call prints. Does the run pause at all? - Call
invoke("customer-1")twice in a row with noresume. Predict the second result. Why is it safe to ask twice?
Pause and think: When we resume, our invoke runs the refund node again from its first line; it does not jump into the middle of the function. Why does the refund still come out right, and what should we keep out of the lines before the pause?
On the second run approval is in the state, so the node skips the pause and finishes. This works because everything before the pause is a pure check. If the node did something with an outside effect before pausing, such as sending an email, that action would happen twice. Keep side effects after the pause, or make them safe to repeat.
Pause and think: customer-1 and customer-2 paused with the same question and then got different endings. Where did each run keep its place while it waited? What would happen if both used one thread id?
In saved, under its own thread id: the index of the next node plus the full state. That entry is the only thing that links the second call to the first. With a shared thread id, the second customer's run would overwrite the first one's saved state, and one manager's answer could be applied to the other customer's refund. Each conversation needs its own thread id.
Key takeaways
- LangGraph builds agents as graphs of nodes, edges and a typed shared state, with cycles allowed.
- Nodes return updates; reducers decide how updates merge into the state.
- Conditional edges hold the decisions, such as “tool call or finish?”.
- The model requests tools; a node in our graph executes them.
- A checkpointer saves state per thread, enabling memory, recovery, time travel and human approval pauses.
- Use LangGraph when control flow, durability or approvals matter; use simpler tools otherwise.
Key terms
- StateGraph: LangGraph's builder for a graph whose nodes share and update a typed state.
- Reducer: A function that merges a node's update into a state field, such as appending messages.
- Conditional edge: An edge that calls a router function on the state to choose the next node.
- ToolNode: A prebuilt node that executes the tool calls requested in the last AI message.
- Checkpointer: A component that saves state snapshots after each step, keyed by thread id.
- Interrupt: A pause inside a node that waits for external (usually human) input before resuming.
← 12.5 LangChain: Composable Components for LLM Applications · 12.7 Claude Code: AI-Powered Software Engineering at the CLI →