Stateful agents (LangGraph)

A ReAct loop is a while-loop with a model in it, and that is fine until something in the middle fails on step nine of eleven. Framing an agent as a graph — nodes that read and write a shared state object, edges that decide what runs next — buys two things a loop does not have: the run can be checkpointed and resumed, and the control flow is a thing you can look at rather than a thing buried in if statements. Both are worth measuring, because both are about what a failure costs you.

LangGraphshared statecheckpointing cyclesrecursion limit

What a failure costs, with and without checkpoints

Take a pipeline of nodes — research, draft, critique, revise, format, publish — where each node can fail: a tool times out, an API rate-limits, a parse fails, the model returns something unusable. Without persisted state a failure means running the whole thing again from the top. With a checkpointer, the state after each successful node is saved and the retry resumes from there. Below, 3000 runs of each:

Node executions to get one successful run

each one an LLM or tool call
timeouts, rate limits, unparseable output

The gap grows in a particular way that is worth naming. Without checkpoints, the probability that a whole run survives is (1−p)n, so the expected number of full attempts is its reciprocal — which climbs exponentially in the number of nodes. Every node you add makes every other node's failure more expensive. With checkpoints the cost is linear: each node is retried on its own, and a failure late in the graph costs one node, not nine.

That is also the mechanism behind the features people actually want from an agent framework. Human-in-the-loop is a checkpoint you do not automatically resume from — the graph interrupts before a node, the state sits in a database, and a person approves it hours later. Time travel is loading an older checkpoint and running forward with a changed state. Neither is possible if the run only exists as local variables inside a function.

from langgraph.graph import StateGraph, END
from typing_extensions import TypedDict, Annotated
import operator

class State(TypedDict):                       # the shared state — the real interface
    question: str
    notes: Annotated[list, operator.add]      # reducer: nodes APPEND, they don't overwrite
    draft: str
    attempts: int

def critique(state: State) -> dict:
    return {"attempts": state["attempts"] + 1}   # return only what changed

def route(state: State) -> str:               # a conditional edge is just a function
    if state["attempts"] >= 3:
        return END
    return "revise" if needs_work(state["draft"]) else END

g = StateGraph(State)
g.add_node("draft", draft_node); g.add_node("critique", critique); g.add_node("revise", revise_node)
g.add_edge("draft", "critique")
g.add_conditional_edges("critique", route)    # the cycle lives here
g.add_edge("revise", "critique")
app = g.compile(checkpointer=SqliteSaver.from_conn_string("runs.db"))

Two details in that snippet do most of the work. The reducer on notes — Annotated[list, operator.add] — is what makes concurrent branches safe: two nodes writing to the same key would otherwise be a last-write-wins race, and the reducer says how to merge instead. And nodes return a partial dict of what changed rather than the whole state, which is what makes a node independently testable: it is an ordinary function from state to a small diff.

The cycle, and why it needs a leash

The interesting graphs are not pipelines, they are loops: draft → critique → revise → critique. That is what makes an agent an agent, and it is also how you spend $400 overnight. Each pass has some chance of producing something the critic accepts; if it doesn't, the loop goes round again:

How many times round?

per pass through the loop
LangGraph's recursion_limit — the cap on loop iterations

The distribution is geometric, which means it has a long tail, which means the mean is a bad summary. Most runs finish in a couple of passes and a small minority go round many times — and it is that minority, not the average, that produces the bill and the timeout. The cap converts an unbounded tail into a bounded cost and a known failure mode: some runs end unfinished, and you get to decide what happens to them.

Which is the design rule: a cycle without a cap is a bug, and a cap without a fallback is a different bug. Set recursion_limit deliberately, and make the "gave up" path do something useful — hand the best draft to a human, return the partial result with a flag, or escalate to a stronger model for one attempt. LangGraph raises GraphRecursionError when the limit is hit, which is the framework refusing to let you not decide.

⚠️ Traps & honesty: node failures are modelled as independent, which they are not — a rate limit fails every node for a minute, and a bad upstream result makes the next node likelier to fail · the retry model assumes retrying a node eventually works, so the checkpointed cost never diverges; a deterministically broken node loops forever in both worlds · "node executions" is a proxy for cost, but nodes are not equally expensive, and the failing ones are often the slow ones · checkpointing is not free: it is a database write per node, and for a fast graph that overhead is real · the loop model gives every pass the same acceptance probability, whereas a real critique loop should be improving — if yours isn't, the cap is hiding a broken revise step rather than protecting you from an expensive one.
Takeaways: a graph is a loop plus persistence — nodes read and write a shared typed state, and edges are functions that choose the next node · without checkpoints the chance a whole run survives is (1−p)ⁿ, so cost climbs exponentially with graph size; with them it is linear, because a late failure costs one node instead of the whole run · that same persistence is what human-in-the-loop approval and time-travel debugging actually are · reducers make concurrent writes to one state key safe, and returning a partial diff is what makes a node testable on its own · loop lengths are geometric, so the mean hides the tail that produces the bill — cap the cycle, and give the "gave up" branch somewhere to go. Next: agent memory.

Second opinion (taught here — these corroborate): LangGraph docs · LangGraph · persistence · Anthropic · building effective agents.