A ReAct loop is a while-loop with a model in it, and that is
fine until something in the middle fails on step nine of eleven. Framing an agent as a graph — nodes
that read and write a shared state object, edges that decide what runs next — buys two things a loop does
not have: the run can be checkpointed and resumed, and the control flow is a thing you can look at
rather than a thing buried in if statements. Both are worth measuring, because both are about
what a failure costs you.
Take a pipeline of nodes — research, draft, critique, revise, format, publish — where each node can fail: a tool times out, an API rate-limits, a parse fails, the model returns something unusable. Without persisted state a failure means running the whole thing again from the top. With a checkpointer, the state after each successful node is saved and the retry resumes from there. Below, 3000 runs of each:
The gap grows in a particular way that is worth naming. Without checkpoints, the probability that a whole
run survives is (1−p)n, so the expected number of full attempts is its reciprocal —
which climbs exponentially in the number of nodes. Every node you add makes every other node's failure more
expensive. With checkpoints the cost is linear: each node is retried on its own, and a failure late in the
graph costs one node, not nine.
That is also the mechanism behind the features people actually want from an agent framework. Human-in-the-loop is a checkpoint you do not automatically resume from — the graph interrupts before a node, the state sits in a database, and a person approves it hours later. Time travel is loading an older checkpoint and running forward with a changed state. Neither is possible if the run only exists as local variables inside a function.
from langgraph.graph import StateGraph, END from typing_extensions import TypedDict, Annotated import operator class State(TypedDict): # the shared state — the real interface question: str notes: Annotated[list, operator.add] # reducer: nodes APPEND, they don't overwrite draft: str attempts: int def critique(state: State) -> dict: return {"attempts": state["attempts"] + 1} # return only what changed def route(state: State) -> str: # a conditional edge is just a function if state["attempts"] >= 3: return END return "revise" if needs_work(state["draft"]) else END g = StateGraph(State) g.add_node("draft", draft_node); g.add_node("critique", critique); g.add_node("revise", revise_node) g.add_edge("draft", "critique") g.add_conditional_edges("critique", route) # the cycle lives here g.add_edge("revise", "critique") app = g.compile(checkpointer=SqliteSaver.from_conn_string("runs.db"))
Two details in that snippet do most of the work. The reducer on notes —
Annotated[list, operator.add] — is what makes concurrent branches safe: two nodes writing to the
same key would otherwise be a last-write-wins race, and the reducer says how to merge instead. And nodes
return a partial dict of what changed rather than the whole state, which is what makes a node
independently testable: it is an ordinary function from state to a small diff.
The interesting graphs are not pipelines, they are loops: draft → critique → revise → critique. That is what makes an agent an agent, and it is also how you spend $400 overnight. Each pass has some chance of producing something the critic accepts; if it doesn't, the loop goes round again:
The distribution is geometric, which means it has a long tail, which means the mean is a bad summary. Most runs finish in a couple of passes and a small minority go round many times — and it is that minority, not the average, that produces the bill and the timeout. The cap converts an unbounded tail into a bounded cost and a known failure mode: some runs end unfinished, and you get to decide what happens to them.
Which is the design rule: a cycle without a cap is a bug, and a cap without a fallback is a different
bug. Set recursion_limit deliberately, and make the "gave up" path do something useful — hand
the best draft to a human, return the partial result with a flag, or escalate to a stronger model for one
attempt. LangGraph raises GraphRecursionError when the limit is hit, which is the framework
refusing to let you not decide.
(1−p)ⁿ, so cost climbs exponentially with graph size; with them it is linear, because
a late failure costs one node instead of the whole run · that same persistence is what human-in-the-loop
approval and time-travel debugging actually are · reducers make concurrent writes to one state key safe, and
returning a partial diff is what makes a node testable on its own · loop lengths are geometric, so the mean
hides the tail that produces the bill — cap the cycle, and give the "gave up" branch somewhere to go. Next:
agent memory.Second opinion (taught here — these corroborate): LangGraph docs · LangGraph · persistence · Anthropic · building effective agents.