Most agents are one while-loop. It calls the model, appends the output, calls again, and keeps going until the context window fills and the answers turn to mush.
That loop was a real step forward. Instead of prompting once and copying the result, you built something that retried, verified, and improved with every pass. It is still the right answer for a single-purpose job.
It stops working the moment the task has more than one concern. Not because the model is weak: because the architecture is wrong. This is the roadmap from that loop to a graph.
A loop is one agent getting smarter. A graph is many agents getting coordinated.
01. Why one loop is not enough
One worker, five jobs
A loop gives one agent autonomy. It retries, it learns from failure, it improves with each pass. Fix this bug, summarise this document, clean this dataset: define the goal, attach a verifier, set a stop condition, let it run.
Trouble starts when the task carries multiple concerns. Research, analysis, code, review. One agent doing all four fills its window with mixed information, loses the thread of earlier steps, and ships work that looks finished and falls apart under inspection. The research bleeds into the analysis. The code ignores the review. Four jobs, one context budget.
The fix is not a bigger context window. It is splitting the work across agents that each do one thing, then wiring the output of one into the input of the next. It is the same move a solo developer makes when they join a team: roles, handoffs, reviews.
02. The four eras
Each one absorbed the last
Prompt engineering taught us to write the instruction. Context engineering taught us to feed the right data. Loop engineering taught us to build one autonomous agent. Graph engineering teaches us to wire many agents into a system.
Most people are still in era one or two: a prompt, some context, one call, and hope. The builders shipping real products are in three and four. They are not writing better prompts. They are designing better flows.
03. The four building blocks
Node, edge, router, state
Every graph, however complex, is assembled from four primitives. You do not need a framework. You need functions, a dictionary, and if-statements.
Node. A function that does one thing. A model call, a tool execution, a query, a file write. State in, state out. It does not know what ran before it or what runs after.
Edge. A connection between two nodes. Unconditional, or conditional on something in state. Conditional edges are how a graph makes decisions.
Router. A node that inspects state and picks the path. It does no work itself. It classifies and dispatches.
State. A dictionary that flows through the whole graph. Every node reads from it and writes to it. Without it, each node starts from scratch and the graph has no memory.
# The entire graph abstraction, in about 15 lines
def node(func):
"""A node takes state and returns state."""
def wrapper(state):
return func(state)
return wrapper
def edge(state, next_node):
"""An edge passes state to the next node."""
return next_node(state)
def router(state, condition, if_true, if_false):
"""A router picks a path based on state."""
if condition(state):
return if_true(state)
return if_false(state)
04. Pattern 01: sequential chain
A to B to C, no branches
The simplest graph. Node A runs, hands off to B, which hands off to C. Most things people call an AI pipeline are exactly this.
It works when every step succeeds and the order never changes: extraction, format conversion, translate then summarise. The moment a step can fail or the path can vary, you need one of the other patterns.
The real risk here is error propagation. Bad output from A becomes the foundation of B, which becomes the foundation of C. Three layers later the error is baked in and invisible. This is why even a trivial chain earns a verification node at the end.
05. Pattern 02: router
Classification, not reasoning
A router inspects the input and picks one of several paths. Not every node runs. This is how one system handles many kinds of task without loading every tool into one agent.
The insight most people miss: routing is a classification task, not a reasoning task. You do not need a frontier model to decide whether a ticket is about billing or a bug. Haiku does it in milliseconds for a fraction of the cost. Spend the expensive model inside the specialist, where the actual work happens.
Every specialist gets its own system prompt and its own tools. A code agent needs run_tests and write_file. A research agent needs web_search and read_doc. Give every agent every tool and you get wrong tool calls, wasted tokens, and confused output.
def router(state):
category = call_claude(
model="claude-haiku-4-5", # cheap, fast, classify
system="Classify this task. One word: code, research, or writing.",
prompt=state["task"],
)
routes = {
"code": code_agent, # own prompt, own tools
"research": research_agent,
"writing": writing_agent,
}
agent = routes.get(category.strip(), research_agent)
return agent(state)
If your router makes bad calls, the fix is almost always a better prompt, not a bigger model. Write five or six real examples of each category into the router prompt. Few-shot on Haiku beats zero-shot on Sonnet for routing.
06. Pattern 03: parallel fan-out
The slowest agent, not the sum
Several nodes run at once on the same input. A collector waits for all of them and merges. Three agents in parallel means your total time is the slowest of the three, not the sum.
The constraint is independence. If B needs A's output, they cannot run together. If both work from the same input and produce separate outputs, they can. Competitive analysis is the clean case: one agent per competitor, all running at once.
The merge node is where almost everyone underinvests. Merging three JSON blobs is trivial. Merging three research reports into one coherent synthesis needs its own prompt and its own quality bar. Treat it as a real agent, not a string concatenation.
I built a market-research graph for a D2C brand. One agent was working through every competitor one at a time, and a full sweep took most of an afternoon. I split it: one agent per competitor, all running at once, then one merge node that wrote the actual comparison. The whole run collapsed to the time of the slowest single agent. The merge was the hard part. Three agents each said a different rival was the cheapest, and the merge had to reconcile that into one honest read.
07. Pattern 04: loop with gate
The builder produces, the reviewer rejects
The builder does the work. A separate reviewer checks it. If the review fails, the builder retries with the failure reason appended to state. This is the pattern behind every self-correcting system.
The rule that makes it work: builder and reviewer must be different agents. Different prompts, often different model tiers. The builder's job is to produce. The reviewer's job is to reject. When one agent reviews its own work it approves its own mistakes, because the reasoning that produced the flaw is the reasoning evaluating it.
Write the reviewer prompt adversarially. Not "is this good", but "find every flaw, inconsistency, and missing edge case; if you find nothing, return passed true; if you are unsure, fail it." A strict reviewer catches real bugs. A polite one waves everything through.
On a claims-triage build for an insurer, the agent drafted the decision and signed off on its own work. It approved a claim it should have flagged, because the logic that made the call was the same logic checking it. I added a second agent with one job: reject. A different prompt, told to hunt for the reason to say no. The bad approvals stopped.
def loop_with_gate(task, max_tries=5):
state = {"task": task, "history": []}
for i in range(max_tries):
result = builder_agent(state) # produces
check = reviewer_agent(result) # rejects
if check["passed"]:
return result
state["history"].append({
"attempt": i + 1,
"issues": check["issues"],
})
return {"status": "max_retries", "last_result": result}
08. Pattern 05: human in the loop
The graph pauses, the human decides
The graph stops at a node and waits for approval. The agent does the research and drafts the action. The human reviews the decision that matters. The agent executes after approval.
This is mandatory for anything expensive, irreversible, or high-stakes: customer emails, production deploys, financial transactions, deleting data. The graph handles the work. The human handles the judgment call.
Make the approval node show exactly what will happen and what it costs. Not "do you approve", but "this will send 2,400 emails to the billing segment with this subject line and this body, estimated cost 12 dollars, approve?" Specific enough to decide in five seconds.
09. State flow and model tiering
Two production concerns
State is a flat dictionary every node reads and writes. Keep it explicit: each node declares which keys it reads and which it writes. When a graph breaks, state is the first thing you check. If you cannot trace which node wrote which key, you cannot debug anything.
Model tiering is the second. Not every node needs the same model. A router classifies: Haiku. A builder reasons and calls tools: Sonnet. A final gate on high-stakes output: Opus, rarely. Using Sonnet to route is hiring a senior engineer to sort mail. It works, and you are burning budget on the wrong task.
Relative cost per node. Match the model to the job.
TIERS = {
"router": "claude-haiku-4-5", # fast, cheap, classify
"builder": "claude-sonnet-4-6", # reasoning, tools
"reviewer": "claude-sonnet-4-6", # strict, adversarial
"final_qa": "claude-opus-4-6", # highest bar, rare
}
# State contract, one line per node
# research_node: reads ["task"] writes ["research"]
# analysis_node: reads ["research"] writes ["analysis"]
# writer_node: reads ["analysis"] writes ["draft"]
# reviewer_node: reads ["draft"] writes ["review"]
10. Fallback paths
The happy path is not the graph
The happy path always works in testing. In production, any unexpected error takes the whole graph down. Every node that can fail needs a fallback edge.
A fallback is not "ignore the error". It is a separate path. If research times out, return a partial result and flag it. If the reviewer crashes, skip automatic review and route to a human. If the graph fails after max retries, save state to disk and alert someone. The user gets something useful instead of a stack trace.
def safe_node(func, state, error_key="error"):
try:
return func(state)
except Exception as e:
state[error_key] = str(e)
state["needs_human"] = True
return state
state = safe_node(research_node, state, "research_error")
if state.get("needs_human"):
notify_human(state)
else:
state = analyze_node(state)
11. A full working graph
Support tickets, end to end
Here is a graph that handles customer support tickets. Haiku classifies. A Sonnet specialist drafts a response. A Sonnet reviewer checks it. If it passes, the response ships. If not, the specialist retries with the reviewer's feedback. After max retries it routes to a human. Every pattern in this article, in about sixty lines, with no framework.
import anthropic, json
client = anthropic.Anthropic()
def classify(state):
r = client.messages.create(
model="claude-haiku-4-5", max_tokens=50,
system="Classify: billing, technical, general. One word.",
messages=[{"role": "user", "content": state["ticket"]}],
)
state["category"] = r.content[0].text.strip().lower()
return state
def draft(state):
prompts = {
"billing": "Handle billing. Be precise on refund policy.",
"technical": "Handle tech. Ask for logs first.",
"general": "Handle general inquiries. Keep it brief.",
}
r = client.messages.create(
model="claude-sonnet-4-6", max_tokens=1024,
system=prompts.get(state["category"], prompts["general"]),
messages=[{"role": "user", "content": state["ticket"]}],
)
state["draft"] = r.content[0].text
return state
def review(state):
r = client.messages.create(
model="claude-sonnet-4-6", max_tokens=200,
system='Find flaws in this support response. JSON: {"passed": bool, "issues": [str]}',
messages=[{"role": "user", "content": state["draft"]}],
)
state["review"] = json.loads(r.content[0].text)
return state
def run(ticket, max_retries=3):
state = {"ticket": ticket, "attempts": 0}
state = classify(state) # haiku router
for _ in range(max_retries):
state = draft(state) # sonnet specialist
state = review(state) # sonnet reviewer
state["attempts"] += 1
if state["review"]["passed"]:
send_response(state["draft"])
return state
state["needs_human"] = True # fallback
return state
12. Five ways graphs break
Seen in almost every first build
Six graphs to build this week
- Ticket triage. Classify by category and urgency, route to a specialist, review before sending.
- Competitive analysis. One sub-agent per competitor, all in parallel, one merge node that synthesises.
- PR review pipeline. Read the diff, run tests, write comments. A second agent screens for false positives before posting.
- Blog pipeline. Researcher gathers, writer drafts, editor checks facts and tone, loop until the editor passes.
- ETL with validation. Extract from PDFs, transform schema, load. A validator samples rows and flags anomalies before commit.
- Incident response. Alert fires the graph. Read logs, find root cause, draft a fix and a postmortem. Human approves the deploy.
A loop is an agent. A graph is a team.
Loop engineering was the breakthrough that turned prompting into building. Making one agent that retries, verifies, and improves is a real skill. It still is.
But every team outperforms every solo operator once the work gets complex. A graph is how you build a team of agents: a researcher who gathers, an analyst who reasons, a writer who drafts, a reviewer who rejects. Each agent stays simple. The graph makes them powerful.
None of the code here has a framework dependency. Functions, dictionaries, if-statements. That is a graph. You do not need LangGraph. You do not need CrewAI. You need the five patterns and the judgment to wire them for your case.
There are two kinds of builders now. The ones still running one loop for everything, and the ones designing the flow. The model is the same. The architecture is not.
// Steal the full map
Get the 10 repos and commands
Every graph in this article as a runnable repo, plus the router and reviewer prompts I actually use with clients. Drop your email and I will send the pack.