Most agents are one while-loop. It calls the model, appends the output, calls again, and keeps going until the context window fills and the answers turn to mush.

That loop was a real step forward. Instead of prompting once and copying the result, you built something that retried, verified, and improved with every pass. It is still the right answer for a single-purpose job.

It stops working the moment the task has more than one concern. Not because the model is weak: because the architecture is wrong. This is the roadmap from that loop to a graph.

A loop is one agent getting smarter. A graph is many agents getting coordinated.

01. Why one loop is not enough

One worker, five jobs

A loop gives one agent autonomy. It retries, it learns from failure, it improves with each pass. Fix this bug, summarise this document, clean this dataset: define the goal, attach a verifier, set a stop condition, let it run.

Trouble starts when the task carries multiple concerns. Research, analysis, code, review. One agent doing all four fills its window with mixed information, loses the thread of earlier steps, and ships work that looks finished and falls apart under inspection. The research bleeds into the analysis. The code ignores the review. Four jobs, one context budget.

The fix is not a bigger context window. It is splitting the work across agents that each do one thing, then wiring the output of one into the input of the next. It is the same move a solo developer makes when they join a team: roles, handoffs, reviews.

LOOP (ONE AGENT) research + analyze + code + review + write all in one context context drowns at step 3 ONE WORKER. FIVE JOBS. GRAPH (FOUR AGENTS) RESEARCH ANALYZE CODE REVIEW FOUR WORKERS. ONE JOB EACH.
Same models. Different architecture. Only one of them holds up past step three.

02. The four eras

Each one absorbed the last

Prompt engineering taught us to write the instruction. Context engineering taught us to feed the right data. Loop engineering taught us to build one autonomous agent. Graph engineering teaches us to wire many agents into a system.

Most people are still in era one or two: a prompt, some context, one call, and hope. The builders shipping real products are in three and four. They are not writing better prompts. They are designing better flows.

2023
Prompt
one call
2024
Context
one session
2025
Loop
one agent
2026
Graph
the system

03. The four building blocks

Node, edge, router, state

Every graph, however complex, is assembled from four primitives. You do not need a framework. You need functions, a dictionary, and if-statements.

Node. A function that does one thing. A model call, a tool execution, a query, a file write. State in, state out. It does not know what ran before it or what runs after.

Edge. A connection between two nodes. Unconditional, or conditional on something in state. Conditional edges are how a graph makes decisions.

Router. A node that inspects state and picks the path. It does no work itself. It classifies and dispatches.

State. A dictionary that flows through the whole graph. Every node reads from it and writes to it. Without it, each node starts from scratch and the graph has no memory.

THE FOUR PRIMITIVES NODE does one thing state in, state out EDGE IF? ROUTER PICKS THE PATH pass fail NODE B NODE C STATE { task, research, analysis, review_passed, final_output } shared
# The entire graph abstraction, in about 15 lines

def node(func):
    """A node takes state and returns state."""
    def wrapper(state):
        return func(state)
    return wrapper

def edge(state, next_node):
    """An edge passes state to the next node."""
    return next_node(state)

def router(state, condition, if_true, if_false):
    """A router picks a path based on state."""
    if condition(state):
        return if_true(state)
    return if_false(state)

04. Pattern 01: sequential chain

A to B to C, no branches

The simplest graph. Node A runs, hands off to B, which hands off to C. Most things people call an AI pipeline are exactly this.

It works when every step succeeds and the order never changes: extraction, format conversion, translate then summarise. The moment a step can fail or the path can vary, you need one of the other patterns.

The real risk here is error propagation. Bad output from A becomes the foundation of B, which becomes the foundation of C. Three layers later the error is baked in and invisible. This is why even a trivial chain earns a verification node at the end.

RESEARCH ANALYZE WRITE VERIFY never skip this one errors compound left to right

05. Pattern 02: router

Classification, not reasoning

A router inspects the input and picks one of several paths. Not every node runs. This is how one system handles many kinds of task without loading every tool into one agent.

The insight most people miss: routing is a classification task, not a reasoning task. You do not need a frontier model to decide whether a ticket is about billing or a bug. Haiku does it in milliseconds for a fraction of the cost. Spend the expensive model inside the specialist, where the actual work happens.

Every specialist gets its own system prompt and its own tools. A code agent needs run_tests and write_file. A research agent needs web_search and read_doc. Give every agent every tool and you get wrong tool calls, wasted tokens, and confused output.

TICKET IN ROUTE haiku, classify only BILLING TECHNICAL GENERAL
One cheap classifier in front. Three specialists behind, each with its own prompt and tools.
def router(state):
    category = call_claude(
        model="claude-haiku-4-5",           # cheap, fast, classify
        system="Classify this task. One word: code, research, or writing.",
        prompt=state["task"],
    )
    routes = {
        "code": code_agent,                 # own prompt, own tools
        "research": research_agent,
        "writing": writing_agent,
    }
    agent = routes.get(category.strip(), research_agent)
    return agent(state)

If your router makes bad calls, the fix is almost always a better prompt, not a bigger model. Write five or six real examples of each category into the router prompt. Few-shot on Haiku beats zero-shot on Sonnet for routing.

06. Pattern 03: parallel fan-out

The slowest agent, not the sum

Several nodes run at once on the same input. A collector waits for all of them and merges. Three agents in parallel means your total time is the slowest of the three, not the sum.

The constraint is independence. If B needs A's output, they cannot run together. If both work from the same input and produce separate outputs, they can. Competitive analysis is the clean case: one agent per competitor, all running at once.

The merge node is where almost everyone underinvests. Merging three JSON blobs is trivial. Merging three research reports into one coherent synthesis needs its own prompt and its own quality bar. Treat it as a real agent, not a string concatenation.

I built a market-research graph for a D2C brand. One agent was working through every competitor one at a time, and a full sweep took most of an afternoon. I split it: one agent per competitor, all running at once, then one merge node that wrote the actual comparison. The whole run collapsed to the time of the slowest single agent. The merge was the hard part. Three agents each said a different rival was the cheapest, and the merge had to reconcile that into one honest read.

TASK ACME researching GLOBEX researching INITECH researching MERGE synthesis prompt same time, not the sum

07. Pattern 04: loop with gate

The builder produces, the reviewer rejects

The builder does the work. A separate reviewer checks it. If the review fails, the builder retries with the failure reason appended to state. This is the pattern behind every self-correcting system.

The rule that makes it work: builder and reviewer must be different agents. Different prompts, often different model tiers. The builder's job is to produce. The reviewer's job is to reject. When one agent reviews its own work it approves its own mistakes, because the reasoning that produced the flaw is the reasoning evaluating it.

Write the reviewer prompt adversarially. Not "is this good", but "find every flaw, inconsistency, and missing edge case; if you find nothing, return passed true; if you are unsure, fail it." A strict reviewer catches real bugs. A polite one waves everything through.

On a claims-triage build for an insurer, the agent drafted the decision and signed off on its own work. It approved a claim it should have flagged, because the logic that made the call was the same logic checking it. I added a second agent with one job: reject. A different prompt, told to hunt for the reason to say no. The bad approvals stopped.

BUILDER sonnet, creative REVIEWER sonnet, adversarial PASS SHIP FAIL, RETRY WITH THE REASON two agents, two prompts, never one
def loop_with_gate(task, max_tries=5):
    state = {"task": task, "history": []}

    for i in range(max_tries):
        result = builder_agent(state)       # produces
        check  = reviewer_agent(result)     # rejects

        if check["passed"]:
            return result

        state["history"].append({
            "attempt": i + 1,
            "issues": check["issues"],
        })

    return {"status": "max_retries", "last_result": result}

08. Pattern 05: human in the loop

The graph pauses, the human decides

The graph stops at a node and waits for approval. The agent does the research and drafts the action. The human reviews the decision that matters. The agent executes after approval.

This is mandatory for anything expensive, irreversible, or high-stakes: customer emails, production deploys, financial transactions, deleting data. The graph handles the work. The human handles the judgment call.

Make the approval node show exactly what will happen and what it costs. Not "do you approve", but "this will send 2,400 emails to the billing segment with this subject line and this body, estimated cost 12 dollars, approve?" Specific enough to decide in five seconds.

AGENT drafts the action HUMAN graph pauses here EXECUTE action: send 2,400 emails segment: billing, overdue cost: $12.00 approve? [y/n] y

09. State flow and model tiering

Two production concerns

State is a flat dictionary every node reads and writes. Keep it explicit: each node declares which keys it reads and which it writes. When a graph breaks, state is the first thing you check. If you cannot trace which node wrote which key, you cannot debug anything.

Model tiering is the second. Not every node needs the same model. A router classifies: Haiku. A builder reasons and calls tools: Sonnet. A final gate on high-stakes output: Opus, rarely. Using Sonnet to route is hiring a senior engineer to sort mail. It works, and you are burning budget on the wrong task.

Haiku
route, classify
Sonnet
build, review
Opus
final qa, rare

Relative cost per node. Match the model to the job.

TIERS = {
    "router":   "claude-haiku-4-5",     # fast, cheap, classify
    "builder":  "claude-sonnet-4-6",    # reasoning, tools
    "reviewer": "claude-sonnet-4-6",    # strict, adversarial
    "final_qa": "claude-opus-4-6",      # highest bar, rare
}

# State contract, one line per node
# research_node:  reads ["task"]      writes ["research"]
# analysis_node:  reads ["research"]  writes ["analysis"]
# writer_node:    reads ["analysis"]  writes ["draft"]
# reviewer_node:  reads ["draft"]     writes ["review"]

10. Fallback paths

The happy path is not the graph

The happy path always works in testing. In production, any unexpected error takes the whole graph down. Every node that can fail needs a fallback edge.

A fallback is not "ignore the error". It is a separate path. If research times out, return a partial result and flag it. If the reviewer crashes, skip automatic review and route to a human. If the graph fails after max retries, save state to disk and alert someone. The user gets something useful instead of a stack trace.

RESEARCH ok ANALYZE WRITE on error FALLBACK, NOTIFY HUMAN
def safe_node(func, state, error_key="error"):
    try:
        return func(state)
    except Exception as e:
        state[error_key] = str(e)
        state["needs_human"] = True
        return state

state = safe_node(research_node, state, "research_error")
if state.get("needs_human"):
    notify_human(state)
else:
    state = analyze_node(state)

11. A full working graph

Support tickets, end to end

Here is a graph that handles customer support tickets. Haiku classifies. A Sonnet specialist drafts a response. A Sonnet reviewer checks it. If it passes, the response ships. If not, the specialist retries with the reviewer's feedback. After max retries it routes to a human. Every pattern in this article, in about sixty lines, with no framework.

CLASSIFY haiku DRAFT sonnet specialist REVIEW sonnet, adversarial pass SHIP FAIL, RETRY x3 HUMAN FALLBACK
import anthropic, json

client = anthropic.Anthropic()

def classify(state):
    r = client.messages.create(
        model="claude-haiku-4-5", max_tokens=50,
        system="Classify: billing, technical, general. One word.",
        messages=[{"role": "user", "content": state["ticket"]}],
    )
    state["category"] = r.content[0].text.strip().lower()
    return state

def draft(state):
    prompts = {
        "billing":   "Handle billing. Be precise on refund policy.",
        "technical": "Handle tech. Ask for logs first.",
        "general":   "Handle general inquiries. Keep it brief.",
    }
    r = client.messages.create(
        model="claude-sonnet-4-6", max_tokens=1024,
        system=prompts.get(state["category"], prompts["general"]),
        messages=[{"role": "user", "content": state["ticket"]}],
    )
    state["draft"] = r.content[0].text
    return state

def review(state):
    r = client.messages.create(
        model="claude-sonnet-4-6", max_tokens=200,
        system='Find flaws in this support response. JSON: {"passed": bool, "issues": [str]}',
        messages=[{"role": "user", "content": state["draft"]}],
    )
    state["review"] = json.loads(r.content[0].text)
    return state

def run(ticket, max_retries=3):
    state = {"ticket": ticket, "attempts": 0}
    state = classify(state)                  # haiku router

    for _ in range(max_retries):
        state = draft(state)                 # sonnet specialist
        state = review(state)                # sonnet reviewer
        state["attempts"] += 1

        if state["review"]["passed"]:
            send_response(state["draft"])
            return state

    state["needs_human"] = True              # fallback
    return state

12. Five ways graphs break

Seen in almost every first build

×
Giant monolith node
One node doing everything. If it fails, everything fails. Break it into single responsibilities.
×
No state between nodes
Each node starts from scratch. The graph forgets what the last node learned.
×
No fallback path
The happy path works. Any error crashes the graph. Always build the fallback edge.
×
Sonnet for routing
Routing is classification. Haiku is faster and cheaper. Spend Sonnet on building and reviewing.
×
Skipping the gate
No review node. The first wrong answer sent to a customer is the last time anyone trusts the system.

Six graphs to build this week

  1. Ticket triage. Classify by category and urgency, route to a specialist, review before sending.
  2. Competitive analysis. One sub-agent per competitor, all in parallel, one merge node that synthesises.
  3. PR review pipeline. Read the diff, run tests, write comments. A second agent screens for false positives before posting.
  4. Blog pipeline. Researcher gathers, writer drafts, editor checks facts and tone, loop until the editor passes.
  5. ETL with validation. Extract from PDFs, transform schema, load. A validator samples rows and flags anomalies before commit.
  6. Incident response. Alert fires the graph. Read logs, find root cause, draft a fix and a postmortem. Human approves the deploy.

A loop is an agent. A graph is a team.

Loop engineering was the breakthrough that turned prompting into building. Making one agent that retries, verifies, and improves is a real skill. It still is.

But every team outperforms every solo operator once the work gets complex. A graph is how you build a team of agents: a researcher who gathers, an analyst who reasons, a writer who drafts, a reviewer who rejects. Each agent stays simple. The graph makes them powerful.

None of the code here has a framework dependency. Functions, dictionaries, if-statements. That is a graph. You do not need LangGraph. You do not need CrewAI. You need the five patterns and the judgment to wire them for your case.

There are two kinds of builders now. The ones still running one loop for everything, and the ones designing the flow. The model is the same. The architecture is not.

// Steal the full map

Get the 10 repos and commands

Every graph in this article as a runnable repo, plus the router and reviewer prompts I actually use with clients. Drop your email and I will send the pack.