← Back to Agentic AI map
Lesson 3.8 · Building with LangGraph

Multi-Agent Graphs

Split one job between small agents - a researcher, a writer and a reviewer - and let a supervisor decide who works next. Real runs, two real bugs, and when one agent is the better choice.

langgraph

What you will be able to do

  • Explain what a multi-agent system is, using an everyday example
  • Say when several agents help, and when one agent is better
  • Build the supervisor pattern: workers around one supervisor node
  • Move between agents with Command(goto=...) and edges back to the supervisor
  • Share work through state, and keep private data inside a subgraph
  • Choose between routing rules in code and an LLM supervisor - and check what the LLM says
  • Stop a team that never finishes: your own limit first, recursion_limit as the safety net

The idea, in plain English

Until now, one agent did the whole job. It searched, it wrote, and it checked its own work. That works for small jobs. But when a job has several different parts, one agent with one long prompt starts to mix them up.

A multi-agent system splits the job. Each agent gets one small job and a short, clear prompt: a researcher finds facts, a writer writes, a reviewer checks. In LangGraph each agent is simply a node. Nothing new is needed - you already know nodes, state and edges.

Someone must decide who works next. That is the supervisor: a node that does no work itself. It looks at the state and sends the job to the right worker. When the worker finishes, it reports back to the supervisor, and the supervisor decides again. This shape - one supervisor in the middle, workers around it - is called the supervisor pattern.

In this lesson we build a three-agent team with llama3, run it, and look at what really happened - including two bugs we made while writing it. All code was run with LangGraph 1.2.14, langchain-core 1.6.9 and llama3 on Ollama.

Worked example: A researcher, a writer and a reviewer write a short paragraph about checkpoints. The reviewer sends the first draft back.

workflowA writing team with a supervisorstep 1 / 5

1 - No facts yet: the researcher

The state has only the task. The supervisor’s first rule says: no notes yet, so send the job to the researcher. The researcher finds 3 matching notes and reports back.

supervisor said
researcher
notes found
3 of 5
LLM calls
0
history
researcher

A real run. Task: "Explain what a checkpoint and a thread are, and how resume works." The reviewer sent the first draft back because it was 2 words too long.

Words you will see in this lesson

A few new words. Each one is simple - most of them are names for things you already built in Lessons 3.2 to 3.5.

Small dictionary
AgentHere: one node with one job and its own prompt (and maybe its own tools).
Multi-agent systemSeveral agents that together finish one job.
WorkerAn agent that does real work: research, writing, reviewing.
SupervisorA node that does no work. It only decides which worker goes next.
Shared stateThe state all agents read and write - like a whiteboard in an office.
HandoffMoving the job from one agent to another. Here: Command(goto=...).
SubgraphA whole compiled graph used as one node inside a bigger graph.

An everyday example: a newspaper office

Think of a small newspaper. A reporter collects facts. A writer turns the facts into an article. A checker reads the article and says "fine" or "fix this". And an editor decides who works next - the editor does not write anything.

Nobody does everything. Each person is good at one thing, and their notes go on a shared board so the next person can continue. If the checker says "too long", the editor gives the article back to the writer. That is exactly the graph in this lesson.

The office and the graph
Editorsupervisor - decides who works next, does no writing.
Reporterresearcher - collects facts.
Writerwriter - turns facts into text.
Checkerreviewer - says APPROVE or REVISE.
The shared boardthe graph state.
"Give it back to the writer"Command(goto="writer") with the review as feedback.

Why split the work - and what it costs

Small agents have real advantages. Each prompt is short and about one thing, so the model is less confused. Each agent can have its own tools - only the researcher needs search. You can test each agent alone. You can even give each agent a different model: a big one for writing, a small fast one for checking.

But a team is slower and more expensive. Every agent is at least one more model call. We gave the same task to one agent (one prompt that does everything) and to our three-agent team. The single agent was 3 times faster - but its paragraph had 69 words, over the 60-word limit, and nobody noticed. The team was slower, but its reviewer caught the long draft and got it fixed.

So: start with one agent. Split the job only when one agent clearly struggles, or when the parts really need different prompts, tools or checks.

Measured - the same task, one agent vs the team
one agent 1 llm call 5.9 s 69 words (limit 60 - nobody checked) team 3 llm calls 17.3 s 53 words (reviewer sent the 62-word draft back once)

Three common shapes

There are three common ways to connect agents. This lesson builds the supervisor, because it is the most flexible and the easiest to understand.

The pipeline is the simplest: fixed edges, A then B then C, like a chain. Use it when the order never changes. The supervisor adds a decision in the middle, so the order can change - for example, back to the writer after a bad review. In a network, agents hand the job directly to each other with no supervisor; it is powerful but hard to follow and hard to stop.

Which shape?
Pipelineresearcher -> writer -> reviewer, always. Plain add_edge. Order never changes.
SupervisorOne node in the middle picks the next worker each time. Order can change.
NetworkEach agent picks the next agent itself. Flexible, but hard to follow and to stop.

The shared state: the team’s whiteboard

The agents never call each other. They talk only through the state. The researcher writes notes, the writer reads notes and writes draft, the reviewer reads draft and writes review. The supervisor reads everything and decides.

Two fields are only for the team’s bookkeeping. revisions counts how many times we sent the draft back - this is how we stop. feedback carries the last review to the writer. history records who worked, so we can print the path at the end.

TeamState - one field per piece of work
class TeamState(TypedDict, total=False): task: str # what the user asked for notes: list[str] # written by the researcher draft: str # written by the writer review: str # written by the reviewer: "APPROVE" or "REVISE: why" feedback: str # the last review, handed to the writer for a rewrite revisions: int # how many times we sent the draft back history: list[str] # who worked, in order - only for us to read

The supervisor and Command(goto=...)

The supervisor is a normal node. Instead of returning a dict, it returns Command(goto=next_worker, update={...}). goto says which node runs next. update changes the state at the same time - here it hands the review to the writer as feedback and counts the rewrite.

The return type Command[Literal["researcher", "writer", "reviewer", "__end__"]] is not only for reading. LangGraph uses it to know where the supervisor can go, so it can draw the graph. In the drawing below, dotted lines (-.->) are the supervisor’s possible choices, and solid lines (-->) are the fixed edges from each worker back to the supervisor.

Our supervisor uses plain rules, checked from top to bottom: no notes -> researcher; no draft -> writer; no review -> reviewer; REVISE and fewer than 2 rewrites -> writer; otherwise END. Rules like these are fast, free and predictable.

What LangGraph drew: graph.get_graph().draw_mermaid()
__start__ --> supervisor; researcher --> supervisor; reviewer --> supervisor; writer --> supervisor; supervisor -.-> __end__; supervisor -.-> researcher; supervisor -.-> reviewer; supervisor -.-> writer;

Example 1 - the whole team, line by line

Here is the complete program. Read the comments - every line that matters has one. Then look at the output: it is a real run, printed exactly as it came out.

Example 1 - team.py
from typing import Literal, TypedDict from langchain_ollama import ChatOllama from langgraph.graph import END, START, StateGraph from langgraph.types import Command llm = ChatOllama(model="llama3", temperature=0) # the model every agent shares # A small "library". The researcher takes facts from here, not from the model's memory. NOTES = { "checkpoint": "A checkpointer saves the graph state after every step.", "thread": "Each conversation or run has a thread_id; the saved state is stored under it.", "resume": "With the same thread_id, a run can continue later, even after a restart.", "sqlite": "SqliteSaver keeps checkpoints in a file; InMemorySaver loses them when the process stops.", "weather": "It is sunny in Hyderabad today.", } # The shared state: the whiteboard every agent reads and writes. class TeamState(TypedDict, total=False): task: str # what the user asked for notes: list[str] # written by the researcher draft: str # written by the writer review: str # written by the reviewer: "APPROVE" or "REVISE: why" feedback: str # the last review, handed to the writer for a rewrite revisions: int # how many times we sent the draft back history: list[str] # who worked, in order - only for us to read # --- The three workers. Each one does ONE job. --------------------------- def researcher(state: TeamState): words = state["task"].lower() found = [text for topic, text in NOTES.items() if topic in words] # keep matching notes print(f" researcher: found {len(found)} notes") return {"notes": found, "history": state.get("history", []) + ["researcher"]} def writer(state: TeamState): feedback = "" if state.get("feedback"): # is this a rewrite? feedback = "\nThe reviewer said: " + state["feedback"] + "\nFix this in the new version." text = llm.invoke( "Write ONE short paragraph (max 60 words) for new developers.\n" "Use only these facts:\n- " + "\n- ".join(state["notes"]) + "\nTask: " + state["task"] + feedback + "\nReturn only the paragraph." ).content.strip() print(f" writer: draft of {len(text.split())} words") return {"draft": text, "history": state.get("history", []) + ["writer"]} def reviewer(state: TeamState): words = len(state["draft"].split()) if words > 60: # cheap checks in code come first verdict = f"REVISE: it has {words} words, the limit is 60" elif "thread_id" not in state["draft"]: verdict = "REVISE: it must mention thread_id" else: # then ask the model to judge answer = llm.invoke( "Does this paragraph only use these facts?\n" + "\n".join(state["notes"]) + "\n\nParagraph:\n" + state["draft"] + "\n\nAnswer with one word: APPROVE or REVISE." ).content.strip() verdict = "APPROVE" if answer.upper().startswith("APPROVE") else "REVISE: " + answer print(f" reviewer: {verdict}") return {"review": verdict, "history": state.get("history", []) + ["reviewer"]} # --- The supervisor. It does no work itself - it only decides who goes next. --- def supervisor(state: TeamState) -> Command[Literal["researcher", "writer", "reviewer", "__end__"]]: if not state.get("notes"): next_worker = "researcher" # 1. no facts yet elif not state.get("draft"): next_worker = "writer" # 2. facts, but no draft elif not state.get("review"): next_worker = "reviewer" # 3. a draft nobody checked elif state["review"].startswith("REVISE") and state.get("revisions", 0) < 2: next_worker = "writer" # 4. rejected: rewrite (max 2 times) else: next_worker = END # 5. approved, or out of rewrites print(f"supervisor -> {next_worker}") update = {} if next_worker == "writer" and state.get("review"): # Hand the review to the writer as feedback, count the rewrite, # and clear the old review so the new draft gets a fresh one. update = {"feedback": state["review"], "review": "", "revisions": state.get("revisions", 0) + 1} return Command(goto=next_worker, update=update) # go there, and change state # --- Build the graph: a star with the supervisor in the middle. --- builder = StateGraph(TeamState) builder.add_node("supervisor", supervisor) builder.add_node("researcher", researcher) builder.add_node("writer", writer) builder.add_node("reviewer", reviewer) builder.add_edge(START, "supervisor") for worker in ["researcher", "writer", "reviewer"]: builder.add_edge(worker, "supervisor") # every worker reports back graph = builder.compile() if __name__ == "__main__": result = graph.invoke( {"task": "Explain what a checkpoint and a thread are, and how resume works."}, {"recursion_limit": 20}, # the safety net (default is 10,007) ) print("\npath:", " -> ".join(result["history"])) print("rewrites:", result.get("revisions", 0), "| final review:", result["review"]) print("final draft:\n" + result["draft"])
Output - a real run (about 20 seconds)
supervisor -> researcher researcher: found 3 notes supervisor -> writer writer: draft of 62 words supervisor -> reviewer reviewer: REVISE: it has 62 words, the limit is 60 supervisor -> writer writer: draft of 53 words supervisor -> reviewer reviewer: APPROVE supervisor -> __end__ path: researcher -> writer -> reviewer -> writer -> reviewer rewrites: 1 | final review: APPROVE final draft: Here is the revised paragraph: A checkpoint is a saved state of the graph after every step, stored under a unique thread_id. This thread_id identifies each conversation or run. When a run is interrupted, it can be resumed later with the same thread_id, allowing the graph to pick up where it left off.

Look closely: the reviewer approved a mistake

Read the final draft again. It starts with "Here is the revised paragraph:". That line is llama3 talking to us, not part of the paragraph. A reader would see it. Our reviewer agent still said APPROVE. This is normal: a reviewer made of an LLM misses things, just like the writer does. So we added one more check in code.

The new rule caught the line - and then something new went wrong. The writer removed the line but wrote 62 words again. Why? It only received the latest feedback ("no introduction line") and forgot the earlier one ("the limit is 60"). After 2 rewrites the supervisor stopped, and the run ended with a draft that was NOT approved.

The fix: give the writer every review so far, not only the last one. With that change the team needed 2 rewrites, and the final 52-word paragraph was clean and approved.

Two more lessons from this. First, decide what happens when the team runs out of rewrites - here the run ends with review = "REVISE: ...", so your app must check it and maybe ask a human (Lesson 3.7). Second, notice that the weather note was never used: the researcher only keeps notes whose topic appears in the task. Choosing what each agent sees is part of designing the team.

Step 1 - one more rule in the reviewer
elif state["draft"].lower().startswith(("here is", "here's")): verdict = "REVISE: start with the paragraph itself, no introduction line"
Output - the rule works, but the writer forgets the word limit
writer: draft of 62 words reviewer: REVISE: it has 62 words, the limit is 60 writer: draft of 53 words reviewer: REVISE: start with the paragraph itself, no introduction line writer: draft of 62 words reviewer: REVISE: it has 62 words, the limit is 60 supervisor -> __end__ rewrites: 2 | final review: REVISE: it has 62 words, the limit is 60
Step 2 - keep every review, not only the last one
feedback: list[str] # in TeamState: every review so far # in the writer feedback = "\nThe reviewer said:\n- " + "\n- ".join(state["feedback"]) + "\nFix ALL of these in the new version." # in the supervisor update = {"feedback": state.get("feedback", []) + [state["review"]], "review": "", "revisions": state.get("revisions", 0) + 1}
Output - approved, and clean
writer: draft of 62 words reviewer: REVISE: it has 62 words, the limit is 60 writer: draft of 58 words reviewer: REVISE: start with the paragraph itself, no introduction line writer: draft of 52 words reviewer: APPROVE rewrites: 2 | final review: APPROVE final draft: A checkpoint saves the graph state after every step, storing it under a unique thread_id. This thread_id identifies each conversation or run, allowing a run to continue later even after a restart. When resuming, the saved state is retrieved and the run picks up where it left off, thanks to the thread_id.

Rules in code, or an LLM supervisor?

Many tutorials let an LLM be the supervisor: you describe the team in a prompt and ask "who should work next?". We tried that with llama3 on the five moments of our job. It chose right three times and wrong two times.

The two mistakes are dangerous. After "REVISE" it picked the reviewer again - the same draft would be reviewed again and again. After "APPROVE" it picked the writer - the finished job would start again. Both are loops with no exit.

Use rules when the next step can be read from the state, as in our team. Use an LLM supervisor only when choosing really needs judgment - for example, which of six specialists fits a customer’s question. Then always check its answer: lowercase it, strip punctuation, accept only names from your list, and fall back to a safe rule when the answer is not allowed.

Example 2 - asking llama3 to supervise, and checking its answer
from langchain_ollama import ChatOllama llm = ChatOllama(model="llama3", temperature=0) ALLOWED = {"researcher", "writer", "reviewer", "finish"} # the only valid answers PROMPT = """You manage a team: researcher (finds facts), writer (writes the paragraph), reviewer (checks the paragraph). Look at the progress and choose who works next. Progress: - notes found: {notes} - draft written: {draft} - review: {review} Answer with exactly one word from: researcher, writer, reviewer, FINISH.""" def ask_llm(notes, draft, review): raw = llm.invoke(PROMPT.format(notes=notes, draft=draft, review=review)).content.strip() word = raw.split()[0].strip(".,:*\"'").lower() if raw else "" # "Writer." -> "writer" return raw, (word if word in ALLOWED else None) # None = not allowed # Five moments in the job, and the right answer for each. cases = [ ("start", "no", "no", "none", "researcher"), ("notes ready", "yes", "no", "none", "writer"), ("draft ready", "yes", "yes", "none", "reviewer"), ("review: REVISE", "yes", "yes", "REVISE: too long", "writer"), ("review: APPROVE", "yes", "yes", "APPROVE", "finish"), ] for name, notes, draft, review, right in cases: raw, pick = ask_llm(notes, draft, review) print(f"{name:16} llama3: {raw!r:13} -> {pick!s:10} right: {right:10} {'OK' if pick == right else 'WRONG'}")
Output - 3 right, 2 wrong (the same on a second run)
start llama3: 'Researcher' -> researcher right: researcher OK notes ready llama3: 'Writer' -> writer right: writer OK draft ready llama3: 'Reviewer' -> reviewer right: reviewer OK review: REVISE llama3: 'Reviewer' -> reviewer right: writer WRONG review: APPROVE llama3: 'Writer' -> writer right: finish WRONG

Watch out: llama3 answered "Researcher" with a capital R. Your node is called "researcher". Without the .lower() step, that answer is a name that does not exist - and the next section shows what LangGraph does with it.

A wrong name is ignored - quietly

What happens if the supervisor sends the job to a node that does not exist - a typo, a capital letter, or a name the LLM invented? We tested it. LangGraph did not raise an error. It printed a warning, ignored the goto, and the run simply ended. Nothing was written. Nothing failed.

This is why you must check the supervisor’s answer against your list of workers. A silent stop is harder to notice than a crash.

Command(goto="editor") - but there is no editor node
def boss(state): return Command(goto="editor") # there is no node called "editor" builder.add_node("boss", boss) builder.add_node("writer", writer) builder.add_edge(START, "boss") print(builder.compile().invoke({"x": 0}))
Output
Task boss with path ('__pregel_pull', 'boss') wrote to unknown channel branch:to:editor, ignoring it. {'x': 0}

Our first version never stopped

Here is a real mistake we made while writing Example 1. In the first version, the supervisor cleared the review before sending the draft back - and the writer looked at the review to see if this was a rewrite. The review was already empty, so the writer never saw the feedback, and the rewrite counter never went up.

The result: writer, reviewer, writer, reviewer... forever. We stopped it by hand after 131 supervisor turns and about five minutes. LangGraph did not stop it, because the default recursion_limit is 10,007 steps.

Two lessons. First, your own limit (here: at most 2 rewrites) only works if the counter really changes - count in the supervisor, which always runs. Second, always pass a recursion_limit as a safety net. With recursion_limit=12, the same buggy graph stopped after 3 LLM calls with GraphRecursionError.

The bug, and the fix
# WRONG - the supervisor empties "review", then the writer checks "review" update = {"review": ""} # in the supervisor if state.get("review", "").startswith("REVISE"): ... # in the writer: always False # RIGHT - hand the review over in its own field, and count in the supervisor update = {"feedback": state["review"], "review": "", "revisions": state.get("revisions", 0) + 1}
Output - the buggy graph with recursion_limit=12
without a limit: 131 supervisor turns, about 5 minutes, stopped by hand with {"recursion_limit": 12}: GraphRecursionError: Recursion limit of 12 reached without hitting a stop condition. llm calls before it stopped: 3

An agent can be a whole graph

Sometimes one agent needs several steps of its own. The researcher might search widely, then keep only the useful notes. You can build those steps as a small graph, compile it, and add the compiled graph as one node. A graph used like this is called a subgraph.

The two graphs share data through keys with the same name. Below, task and notes exist in both, so the parent sends task in and gets notes back. The key found exists only in the researcher’s graph: it is the researcher’s private scratch paper, and the parent never sees it.

To see the inside of a subgraph while it runs, stream with subgraphs=True. Each update comes with a namespace: an empty () means the parent graph; ("researcher:...",) means inside the researcher. Lesson 3.9 explains streaming in detail.

Example 3 - the researcher as its own graph
from typing import TypedDict from langgraph.graph import END, START, StateGraph from team import NOTES, TeamState # The researcher as its own small graph: search widely, then keep the useful notes. class ResearchState(TypedDict, total=False): task: str # shared with the parent graph (same key name) found: dict # private to the researcher - the parent never sees it notes: list[str] # shared - this is the researcher's answer def search(state: ResearchState): return {"found": dict(NOTES)} # a wide search: every note, 5 of them def keep_relevant(state: ResearchState): words = state["task"].lower() return {"notes": [text for topic, text in state["found"].items() if topic in words]} r = StateGraph(ResearchState) r.add_node("search", search) r.add_node("keep_relevant", keep_relevant) r.add_edge(START, "search") r.add_edge("search", "keep_relevant") r.add_edge("keep_relevant", END) research_graph = r.compile() # Parent: the compiled researcher graph is just a node. p = StateGraph(TeamState) p.add_node("researcher", research_graph) p.add_edge(START, "researcher") p.add_edge("researcher", END) parent = p.compile() out = parent.invoke({"task": "Explain what a checkpoint and a thread are, and how resume works."}) print("parent keys:", sorted(out)) print("notes:", len(out["notes"]), "| 'found' in parent state:", "found" in out) for ns, chunk in parent.stream({"task": "checkpoint thread resume"}, stream_mode="updates", subgraphs=True): print(ns, list(chunk))
Output
parent keys: ['notes', 'task'] notes: 3 | 'found' in parent state: False ('researcher:a841cad7-3f1b-dfdc-9236-eb9a99f6bb46',) ['search'] ('researcher:a841cad7-3f1b-dfdc-9236-eb9a99f6bb46',) ['keep_relevant'] () ['researcher']

Tip: Our first keep_relevant dropped notes whose text contained the word "weather". The weather note says "It is sunny in Hyderabad today" - no word "weather" - so it slipped through. Filter on the topic you control, not on text you hope looks a certain way.

One agent or a team? A quick test

Before you build a team, answer these questions. If most answers are "no", one agent - maybe with a review step - is simpler, faster and cheaper.

Do you need several agents?
Do the parts need different instructions?Writing and checking need different prompts - yes, split.
Do the parts need different tools?Only the researcher searches - a good reason to split.
Must one part check another?A separate reviewer catches what the writer misses.
Is one prompt getting long and confused?Split it into focused agents.
Is speed or cost most important?Then stay with one agent: every agent adds model calls.

Multi-agent graphs at a glance

A worker

A normal node with one job.

builder.add_node("writer", writer)
Report back

Each worker returns to the supervisor.

builder.add_edge("writer", "supervisor")
Choose the next worker

Go there and update state in one return.

return Command(goto="writer", update={...})
Declare the choices

Lets LangGraph draw and check the routes.

-> Command[Literal["writer", "__end__"]]
Finish

Go to the end of the graph.

Command(goto=END)
An agent as a graph

A compiled graph added as a node.

builder.add_node("researcher", research_graph)
See inside subgraphs

Updates come with a namespace.

graph.stream(x, stream_mode="updates", subgraphs=True)
Safety net

Stop a team that never finishes.

graph.invoke(x, {"recursion_limit": 20})

Try it yourself

The code does not change. Swap the content string and the program does something else entirely.

Add the rule

“Add the "Here is" rule to the reviewer in Example 1. Run it with only the last feedback, then with every review. Compare the endings.”

Stricter limit

“Change the word limit to 40 in both the writer prompt and the reviewer. Does the team finish within 2 rewrites?”

A fourth agent

“Add a translator agent that turns the approved paragraph into Hindi or Telugu. Which supervisor rule do you need?”

LLM supervisor

“Replace the supervisor rules with Example 2. Keep the .lower() check and fall back to the rules when the answer is not allowed.”

Break it on purpose

“Make the supervisor return Command(goto="Writer") with a capital W. Read the warning, and look at what invoke() returns.”

What usually goes wrong

A rewrite loop with no working exit

Our own first version. The counter that should stop the loop never changed, and the default recursion_limit of 10,007 let it run for minutes. Count in the supervisor, and always set recursion_limit.

✗ update = {"review": ""}        # writer never sees REVISE, revisions stays 0
✓ update = {"feedback": state["review"], "review": "",
          "revisions": state.get("revisions", 0) + 1}
Trusting the LLM supervisor’s answer

llama3 chose wrong 2 times out of 5, and wrote "Researcher" with a capital R. A goto to an unknown name is ignored without an error. Normalize, check against your list, and fall back to a rule.

✗ return Command(goto=llm.invoke(prompt).content)
✓ pick = raw.split()[0].strip(".,:").lower()
if pick not in ALLOWED: pick = rule_based_next(state)
Giving the writer only the last review

In our run the writer fixed the newest problem and brought back an old one (62 words again). Keep a list of every review and send all of it.

✗ update = {"feedback": state["review"], ...}
✓ update = {"feedback": state.get("feedback", []) + [state["review"]], ...}
Letting an LLM review alone

Our reviewer approved a draft that began with "Here is the revised paragraph:". Put the checks you can write in code first, and ask the model only what code cannot check.

Agents calling each other directly

If the writer calls the reviewer function itself, the graph cannot see it, stream it, checkpoint it or limit it. Let the graph move the job: return to the supervisor, and let it decide.

Too many agents

Every agent is at least one more model call. Our team was 3 times slower than one agent. Split only when the parts really need different prompts, tools or checks.

Key points

  • A multi-agent system splits one job between agents; in LangGraph each agent is a node.
  • The supervisor does no work - it reads the state and picks the next worker.
  • Command(goto=..., update=...) moves the job and changes the state in one return.
  • Agents share work only through the state; a subgraph can keep private keys.
  • Use rules in code when the next step is clear; an LLM supervisor must be checked.
  • A goto to an unknown node is ignored quietly - validate names against your list.
  • Every loop needs your own limit that really counts, plus recursion_limit (default 10,007).
  • When the team runs out of rewrites, the result is not approved - check it, or ask a human.
  • A team is slower and costs more calls - start with one agent and split when needed.

Quick check before you move on

What does the supervisor do?
It does no work itself. It reads the state and decides which worker runs next, or ends the run.
How do the agents share their work?
Through the shared state. Each worker writes its result into a field, and the next agent reads it.
What does Command(goto="writer", update={...}) do?
It sends the run to the writer node next and changes the state at the same time.
Why did our first version never stop?
The rewrite counter never increased, so the "max 2 rewrites" rule never fired, and the default recursion_limit (10,007) is too high to stop it quickly.
When is one agent better than a team?
When the job does not need different prompts, tools or checks - one agent is faster and cheaper. In our test it was 3 times faster.

Quiz

  1. 1.

    The supervisor returns Command(goto="Reviewer") but the node is named "reviewer". What happens?

  2. 2.

    In our test, what did llama3 choose as supervisor after the review said APPROVE?

  3. 3.

    A subgraph has the keys task, found and notes. The parent has task and notes. What does the parent see after the subgraph runs?

  4. 4.

    Why does every worker have an edge back to the supervisor?

  5. 5.

    The reviewer is an LLM. How do you make it more reliable?

Interview questions

How would you build a supervisor-based multi-agent system in LangGraph?

Each agent is a node with one responsibility. A supervisor node reads the shared state and returns Command(goto=worker) with an optional state update; every worker has an edge back to the supervisor. The state holds each agent’s output plus bookkeeping such as a revision counter. Add a hard limit on loops and a recursion_limit as a safety net.

Rule-based routing or an LLM router - how do you choose?

If the next step follows from the state (no draft yet, review failed), use rules: they are free, fast and testable. Use an LLM when routing needs judgment, then constrain it - a fixed list or structured output - normalize the answer, validate it, and fall back to a safe default.

What are the costs of a multi-agent design?

More model calls, more latency, more state to design, and more ways to loop. In our measurement a three-agent team took 17.3 s and 3 calls against 5.9 s and 1 call for one agent. The benefit has to be worth it - focused prompts, separate tools, independent review.

How do you debug which agent did what?

Record a history in state, stream updates (with subgraphs=True for nested agents), keep a checkpointer to inspect each step, and trace with LangSmith (Module 5).

How do you keep agents from seeing data they should not?

Give each agent only the fields it needs in its prompt, and use subgraphs with their own state for private working data; only keys shared with the parent flow in and out.

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...