← Back to Agentic AI map
Lesson 3.9 · Building with LangGraph

Streaming

Show the user what is happening while the graph runs: progress messages, finished steps, and the answer word by word - in the console and in a browser. Measured with real runs.

langgraph

What you will be able to do

  • Explain why streaming makes a slow agent feel fast
  • Use stream() instead of invoke(), and read what it gives you
  • Choose a stream mode: updates, values, messages or custom
  • Send your own progress messages from inside a node with get_stream_writer()
  • Show the answer word by word, and know which node each word came from
  • Use several stream modes in one loop
  • Stream to a browser with FastAPI and Server-Sent Events
  • Avoid the common surprises: lazy streams, leaving the loop early, empty pieces

The idea, in plain English

An agent is slow. It searches, it thinks, it calls a model - and every model call takes seconds. If your app uses graph.invoke(), the user sees nothing until everything is finished. In our test that was 8.5 seconds of a blank screen. People think the app is broken.

Streaming means sending pieces of the result while the work is still going on. The user sees "searching docs...", then "searching blog...", then the answer appearing word by word. The total time is the same. But the user sees something at once, so the wait feels much shorter.

In LangGraph you stream by calling graph.stream() instead of graph.invoke(). It gives you pieces in a for loop. A stream mode tells LangGraph which pieces you want: finished steps, the whole state, the model’s words, or your own progress messages.

In this lesson we measure each stream mode on one small graph, print the answer live in the console, and then send it to a browser. All code was run with LangGraph 1.2.14, langchain-core 1.6.9, FastAPI and llama3 on Ollama.

Worked example: A search-then-answer graph: progress in 0.00 s, the first word at 3.6 s - instead of a blank screen for 8.5 s.

workflowWhat the user sees, and whenstep 1 / 5

1 - 0.00 s: progress at once

The search node sends "searching docs..." with get_stream_writer() before it has found anything. The screen is not blank.

custom
searching docs...
time
0.00 s
with invoke()
blank screen
graph state
question only

One run of the graph below, streamed. The times are real. With invoke() the user would see nothing until 8.47 s.

Words you will see in this lesson

Streaming has a few words of its own. Here they are, in plain English.

Small dictionary
StreamSend the result in pieces while the work is still going on.
ChunkOne piece that the stream gives you in one turn of the for loop.
TokenA small piece of text from the model - a word or part of a word.
Stream modeWhich pieces you want: updates, values, messages or custom.
GeneratorA Python object that makes values one by one, only when the loop asks.
SSEServer-Sent Events: a simple way for a web server to push text to a browser.

An everyday example: a food delivery app

You order food in an app. Imagine the app shows nothing for 40 minutes, and then the food arrives. Even if it is on time, those 40 minutes feel long and you start to worry.

Real delivery apps show each step: "order received", "cooking", "out for delivery", and a map that moves. The food does not come faster. But you always know what is happening, so the wait is fine.

Streaming does the same for your agent. "searching docs...", "searching blog...", then words appearing one by one. Same total time - a much better wait.

invoke() or stream()?

invoke() runs the whole graph and then returns the final state. Simple, and right for scripts, tests and background jobs where nobody is watching.

stream() returns a generator. You loop over it, and each turn of the loop gives you the next piece as soon as it is ready. Use it whenever a person is waiting for the result.

Below is the graph we use for every example in this lesson. The search node pretends to search three sources (0.5 seconds each) and sends a progress message for each. The answer node asks llama3 a question. Then we ran it once with invoke() and once in each stream mode, and printed the time each piece arrived.

Example 1 - stream_demo.py: one graph, every mode
import sys import time from typing import TypedDict from langchain_ollama import ChatOllama from langgraph.config import get_stream_writer from langgraph.graph import END, START, StateGraph llm = ChatOllama(model="llama3", temperature=0) class State(TypedDict, total=False): question: str facts: list[str] answer: str def search(state: State): write = get_stream_writer() # lets this node send its own messages facts = [] for source in ["docs", "blog", "forum"]: write({"progress": f"searching {source}..."}) # goes out at once, mid-node time.sleep(0.5) # pretend each source takes 0.5 s facts.append(f"fact from {source}") return {"facts": facts} def answer(state: State): # A normal invoke(). In stream_mode="messages" LangGraph still sends the words one by one. reply = llm.invoke("In two short sentences, explain what streaming means in a chat app.") return {"answer": reply.content} builder = StateGraph(State) builder.add_node("search", search) builder.add_node("answer", answer) builder.add_edge(START, "search") builder.add_edge("search", "answer") builder.add_edge("answer", END) graph = builder.compile() if __name__ == "__main__": mode = sys.argv[1] # invoke, updates, values, custom or messages start = time.perf_counter() clock = lambda: f"{time.perf_counter() - start:5.2f}s" question = {"question": "What is streaming?"} if mode == "invoke": result = graph.invoke(question) # waits for the whole run print(clock(), "invoke() returned:", sorted(result)) elif mode == "messages": count = 0 for token, meta in graph.stream(question, stream_mode="messages"): count += 1 if count <= 3: # print the first three pieces print(clock(), f"token from {meta['langgraph_node']}: {token.content!r}") print(clock(), f"last token - {count} pieces in total") else: for chunk in graph.stream(question, stream_mode=mode): if mode == "custom": print(clock(), chunk) elif mode == "updates": print(clock(), {node: sorted(change) for node, change in chunk.items()}) else: # values: the whole state each time print(clock(), "state keys:", sorted(chunk))
Output - python stream_demo.py invoke
$ python stream_demo.py invoke 8.47s invoke() returned: ['answer', 'facts', 'question']

Tip: With invoke(), the first and only thing the user gets arrives at 8.47 s. Keep that number in mind while you read the four modes below.

Mode 1 - "updates": what each step changed

In "updates" mode you get one chunk each time a node finishes. The chunk is a dict: the node’s name, and what that node returned. Not the whole state - only the change.

This is the most useful mode for showing progress: "search finished", "answer finished". It is also how you see a human-in-the-loop pause (Lesson 3.7): it arrives as a chunk with the key "__interrupt__".

Output - python stream_demo.py updates
1.51s {'search': ['facts']} 6.92s {'answer': ['answer']}

Mode 2 - "values": the whole state after each step

In "values" mode you get the complete state: once at the start, and again after every step. Look at the keys growing: first only question, then facts is added, then answer.

Use it when your screen shows the whole state - for example a debug panel. For big states it sends a lot of data again and again, so "updates" is usually lighter.

Output - python stream_demo.py values
0.00s state keys: ['question'] 1.51s state keys: ['facts', 'question'] 7.06s state keys: ['answer', 'facts', 'question']

Mode 3 - "messages": the answer word by word

In "messages" mode you get the model’s output in small pieces, while it is still writing. Each chunk is a pair: (token, metadata). token.content is the text piece. metadata["langgraph_node"] tells you which node is talking.

Something surprising: our answer node calls llm.invoke(), not llm.stream(). LangGraph still sent the words one by one. When you stream a graph in "messages" mode, LangGraph listens to the chat model inside the node and passes on each piece. You do not have to change the node.

The first word arrived at 3.64 s; the last at 7.16 s. The user can start reading 3.5 seconds before the answer is complete.

Output - python stream_demo.py messages
3.64s token from answer: 'In' 3.69s token from answer: ' a' 3.74s token from answer: ' chat' 7.16s last token - 67 pieces in total

Watch out: Not every piece has text. In our run the last 2 of the 67 pieces were empty. Skip pieces where token.content is empty before you send them to a screen.

Mode 4 - "custom": your own progress messages

The other modes only send something when a step finishes or the model writes. But our search node is busy for 1.5 seconds before either happens. To show progress inside a node, use get_stream_writer().

Call write = get_stream_writer() at the top of the node. Then call write({...}) whenever you want. The value goes out at once, in "custom" mode, while the node keeps working. Our three messages arrived at 0.00, 0.51 and 1.01 seconds - long before search finished.

Output - python stream_demo.py custom
0.00s {'progress': 'searching docs...'} 0.51s {'progress': 'searching blog...'} 1.01s {'progress': 'searching forum...'}

Which mode should I use?

You can use one mode, or several at once (next section). Most chat apps use "messages" for the answer, plus "custom" or "updates" for progress.

The four stream modes
"updates"After each node: {node: what it returned}. Progress, and __interrupt__ pauses.
"values"After each node: the whole state. Debug panels, small states.
"messages"While the model writes: (token, metadata). Show the answer word by word.
"custom"Whenever the node calls write(...). Progress from inside a long node.

Several modes at once: live console output

Pass a list: stream_mode=["custom", "messages"]. Now every chunk is a pair (mode, chunk), so you know which kind of piece you got.

To show the answer live in a terminal, print each token with end="" (stay on the same line) and flush=True (show it now, do not wait). Without flush=True, Python may hold the text back and print it all at once - and the effect is lost.

Example 2 - live.py
from stream_demo import graph # Live console output: show progress lines and the answer as it is written. for mode, chunk in graph.stream({"question": "What is streaming?"}, stream_mode=["custom", "messages"]): if mode == "custom": print("..." + chunk["progress"]) # e.g. ...searching docs... else: token, meta = chunk print(token.content, end="", flush=True) # same line, shown at once print()
Output - the three lines appear first, then the answer grows word by word
...searching docs... ...searching blog... ...searching forum... In a chat app, streaming refers to the ability to share a live video or audio feed with others in real-time, allowing for interactive and immersive communication. This can include features like live video conferencing, screen sharing, or even live audio broadcasts, enabling users to engage with each other in a more dynamic and spontaneous way.

Tip: llama3 explained the other meaning of "streaming" - live video. Streaming changes when the words arrive, not what they say. A clearer prompt fixes the answer; that is a prompting problem, not a streaming one.

Things that surprise people - tested

We tested five behaviours that often confuse people, on a graph with two LLM nodes (title and body) and a human approval step from Lesson 3.7.

First: stream() is lazy. Calling it does nothing. The graph only starts when you begin the for loop. Second: if you leave the loop early with break, the run stops - only the title node ran. Third: when two nodes use the model, metadata["langgraph_node"] tells you which node each piece came from, so you can show only the body. Fourth: a pause for a human arrives as an "__interrupt__" chunk, and you resume with Command(resume=...) as in Lesson 3.7. Fifth: in async code - a web server - use astream() with async for; it gives the same chunks.

Example 3 - surprises.py
import asyncio from typing import TypedDict from langchain_ollama import ChatOllama from langgraph.checkpoint.memory import InMemorySaver from langgraph.graph import END, START, StateGraph from langgraph.types import Command, interrupt llm = ChatOllama(model="llama3", temperature=0) RAN = [] # which nodes really ran class State(TypedDict, total=False): topic: str title: str body: str ok: str def title(state: State): # LLM node 1 RAN.append("title") return {"title": llm.invoke(f"Give a 4-word title about {state['topic']}. Only the title.").content} def body(state: State): # LLM node 2 RAN.append("body") return {"body": llm.invoke(f"One sentence about {state['topic']}.").content} def approve(state: State): # waits for a human (Lesson 3.7) RAN.append("approve") return {"ok": interrupt({"title": state["title"]})} builder = StateGraph(State) builder.add_node("title", title) builder.add_node("body", body) builder.add_node("approve", approve) builder.add_edge(START, "title") builder.add_edge("title", "body") builder.add_edge("body", "approve") builder.add_edge("approve", END) graph = builder.compile() # 1. stream() is lazy: nothing runs until you loop over it. RAN.clear() stream = graph.stream({"topic": "tea"}, stream_mode="updates") print("1. stream() called, no loop yet. Nodes that ran:", RAN) # 2. Leaving the loop early stops the run. RAN.clear() for chunk in graph.stream({"topic": "tea"}, stream_mode="updates"): break print("2. break after the first update. Nodes that ran:", RAN) # 3. Two LLM nodes stream tokens. metadata["langgraph_node"] says which node each came from. pieces = {} for token, meta in graph.stream({"topic": "tea"}, stream_mode="messages"): pieces.setdefault(meta["langgraph_node"], []).append(token.content) for node, parts in pieces.items(): print(f"3. {node}: {len(parts)} pieces -> {''.join(parts).strip()[:55]!r}") # 4. A pause for a human arrives in the stream as "__interrupt__". saved = builder.compile(checkpointer=InMemorySaver()) config = {"configurable": {"thread_id": "post-1"}} for chunk in saved.stream({"topic": "tea"}, config, stream_mode="updates"): print("4.", list(chunk)) for chunk in saved.stream(Command(resume="yes"), config, stream_mode="updates"): print("4. after resume:", chunk) # 5. In async code (a web server), use astream with "async for". async def main(): async for chunk in graph.astream({"topic": "tea"}, stream_mode="updates"): print("5. astream:", list(chunk)) asyncio.run(main())
Output
1. stream() called, no loop yet. Nodes that ran: [] 2. break after the first update. Nodes that ran: ['title'] 3. title: 10 pieces -> '"Steeped in Serenity"' 3. body: 36 pieces -> 'Tea is a popular beverage made from the leaves of the C' 4. ['title'] 4. ['body'] 4. ['__interrupt__'] 4. after resume: {'approve': {'ok': 'yes'}} 5. astream: ['title'] 5. astream: ['body'] 5. astream: ['__interrupt__']

Streaming to a browser

A console is good for learning. A real app streams to a browser. The simplest way is Server-Sent Events (SSE): the browser opens one normal HTTP request, and the server keeps it open and writes small text messages into it, one after another.

Each message is two lines - "event: name" and "data: value" - followed by an empty line. In the browser, the built-in EventSource object reads them. In FastAPI, a StreamingResponse with media_type="text/event-stream" sends them.

Inside the endpoint we use astream() (the async version, because FastAPI is async) with three modes: custom for progress, updates for finished steps, and messages for the words. Each piece becomes one SSE message. At the end we send a "done" event so the browser knows to stop.

Example 4 - server.py (FastAPI)
import json from fastapi import FastAPI from fastapi.responses import StreamingResponse from stream_demo import graph # the graph from Example 1 app = FastAPI() def sse(event: str, data) -> str: # One Server-Sent Event: an "event:" line, a "data:" line, then an empty line. return f"event: {event}\ndata: {json.dumps(data)}\n\n" @app.get("/ask") async def ask(q: str): async def events(): modes = ["custom", "updates", "messages"] # progress, finished steps, tokens async for mode, chunk in graph.astream({"question": q}, stream_mode=modes): if mode == "custom": yield sse("progress", chunk["progress"]) elif mode == "updates": yield sse("step", list(chunk)) # e.g. ["search"] else: token, meta = chunk if token.content: # skip empty pieces yield sse("token", token.content) yield sse("done", "") # tell the browser we finished return StreamingResponse(events(), media_type="text/event-stream")
Run it, and look at the raw stream with curl -N (no buffering)
$ uvicorn server:app --port 8765 $ curl -N "localhost:8765/ask?q=hi" event: progress data: "searching docs..." event: progress data: "searching blog..." event: progress data: "searching forum..." ...
Output - a Python client printing when each event arrived
content-type: text/event-stream; charset=utf-8 0.03s progress "searching docs..." 0.54s progress "searching blog..." 1.04s progress "searching forum..." 1.55s step ["search"] 3.36s token "In" 3.41s token " a" 3.47s token " chat" 6.94s step ["answer"] 6.94s done "" tokens received: 65
In the browser - EventSource reads the events
const source = new EventSource("/ask?q=" + encodeURIComponent(question)); source.addEventListener("progress", (e) => status.textContent = JSON.parse(e.data)); source.addEventListener("token", (e) => answer.textContent += JSON.parse(e.data)); source.addEventListener("done", () => source.close()); // stop, or it reconnects

Watch out: Call source.close() when you receive "done". EventSource reconnects automatically when a stream ends. We tested it in Chrome: without close() the page asked the same question 3 times in 25 seconds - and each time the whole graph ran again. With close(), once.

Streaming a multi-agent graph

In Lesson 3.8 the researcher was a subgraph. By default, stream() shows only the parent graph’s steps. Add subgraphs=True to see inside: each chunk becomes a pair (namespace, chunk). An empty namespace () is the parent; ("researcher:...",) is inside the researcher.

For a team, "updates" mode is a ready-made activity feed: "researcher finished", "writer finished", "reviewer finished". Users like to see who is working.

From Lesson 3.8 - updates with subgraphs=True
('researcher:a841cad7-3f1b-dfdc-9236-eb9a99f6bb46',) ['search'] ('researcher:a841cad7-3f1b-dfdc-9236-eb9a99f6bb46',) ['keep_relevant'] () ['researcher']

Good streaming, for real users

Show words a person understands. "Searching the docs..." is good; "node search_v2 started" is not. Send progress with the custom writer and choose the words yourself.

Be careful what you stream. "messages" mode sends everything every model in the graph writes - including an internal reviewer or a supervisor’s reasoning. Filter on metadata["langgraph_node"] and send only the nodes the user should see.

Remember that streaming does not make the work faster. Our streamed run and our invoke() run did the same work. Streaming only changes when the user sees it. If the total is too slow, you still need fewer or faster model calls.

Streaming at a glance

Stream a run

A for loop instead of one result.

for chunk in graph.stream(inputs, stream_mode="updates"):
Step results

After each node: {node: change}.

stream_mode="updates"
Whole state

After each node: the full state.

stream_mode="values"
Words

(token, metadata) while the model writes.

stream_mode="messages"
Own progress

Send from inside a node.

write = get_stream_writer(); write({"progress": "..."})
Which node?

The node a token came from.

metadata["langgraph_node"]
Several modes

Chunks become (mode, chunk).

stream_mode=["custom", "messages"]
Async

For web servers.

async for chunk in graph.astream(...):
Inside subgraphs

Chunks become (namespace, chunk).

graph.stream(x, subgraphs=True)

Try it yourself

The code does not change. Swap the content string and the program does something else entirely.

Feel the difference

“Run Example 1 with invoke, then with messages. Watch the screen, not the numbers: which one feels faster?”

Better progress

“Add a fourth source to search and a "found 3 facts" message at the end of the node.”

Only the body

“In Example 3, print only the tokens from the body node, live, with end="" and flush=True.”

Stop early

“Break out of the messages loop after 10 pieces. Did the answer node finish? Check with a print inside the node.”

A web page

“Run Example 4 and write a tiny HTML page with EventSource that shows progress above the answer.”

What usually goes wrong

Expecting stream() to start the run

stream() returns a generator. Nothing runs until you loop over it. In our test, calling stream() alone ran zero nodes.

✗ graph.stream(inputs)          # nothing happens
✓ for chunk in graph.stream(inputs, stream_mode="updates"):
    show(chunk)
Waiting for words in "updates" mode

"updates" sends one chunk when a node finishes - the whole answer at once. For words as they are written, use "messages".

✗ graph.stream(inputs, stream_mode="updates")   # answer arrives at 6.92 s, all at once
✓ graph.stream(inputs, stream_mode="messages")  # first word at 3.64 s
Forgetting flush=True in the console

Python may hold printed text in a buffer. Print tokens with end="" and flush=True so they appear at once.

✗ print(token.content, end="")
✓ print(token.content, end="", flush=True)
Sending empty pieces

The last 2 of our 67 pieces had no text. Check token.content before sending.

✗ yield sse("token", token.content)
✓ if token.content:
    yield sse("token", token.content)
Streaming every node to the user

"messages" mode includes every model call in the graph - reviewers, supervisors, internal steps. Filter by metadata["langgraph_node"].

Leaving the loop early by accident

A break or an exception inside the for loop stops the run. In our test, break after the first update meant only the first node ran.

Key points

  • Streaming sends pieces while the graph runs; the work is the same, the wait feels shorter.
  • graph.stream() gives a generator - nothing runs until you loop over it.
  • "updates" = what each node returned; "values" = the whole state after each node.
  • "messages" = (token, metadata) while the model writes - even if the node calls llm.invoke().
  • "custom" = your own messages from inside a node, with get_stream_writer().
  • Several modes at once give (mode, chunk); subgraphs=True gives (namespace, chunk).
  • A human-in-the-loop pause arrives as an "__interrupt__" chunk in "updates" mode.
  • For browsers: astream() + Server-Sent Events, skip empty tokens, send a "done" event.

Quick check before you move on

What is the difference between invoke() and stream()?
invoke() waits and returns the final state. stream() gives pieces in a for loop while the graph runs.
Which mode shows the answer word by word?
"messages". Each chunk is (token, metadata).
How do you send "searching docs..." from inside a node?
write = get_stream_writer(), then write({"progress": "searching docs..."}). It arrives in "custom" mode.
Two nodes call the model. How do you show only one node’s words?
Check metadata["langgraph_node"] for each token and keep only that node.
Does streaming make the graph finish sooner?
No. The same work is done. The user just sees pieces earlier.

Quiz

  1. 1.

    In our test, when did the user first see something with invoke(), and with streaming?

  2. 2.

    You stream in "updates" mode and the graph pauses with interrupt(). What chunk do you get?

  3. 3.

    What do you get from stream_mode=["custom", "messages"]?

  4. 4.

    Your SSE endpoint ends, and the browser asks the question again. Why?

  5. 5.

    Why does the answer node stream tokens even though it calls llm.invoke()?

Interview questions

Explain the LangGraph stream modes and when you would use each.

updates gives each node’s returned changes - good for progress and for catching interrupts. values gives the full state after each step - good for debugging or small states. messages gives LLM tokens with metadata - for typing-style output. custom gives arbitrary data written from inside nodes with get_stream_writer - for fine-grained progress. Modes can be combined, yielding (mode, chunk) pairs.

How would you stream a LangGraph agent to a web frontend?

An async endpoint calls graph.astream with the needed modes and turns each chunk into a Server-Sent Event (or WebSocket message): progress, tokens filtered by node, step completions, and a final done event. Return it as text/event-stream; the browser uses EventSource and closes it on done.

How do you avoid leaking internal reasoning through streaming?

messages mode includes every model call, so filter by metadata["langgraph_node"] (or tags) and forward only user-facing nodes. Send progress through custom events whose text you control.

What happens if the client disconnects in the middle of a stream?

The loop over the stream stops, and with it the run - in our test, leaving the loop early meant later nodes never ran. If the work must finish anyway, run it in the background with a checkpointer and let the client reconnect to the thread.

Does streaming reduce latency?

It reduces time to first output, not total time. In our run the first word arrived at 3.6 s instead of everything at 8.5 s; the total work was unchanged.

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...