← Back to Agentic AI map
Lesson 3.6 · Building with LangGraph

Persistence & Checkpointing

Give a graph a checkpointer and a thread_id, and its state survives between calls, across a crash, and - with a database - across a restart.

langgraph

What you will be able to do

  • Explain what a checkpoint is and when LangGraph saves one
  • Compile a graph with a checkpointer and run it on a thread
  • Read a thread back with get_state and get_state_history
  • Resume a crashed run with invoke(None, config) - in the same process and in a new one
  • Choose between InMemorySaver and a database-backed checkpointer
  • Treat thread IDs and checkpoint data as application data: scoped, authorized, retained deliberately

The idea, in plain English

Until now every invoke() started from nothing. The state lived in the call and disappeared with it: a second question did not know about the first, and a crash halfway through a five-step research run meant starting over at step one.

A checkpointer fixes that. Compile the graph with one and LangGraph saves a checkpoint after every step - the state values, plus where execution is: which node runs next, and any error from the step that failed. Like a save point in a game, it is enough to look at the run later or carry on from it.

Checkpoints are grouped by thread. You pass a thread_id in the config on every call; each thread is one independent conversation or workflow with its own history. The same compiled graph serves every thread - the definition is shared, the state is not.

Where checkpoints live is the checkpointer’s choice. InMemorySaver keeps them in a Python dict, gone when the process exits. SqliteSaver and PostgresSaver keep them in a database, so a new process can load the thread and continue. We ran everything below on LangGraph 1.2.14 with langgraph-checkpoint-sqlite 3.1.1; PostgresSaver (langgraph-checkpoint-postgres 3.1.2) has the same interface - we checked its API, but had no Postgres server running, so that block was not run.

Worked example: A research graph that crashes in search, then finishes in a new process without re-running plan.

request flowCrash in one process, resume in anotherstep 1 / 5

1 - The run starts on a thread

Process 1 calls invoke({"topic": "LangGraph"}, config) with thread_id research-42. The input itself is saved as the first checkpoint.

thread_id
research-42
checkpoint
step -1, input
next
__start__
process
1

A real run: plan -> search -> write on thread research-42, with SqliteSaver. search fails in process 1; process 2 finishes the job.

State, checkpoint, thread

State is the data flowing through the graph right now. A checkpoint is a saved snapshot of a run: the state values plus the execution position - the next node, the step number, the error of a step that failed. A thread is the ID that groups a run’s checkpoints together, so the next call on that thread starts from the latest one.

get_state(config) returns the latest checkpoint of a thread as a StateSnapshot: .values, .next and, after a crash, .tasks with the error. get_state_history(config) returns every checkpoint, newest first.

Three words to keep apart
StateWhat is happening - the values in the graph.
Thread IDWhich conversation or workflow - the key the checkpoints are filed under.
CheckpointWhere it got to - a saved snapshot of state and position after a step.

When checkpoints are saved

One per step, not just one at the end. A two-node graph a -> b left four checkpoints on its thread: step -1 (the input, next __start__), step 0 (next a), step 1 (count 1, next b), step 2 (count 2, next nothing). That is why a crash in the middle loses only the step that failed.

A step that crashes saves no new checkpoint. When search raised, the newest checkpoint was still step 1 with next ('search',), and its tasks recorded ConnectionError('search service down').

Resume, or start again

To continue a thread from where it stopped, invoke it with None as the input: graph.invoke(None, config). In our run search ran again and write followed; plan ran once in total.

Passing input instead starts a new run on the same thread. Re-invoking the crashed research thread with {"topic": "LangGraph"} ran plan a second time. The same rule surprises people with the counter example: invoke({"count": 0}) on a thread that already holds 1 returns 1 again - the input overwrote the saved count. To build on what the thread holds, send only what is new: invoke({}, config) returned the saved count plus the graph’s work, and in a chat you send only the new message.

Watch out: Resume works per node: the node that failed runs again from its first line. If it sent an email or charged a card before crashing, it will do so twice. Keep side effects idempotent, or split them into their own node.

Conversations on a thread

With MessagesState the messages field appends, so every invoke on the same thread adds to the conversation the checkpointer already holds. We ran llama3 through a one-node chat graph. On thread chat-123: "My order number is ORD1001." then "What is my order number?" - the answer was "Your order number is ORD1001." with four messages in state. The same question on chat-456 got "Can you please provide me with your email address or order details...", because that thread had only its own message.

That is short-term memory: the history of one conversation, restored by thread. Long-term memory - facts about a user that outlive any one conversation, such as "prefers Telugu" - is a separate store, covered later in the course.

In memory or in a database

InMemorySaver keeps checkpoints in the Python process. We ran demo-1 in one process (count 1), then a second process with an identical graph: get_state returned {}. The same is true for workers behind a load balancer - each has its own memory, so a thread started on worker 1 is empty on worker 2.

SqliteSaver writes to a file - two tables, checkpoints and writes - which is why our process 2 found the crashed run. It suits a single machine. For a service with several workers, PostgresSaver (package langgraph-checkpoint-postgres) stores the same checkpoints in a shared database; call setup() once to create its tables.

Checkpointers compared
InMemorySaverLearning, tests, notebooks. Lost when the process exits; not shared between workers.
SqliteSaverA local file. Survives a restart; one machine.
PostgresSaverA shared database. Survives restarts and deploys; every worker sees every thread.

Thread IDs are not security

Any caller who passes a thread_id gets that thread’s state - LangGraph does not know who owns it. An unknown thread_id is not an error either: get_state on "nobody" returned {}, so a typo silently starts an empty conversation.

Generate thread IDs on the server, record which tenant and user own each one, and check ownership before you build the config. Checkpoints hold everything the state holds - user messages, tool results, personal data - so treat the store like any application database: access control, encryption, backups, and a retention policy. delete_thread(thread_id) on the checkpointer removes a thread; after it, get_state returned {}.

Why this comes before human approval

A run that can stop and be continued later is exactly what human-in-the-loop needs: draft the email, save a checkpoint, wait - minutes or days - for approval, then resume and send. Lesson 3.7 builds that pause on top of the checkpointer from this lesson.

Step-by-step code

A graph with a checkpointer
from typing import TypedDict from langgraph.checkpoint.memory import InMemorySaver from langgraph.graph import END, START, StateGraph class State(TypedDict): count: int def increment(state: State): return {"count": state["count"] + 1} builder = StateGraph(State) builder.add_node("increment", increment) builder.add_edge(START, "increment") builder.add_edge("increment", END) graph = builder.compile(checkpointer=InMemorySaver()) # the one change demo1 = {"configurable": {"thread_id": "demo-1"}} demo2 = {"configurable": {"thread_id": "demo-2"}} print(graph.invoke({"count": 0}, demo1)) print(graph.invoke({"count": 0}, demo1)) # input overwrites the saved count print(graph.invoke({"count": 5}, demo1)) print(graph.invoke({"count": 100}, demo2)) snapshot = graph.get_state(demo1) print(snapshot.values, snapshot.next) print(len(list(graph.get_state_history(demo1))), "checkpoints on demo-1") graph.invoke({"count": 0}) # no thread_id
Output - a real run
{'count': 1} {'count': 1} {'count': 6} {'count': 101} {'count': 6} () 9 checkpoints on demo-1 (3 runs x input, start, increment) ValueError: Checkpointer requires one or more of the following 'configurable' keys: thread_id, checkpoint_ns, checkpoint_id graph.get_state({"configurable": {"thread_id": "nobody"}}).values -> {}
Crash, then resume in a new process
import sqlite3 import sys from typing import TypedDict from langgraph.checkpoint.sqlite import SqliteSaver from langgraph.graph import END, START, StateGraph SEARCH_IS_DOWN = sys.argv[1] == "crash" class Research(TypedDict, total=False): topic: str plan: str notes: str report: str def plan(state): print(" running plan") return {"plan": f"plan for {state['topic']}"} def search(state): print(" running search") if SEARCH_IS_DOWN: raise ConnectionError("search service down") return {"notes": "3 sources found"} def write(state): print(" running write") return {"report": f"{state['plan']} + {state['notes']}"} builder = StateGraph(Research) builder.add_node("plan", plan) builder.add_node("search", search) builder.add_node("write", write) builder.add_edge(START, "plan") builder.add_edge("plan", "search") builder.add_edge("search", "write") builder.add_edge("write", END) conn = sqlite3.connect("checkpoints.db", check_same_thread=False) graph = builder.compile(checkpointer=SqliteSaver(conn)) config = {"configurable": {"thread_id": "research-42"}} if SEARCH_IS_DOWN: try: graph.invoke({"topic": "LangGraph"}, config) except ConnectionError as error: print(" crashed:", error) print(" next:", graph.get_state(config).next) else: state = graph.get_state(config) print(" found:", state.values, "next:", state.next) print(" result:", graph.invoke(None, config)["report"]) # None = resume
Output - two separate processes
$ python research.py crash running plan running search crashed: search service down next: ('search',) $ python research.py resume found: {'topic': 'LangGraph', 'plan': 'plan for LangGraph'} next: ('search',) running search running write result: plan for LangGraph + 3 sources found checkpoints.db tables: checkpoints, writes The same experiment with InMemorySaver: process 1: {'count': 1} process 2 get_state: {}
A conversation that remembers - llama3
from langchain_ollama import ChatOllama from langgraph.checkpoint.memory import InMemorySaver from langgraph.graph import END, START, MessagesState, StateGraph llm = ChatOllama(model="llama3", temperature=0) def chat(state: MessagesState): system = ("system", "You are a support agent. Answer in one short sentence.") return {"messages": [llm.invoke([system] + state["messages"])]} builder = StateGraph(MessagesState) builder.add_node("chat", chat) builder.add_edge(START, "chat") builder.add_edge("chat", END) app = builder.compile(checkpointer=InMemorySaver()) a = {"configurable": {"thread_id": "chat-123"}} b = {"configurable": {"thread_id": "chat-456"}} for config, text in [(a, "My order number is ORD1001."), (a, "What is my order number?"), (b, "What is my order number?")]: result = app.invoke({"messages": [{"role": "user", "content": text}]}, config) print(config["configurable"]["thread_id"], "|", result["messages"][-1].content, f"({len(result['messages'])} messages)")
Output - real llama3 runs
chat-123 | I've located your order, ORD1001, and I'm happy to assist you ... (2 messages) chat-123 | Your order number is ORD1001. (4 messages) chat-456 | I'm happy to help! Can you please provide me with your email address or order details so I can look up your order number for you? (2 messages)
Postgres for production - same interface (not run here)
# pip install langgraph-checkpoint-postgres from langgraph.checkpoint.postgres import PostgresSaver DB_URI = "postgresql://agent:secret@localhost:5432/agents" with PostgresSaver.from_conn_string(DB_URI) as checkpointer: checkpointer.setup() # create the tables - once graph = builder.compile(checkpointer=checkpointer) graph.invoke(None, {"configurable": {"thread_id": "research-42"}})

Tip: When a run looks wrong, read the thread before reading the code: get_state(config).values for the data, .next for where it stopped, .tasks for the error, get_state_history(config) for how it got there.

Checkpointing at a glance

Attach a checkpointer

Saves a checkpoint after every step.

builder.compile(checkpointer=InMemorySaver())
Pick the thread

Required on every call once a checkpointer is set.

{"configurable": {"thread_id": "t-1"}}
Latest checkpoint

values, next, tasks (errors).

graph.get_state(config)
All checkpoints

Newest first, with step numbers.

graph.get_state_history(config)
Resume

Continue from the saved position.

graph.invoke(None, config)
Durable, one machine

A SQLite file.

SqliteSaver(sqlite3.connect(path, check_same_thread=False))
Durable, shared

Postgres; run setup() once.

PostgresSaver.from_conn_string(uri)
Delete a thread

Remove all its checkpoints.

checkpointer.delete_thread("t-1")

Try it yourself

The code does not change. Swap the content string and the program does something else entirely.

Count on

“Run the counter twice on one thread with invoke({}, config). Then with {"count": 0}. Explain the difference.”

Read the history

“Print step, source and next for every checkpoint of a two-node graph’s thread.”

Crash and resume

“Make the middle node raise, read get_state().next and .tasks, then fix it and call invoke(None, config).”

Restart

“Run the research script in two terminals - crash, then resume. Swap SqliteSaver for InMemorySaver and repeat.”

Two users

“Run the llama3 chat on two threads and confirm one never sees the other’s order number.”

What usually goes wrong

Expecting state to survive without a checkpointer

A plain compile() keeps nothing between calls. Persistence starts with a checkpointer and a thread_id.

✗ graph = builder.compile()
✓ graph = builder.compile(checkpointer=saver)
Resending the input to resume

Input starts a new run on the thread - plan ran again in ours. None continues from the checkpoint.

✗ graph.invoke({"topic": "LangGraph"}, config)
✓ graph.invoke(None, config)
InMemorySaver in production

It is gone on restart and private to each worker. Use a database-backed checkpointer.

✗ checkpointer=InMemorySaver()
✓ PostgresSaver.from_conn_string(uri)
One thread ID for everyone

Every user then shares one conversation and one state.

✗ {"configurable": {"thread_id": "agent"}}
✓ {"configurable": {"thread_id": conversation.id}}
Treating the thread ID as permission

LangGraph loads whatever thread you name. Check that the signed-in user owns it before building the config.

✗ config = {"configurable": {"thread_id": request.thread_id}}
✓ thread = threads.get_owned(request.thread_id, user.id)   # 404 if not theirs
config = {"configurable": {"thread_id": thread.id}}
Saving secrets without a plan

Checkpoints contain every message and tool result. Apply access control, encryption and a retention policy, and delete threads you no longer need.

Key points

  • A checkpointer saves a checkpoint - state plus position - after every step.
  • thread_id groups checkpoints; every call on a checkpointed graph needs one.
  • get_state shows values, next node and errors; get_state_history shows every step.
  • invoke(None, config) resumes; sending input starts a new run on the thread.
  • The failed node re-runs from its start, so keep side effects idempotent.
  • InMemorySaver is lost on restart; SQLite or Postgres survives it.
  • Thread IDs identify, they do not authorize - check ownership first.
  • Checkpointing is short-term, per-thread memory, and the base for human approval.

Quick check before you move on

What is persistence?
Saving information so it can be retrieved later - here, saving a graph’s execution state outside the running call.
What is a checkpoint?
A saved snapshot of a run after a step: the state values plus where execution is - the next node and any error.
Why is thread_id important?
It identifies one conversation or workflow, so its checkpoints are stored and loaded separately from every other thread.
Does in-memory checkpointing survive a process restart?
No. A new process found {} for the same thread.
How do you continue a crashed run?
Call graph.invoke(None, config) with the same thread_id. It resumes at the node that failed.

Quiz

  1. 1.

    What problem does checkpointing solve?

  2. 2.

    A thread holds count 1. You call invoke({"count": 0}, config). What comes back, and why?

  3. 3.

    search crashed after plan ran. What does get_state(config).next show, and what runs on invoke(None, config)?

  4. 4.

    Which setup lets three API workers serve the same thread?

  5. 5.

    Is checkpointing the same as long-term memory?

Interview questions

What is checkpointing in LangGraph?

Compiling a graph with a checkpointer so it saves a snapshot of state and execution position after every step, filed under a thread_id. Runs can then be inspected, resumed after failure, or continued by later calls.

What is the difference between state and a checkpoint?

State is the data in the graph now. A checkpoint is a persisted snapshot of it plus the position - next node, step number, errors - enough to continue the run.

How do you resume a failed run, and what are the caveats?

invoke(None, config) on the same thread. It continues at the failed node, which runs again from its start - so nodes with side effects must be idempotent. Passing input instead starts a new run.

Why is InMemorySaver not enough in production?

It lives in process memory: lost on restart or deploy, and not shared between workers, so a request on another worker sees an empty thread.

How would you design checkpointing for a multi-tenant SaaS?

Server-generated thread IDs recorded against tenant and user, an ownership check before building the config, a shared durable store such as Postgres, encryption and access control on that store, and a retention policy that deletes old threads.

Which workflows benefit most?

Long-running and multi-step ones - research, tool-heavy agents, multi-agent runs - plus multi-turn conversations and anything that waits for a human approval.

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...