Persistence & Checkpointing
Give a graph a checkpointer and a thread_id, and its state survives between calls, across a crash, and - with a database - across a restart.
What you will be able to do
- Explain what a checkpoint is and when LangGraph saves one
- Compile a graph with a checkpointer and run it on a thread
- Read a thread back with get_state and get_state_history
- Resume a crashed run with invoke(None, config) - in the same process and in a new one
- Choose between InMemorySaver and a database-backed checkpointer
- Treat thread IDs and checkpoint data as application data: scoped, authorized, retained deliberately
The idea, in plain English
Until now every invoke() started from nothing. The state lived in the call and disappeared with it: a second question did not know about the first, and a crash halfway through a five-step research run meant starting over at step one.
A checkpointer fixes that. Compile the graph with one and LangGraph saves a checkpoint after every step - the state values, plus where execution is: which node runs next, and any error from the step that failed. Like a save point in a game, it is enough to look at the run later or carry on from it.
Checkpoints are grouped by thread. You pass a thread_id in the config on every call; each thread is one independent conversation or workflow with its own history. The same compiled graph serves every thread - the definition is shared, the state is not.
Where checkpoints live is the checkpointer’s choice. InMemorySaver keeps them in a Python dict, gone when the process exits. SqliteSaver and PostgresSaver keep them in a database, so a new process can load the thread and continue. We ran everything below on LangGraph 1.2.14 with langgraph-checkpoint-sqlite 3.1.1; PostgresSaver (langgraph-checkpoint-postgres 3.1.2) has the same interface - we checked its API, but had no Postgres server running, so that block was not run.
Worked example: A research graph that crashes in search, then finishes in a new process without re-running plan.
1 - The run starts on a thread
Process 1 calls invoke({"topic": "LangGraph"}, config) with thread_id research-42. The input itself is saved as the first checkpoint.
A real run: plan -> search -> write on thread research-42, with SqliteSaver. search fails in process 1; process 2 finishes the job.
State, checkpoint, thread
State is the data flowing through the graph right now. A checkpoint is a saved snapshot of a run: the state values plus the execution position - the next node, the step number, the error of a step that failed. A thread is the ID that groups a run’s checkpoints together, so the next call on that thread starts from the latest one.
get_state(config) returns the latest checkpoint of a thread as a StateSnapshot: .values, .next and, after a crash, .tasks with the error. get_state_history(config) returns every checkpoint, newest first.
StateWhat is happening - the values in the graph.Thread IDWhich conversation or workflow - the key the checkpoints are filed under.CheckpointWhere it got to - a saved snapshot of state and position after a step.When checkpoints are saved
One per step, not just one at the end. A two-node graph a -> b left four checkpoints on its thread: step -1 (the input, next __start__), step 0 (next a), step 1 (count 1, next b), step 2 (count 2, next nothing). That is why a crash in the middle loses only the step that failed.
A step that crashes saves no new checkpoint. When search raised, the newest checkpoint was still step 1 with next ('search',), and its tasks recorded ConnectionError('search service down').
Resume, or start again
To continue a thread from where it stopped, invoke it with None as the input: graph.invoke(None, config). In our run search ran again and write followed; plan ran once in total.
Passing input instead starts a new run on the same thread. Re-invoking the crashed research thread with {"topic": "LangGraph"} ran plan a second time. The same rule surprises people with the counter example: invoke({"count": 0}) on a thread that already holds 1 returns 1 again - the input overwrote the saved count. To build on what the thread holds, send only what is new: invoke({}, config) returned the saved count plus the graph’s work, and in a chat you send only the new message.
Watch out: Resume works per node: the node that failed runs again from its first line. If it sent an email or charged a card before crashing, it will do so twice. Keep side effects idempotent, or split them into their own node.
Conversations on a thread
With MessagesState the messages field appends, so every invoke on the same thread adds to the conversation the checkpointer already holds. We ran llama3 through a one-node chat graph. On thread chat-123: "My order number is ORD1001." then "What is my order number?" - the answer was "Your order number is ORD1001." with four messages in state. The same question on chat-456 got "Can you please provide me with your email address or order details...", because that thread had only its own message.
That is short-term memory: the history of one conversation, restored by thread. Long-term memory - facts about a user that outlive any one conversation, such as "prefers Telugu" - is a separate store, covered later in the course.
In memory or in a database
InMemorySaver keeps checkpoints in the Python process. We ran demo-1 in one process (count 1), then a second process with an identical graph: get_state returned {}. The same is true for workers behind a load balancer - each has its own memory, so a thread started on worker 1 is empty on worker 2.
SqliteSaver writes to a file - two tables, checkpoints and writes - which is why our process 2 found the crashed run. It suits a single machine. For a service with several workers, PostgresSaver (package langgraph-checkpoint-postgres) stores the same checkpoints in a shared database; call setup() once to create its tables.
InMemorySaverLearning, tests, notebooks. Lost when the process exits; not shared between workers.SqliteSaverA local file. Survives a restart; one machine.PostgresSaverA shared database. Survives restarts and deploys; every worker sees every thread.Thread IDs are not security
Any caller who passes a thread_id gets that thread’s state - LangGraph does not know who owns it. An unknown thread_id is not an error either: get_state on "nobody" returned {}, so a typo silently starts an empty conversation.
Generate thread IDs on the server, record which tenant and user own each one, and check ownership before you build the config. Checkpoints hold everything the state holds - user messages, tool results, personal data - so treat the store like any application database: access control, encryption, backups, and a retention policy. delete_thread(thread_id) on the checkpointer removes a thread; after it, get_state returned {}.
Why this comes before human approval
A run that can stop and be continued later is exactly what human-in-the-loop needs: draft the email, save a checkpoint, wait - minutes or days - for approval, then resume and send. Lesson 3.7 builds that pause on top of the checkpointer from this lesson.
Step-by-step code
from typing import TypedDict
from langgraph.checkpoint.memory import InMemorySaver
from langgraph.graph import END, START, StateGraph
class State(TypedDict):
count: int
def increment(state: State):
return {"count": state["count"] + 1}
builder = StateGraph(State)
builder.add_node("increment", increment)
builder.add_edge(START, "increment")
builder.add_edge("increment", END)
graph = builder.compile(checkpointer=InMemorySaver()) # the one change
demo1 = {"configurable": {"thread_id": "demo-1"}}
demo2 = {"configurable": {"thread_id": "demo-2"}}
print(graph.invoke({"count": 0}, demo1))
print(graph.invoke({"count": 0}, demo1)) # input overwrites the saved count
print(graph.invoke({"count": 5}, demo1))
print(graph.invoke({"count": 100}, demo2))
snapshot = graph.get_state(demo1)
print(snapshot.values, snapshot.next)
print(len(list(graph.get_state_history(demo1))), "checkpoints on demo-1")
graph.invoke({"count": 0}) # no thread_id{'count': 1}
{'count': 1}
{'count': 6}
{'count': 101}
{'count': 6} ()
9 checkpoints on demo-1 (3 runs x input, start, increment)
ValueError: Checkpointer requires one or more of the following 'configurable' keys:
thread_id, checkpoint_ns, checkpoint_id
graph.get_state({"configurable": {"thread_id": "nobody"}}).values -> {}import sqlite3
import sys
from typing import TypedDict
from langgraph.checkpoint.sqlite import SqliteSaver
from langgraph.graph import END, START, StateGraph
SEARCH_IS_DOWN = sys.argv[1] == "crash"
class Research(TypedDict, total=False):
topic: str
plan: str
notes: str
report: str
def plan(state):
print(" running plan")
return {"plan": f"plan for {state['topic']}"}
def search(state):
print(" running search")
if SEARCH_IS_DOWN:
raise ConnectionError("search service down")
return {"notes": "3 sources found"}
def write(state):
print(" running write")
return {"report": f"{state['plan']} + {state['notes']}"}
builder = StateGraph(Research)
builder.add_node("plan", plan)
builder.add_node("search", search)
builder.add_node("write", write)
builder.add_edge(START, "plan")
builder.add_edge("plan", "search")
builder.add_edge("search", "write")
builder.add_edge("write", END)
conn = sqlite3.connect("checkpoints.db", check_same_thread=False)
graph = builder.compile(checkpointer=SqliteSaver(conn))
config = {"configurable": {"thread_id": "research-42"}}
if SEARCH_IS_DOWN:
try:
graph.invoke({"topic": "LangGraph"}, config)
except ConnectionError as error:
print(" crashed:", error)
print(" next:", graph.get_state(config).next)
else:
state = graph.get_state(config)
print(" found:", state.values, "next:", state.next)
print(" result:", graph.invoke(None, config)["report"]) # None = resume$ python research.py crash
running plan
running search
crashed: search service down
next: ('search',)
$ python research.py resume
found: {'topic': 'LangGraph', 'plan': 'plan for LangGraph'} next: ('search',)
running search
running write
result: plan for LangGraph + 3 sources found
checkpoints.db tables: checkpoints, writes
The same experiment with InMemorySaver:
process 1: {'count': 1}
process 2 get_state: {}from langchain_ollama import ChatOllama
from langgraph.checkpoint.memory import InMemorySaver
from langgraph.graph import END, START, MessagesState, StateGraph
llm = ChatOllama(model="llama3", temperature=0)
def chat(state: MessagesState):
system = ("system", "You are a support agent. Answer in one short sentence.")
return {"messages": [llm.invoke([system] + state["messages"])]}
builder = StateGraph(MessagesState)
builder.add_node("chat", chat)
builder.add_edge(START, "chat")
builder.add_edge("chat", END)
app = builder.compile(checkpointer=InMemorySaver())
a = {"configurable": {"thread_id": "chat-123"}}
b = {"configurable": {"thread_id": "chat-456"}}
for config, text in [(a, "My order number is ORD1001."),
(a, "What is my order number?"),
(b, "What is my order number?")]:
result = app.invoke({"messages": [{"role": "user", "content": text}]}, config)
print(config["configurable"]["thread_id"], "|", result["messages"][-1].content,
f"({len(result['messages'])} messages)")chat-123 | I've located your order, ORD1001, and I'm happy to assist you ... (2 messages)
chat-123 | Your order number is ORD1001. (4 messages)
chat-456 | I'm happy to help! Can you please provide me with your email address or
order details so I can look up your order number for you? (2 messages)# pip install langgraph-checkpoint-postgres
from langgraph.checkpoint.postgres import PostgresSaver
DB_URI = "postgresql://agent:secret@localhost:5432/agents"
with PostgresSaver.from_conn_string(DB_URI) as checkpointer:
checkpointer.setup() # create the tables - once
graph = builder.compile(checkpointer=checkpointer)
graph.invoke(None, {"configurable": {"thread_id": "research-42"}})Tip: When a run looks wrong, read the thread before reading the code: get_state(config).values for the data, .next for where it stopped, .tasks for the error, get_state_history(config) for how it got there.
Checkpointing at a glance
Attach a checkpointerSaves a checkpoint after every step.
builder.compile(checkpointer=InMemorySaver())
Pick the threadRequired on every call once a checkpointer is set.
{"configurable": {"thread_id": "t-1"}}Latest checkpointvalues, next, tasks (errors).
graph.get_state(config)
All checkpointsNewest first, with step numbers.
graph.get_state_history(config)
ResumeContinue from the saved position.
graph.invoke(None, config)
Durable, one machineA SQLite file.
SqliteSaver(sqlite3.connect(path, check_same_thread=False))
Durable, sharedPostgres; run setup() once.
PostgresSaver.from_conn_string(uri)
Delete a threadRemove all its checkpoints.
checkpointer.delete_thread("t-1")Try it yourself
The code does not change. Swap the content string and the program does something else entirely.
“Run the counter twice on one thread with invoke({}, config). Then with {"count": 0}. Explain the difference.”
“Print step, source and next for every checkpoint of a two-node graph’s thread.”
“Make the middle node raise, read get_state().next and .tasks, then fix it and call invoke(None, config).”
“Run the research script in two terminals - crash, then resume. Swap SqliteSaver for InMemorySaver and repeat.”
“Run the llama3 chat on two threads and confirm one never sees the other’s order number.”
What usually goes wrong
A plain compile() keeps nothing between calls. Persistence starts with a checkpointer and a thread_id.
✗ graph = builder.compile()✓ graph = builder.compile(checkpointer=saver)Input starts a new run on the thread - plan ran again in ours. None continues from the checkpoint.
✗ graph.invoke({"topic": "LangGraph"}, config)✓ graph.invoke(None, config)It is gone on restart and private to each worker. Use a database-backed checkpointer.
✗ checkpointer=InMemorySaver()✓ PostgresSaver.from_conn_string(uri)Every user then shares one conversation and one state.
✗ {"configurable": {"thread_id": "agent"}}✓ {"configurable": {"thread_id": conversation.id}}LangGraph loads whatever thread you name. Check that the signed-in user owns it before building the config.
✗ config = {"configurable": {"thread_id": request.thread_id}}✓ thread = threads.get_owned(request.thread_id, user.id) # 404 if not theirs
config = {"configurable": {"thread_id": thread.id}}Checkpoints contain every message and tool result. Apply access control, encryption and a retention policy, and delete threads you no longer need.
Key points
- A checkpointer saves a checkpoint - state plus position - after every step.
- thread_id groups checkpoints; every call on a checkpointed graph needs one.
- get_state shows values, next node and errors; get_state_history shows every step.
- invoke(None, config) resumes; sending input starts a new run on the thread.
- The failed node re-runs from its start, so keep side effects idempotent.
- InMemorySaver is lost on restart; SQLite or Postgres survives it.
- Thread IDs identify, they do not authorize - check ownership first.
- Checkpointing is short-term, per-thread memory, and the base for human approval.
Quick check before you move on
Quiz
- 1.
What problem does checkpointing solve?
- 2.
A thread holds count 1. You call invoke({"count": 0}, config). What comes back, and why?
- 3.
search crashed after plan ran. What does get_state(config).next show, and what runs on invoke(None, config)?
- 4.
Which setup lets three API workers serve the same thread?
- 5.
Is checkpointing the same as long-term memory?
Interview questions
What is checkpointing in LangGraph?
Compiling a graph with a checkpointer so it saves a snapshot of state and execution position after every step, filed under a thread_id. Runs can then be inspected, resumed after failure, or continued by later calls.
What is the difference between state and a checkpoint?
State is the data in the graph now. A checkpoint is a persisted snapshot of it plus the position - next node, step number, errors - enough to continue the run.
How do you resume a failed run, and what are the caveats?
invoke(None, config) on the same thread. It continues at the failed node, which runs again from its start - so nodes with side effects must be idempotent. Passing input instead starts a new run.
Why is InMemorySaver not enough in production?
It lives in process memory: lost on restart or deploy, and not shared between workers, so a request on another worker sees an empty thread.
How would you design checkpointing for a multi-tenant SaaS?
Server-generated thread IDs recorded against tenant and user, an ownership check before building the config, a shared durable store such as Postgres, encryption and access control on that store, and a retention policy that deletes old threads.
Which workflows benefit most?
Long-running and multi-step ones - research, tool-heavy agents, multi-agent runs - plus multi-turn conversations and anything that waits for a human approval.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...
AI
System Design
Backend
- GraphQL8 modules · 69 lessons planned
- Core Python13 modules · 75 lessons planned
- FastAPI5 sections · 20 lessons
- Node.js14 modules · 206 lessons planned
- Node.js Performance7 chapters · 36 topics
- Event Loop Lifecycle6 phases · 3 scenarios
- Docker & Containerization11 modules · 144 lessons planned
- AWS for Developers14 modules · 219 lessons planned
- CI/CD & DevOps Automation10 modules · 134 lessons planned