Memory
The chat agent with memory, next to Module 1’s version: ConversationBufferMemory and its relatives, and the checkpointer that replaced them.
What you will be able to do
- Explain why memory is application-managed context, not the model remembering
- Recognise ConversationBufferMemory and its relatives, and why they are deprecated
- Give an agent memory with a checkpointer and a thread_id
- Keep each user’s conversation separate, and inspect what is stored
- Trim or summarize history - and predict what gets forgotten
- Tell conversation memory from persistent memory, and make memory survive a restart
The idea, in plain English
A model call is stateless. Say "My name is Chandu" in one call and ask "What’s my name?" in the next, and the model does not know. Memory is your application keeping the conversation and sending it back each time - Lesson 1.8’s self.messages list.
Classic LangChain wrapped that list in memory classes: ConversationBufferMemory kept everything, ConversationBufferWindowMemory kept the last k exchanges, ConversationSummaryMemory kept a summary. They still run from the langchain-classic package, but every one prints a deprecation warning pointing to create_agent with checkpointing. In LangChain 1.x there is no langchain.memory module at all.
The replacement is a checkpointer. create_agent(..., checkpointer=InMemorySaver()) saves the agent’s messages after every run, under a thread_id you pass in the config. Call again with the same thread_id and the history is loaded and sent with your new message. Different thread_id, different conversation.
Every model output here is from llama3. A chat agent with no tools does not need tool calling, so these runs are the real thing.
Worked example: Chat agent with memory, next to Module 1’s manual version.
1 - Turn 1 is saved
"Hi, my name is Chandu." The model replies, and the checkpointer stores both messages under thread user-123.
create_agent with a checkpointer: the history is loaded before the model call and saved after it.
Memory is a message list
Without memory, every call is a blank slate. We ran create_agent with no checkpointer: turn 1 got "Nice to meet you, Chandu!", turn 2 got "I don’t actually know your name." Same agent object, no history between calls.
Lesson 1.8’s ChatAgent fixed it by appending each user message and reply to self.messages and sending the whole list: "I remember! Your name is Chandu!" The model did not remember anything - the application supplied the history. Every LangChain memory feature is that list, managed for you.
ConversationBufferMemory and its relatives
Classic LangChain had a family of memory classes. ConversationBufferMemory kept every message - a buffer, nothing dropped. ConversationBufferWindowMemory(k=1) kept the last k exchanges: after three, only q2/a2 remained. ConversationSummaryMemory replaced old messages with a model-written summary.
You will meet them in tutorials. In LangChain 1.4, from langchain.memory import ... fails with ModuleNotFoundError. From langchain-classic they still run, with a warning: "ConversationBufferMemory was deprecated in LangChain 0.3.1 and will be removed in 2.0.0. Use langchain.agents.create_agent instead. For agents that need to remember prior interactions, use create_agent with checkpointing or the Store API." Even InMemoryChatMessageHistory in langchain-core is deprecated as of 1.6.4.
ConversationBufferMemoryKeep everything -> a checkpointer with no trimming.ConversationBufferWindowMemory(k)Keep the last k exchanges -> trim_messages in a before_model hook.ConversationSummaryMemorySummarize old messages -> SummarizationMiddleware.One memory object per userOne thread_id per conversation.The checkpointer and thread_id
Pass checkpointer=InMemorySaver() to create_agent and pass {"configurable": {"thread_id": "user-123"}} with every invoke. You send only the new message; the agent loads the thread’s history, runs, and saves the result. Call it without a thread_id and it raises ValueError: Checkpointer requires one or more of the following ‘configurable’ keys: thread_id, ...
agent.get_state(config).values["messages"] shows exactly what is stored - the first thing to check when an agent "forgets". Notice the system prompt is not in it: create_agent adds system_prompt to every call, so trimming the history can never remove it. In Module 1 you had to protect messages[0] yourself.
One thread per conversation
A hundred users cannot share one list. thread_id is the session key: we told thread user-123 the name, then asked thread user-456 - "I don’t actually know your name." Same agent, separate histories.
Memory also lives in the checkpointer, not the agent. A new agent with a new InMemorySaver did not know the name, even with the same thread_id - the same as Lesson 1.8’s new ChatAgent() starting empty.
Memory costs tokens
Every turn resends the whole thread. Over five short turns the input tokens llama3 processed went 35, 57, 88, 120, 141 - and they keep growing with every exchange. Lesson 1.8 showed what happens past the context window: Ollama silently dropped the oldest messages.
So a buffer that keeps everything works for short chats and needs a strategy for long ones: trim, summarize, or move facts somewhere else.
Trimming - by whole exchanges
Module 1’s first trim kept messages[0] plus the last N-1 messages, and could start the kept history with an assistant reply whose question was gone. trim_messages does the same with start_on="human": with 5 exchanges and room for 6 messages, without start_on it kept "assistant 3, user 4 ... assistant 5"; with it, "user 4 ... assistant 5" - whole exchanges only.
In an agent, trim in a before_model hook. We kept the last 5 messages: after three facts, "What’s my name?" got "I don’t know your name" while "Where do I live?" got "You live in Hyderabad!" - the name had been trimmed, the city had not. Trimming saves tokens by forgetting, and it forgets the oldest things first.
Summarizing - and what a summary drops
SummarizationMiddleware asks a model to summarize older messages when a trigger is reached, and replaces them with the summary. With trigger=("messages", 6) and keep=("messages", 2), llama3’s summary read: "The user, Chandu, has introduced themselves and shared their interest in Python."
It kept the name and lost the city. "I live in Hyderabad" was summarized away, and "Where do I live?" got "I don’t have that information." A summary is the model deciding what matters - and it can decide wrong.
Watch out: Every way of shrinking memory forgets something. Trimming forgets the oldest messages; summarizing forgets whatever the summary leaves out. Test what your agent still knows after it happens.
Conversation memory vs persistent memory
InMemorySaver - like self.messages - lives in the running process. Stop the program and the conversation is gone. Persistent memory is stored outside: SqliteSaver (pip install langgraph-checkpoint-sqlite) writes threads to a database file. We ran "Hi, my name is Chandu." in one Python process and "What’s my name?" in a second one: "I remember! Your name is Chandu!" - 4 messages loaded from disk.
Not everything belongs in permanent memory. A thread is short-term: this conversation. Long-term facts - the user prefers Python examples - are worth storing separately and loading when relevant; LangGraph’s Store API, named in the deprecation warning, is built for that. Production systems combine recent messages, summaries and stored facts.
Memory and tools together
Lesson 2.7’s agent plus a checkpointer. Turn 1: "My order number is 12345." Turn 2: "Where is my order?" - on that call the model received both turns, so it could call get_order_status with order_id 12345, which it was never given in the second message. (Scripted model, as in Lesson 2.7; llama3 cannot call tools.)
When the agent forgets - a checklist
Before blaming the model, look at what it was given.
Is it stored?agent.get_state(config).values["messages"]Right thread?The same thread_id on every call of the conversation.Trimmed or summarized?Check what the hook or the summary kept.Too long?Watch input_tokens; past the window, Ollama drops the oldest text.Reset?A new InMemorySaver - or a restart - starts empty.Step-by-step code
import ollama
class ChatAgent:
def __init__(self, model="llama3.1", max_history=10):
self.model = model
self.max_history = max_history
self.messages = [{"role": "system", "content": "You are a friendly assistant."}]
def ask(self, user_input):
self.messages.append({"role": "user", "content": user_input})
if len(self.messages) > self.max_history:
self.messages = [self.messages[0]] + self.messages[-(self.max_history - 1):]
response = ollama.chat(model=self.model, messages=self.messages)
reply = response["message"]["content"]
self.messages.append({"role": "assistant", "content": reply})
return reply
agent = ChatAgent()
print(agent.ask("Hi, my name is Chandu."))
print(agent.ask("What's my name?"))
# One run (the wording changes from run to run):
# Nice to meet you, Chandu! I'm happy to help you with any questions or topics you'd like to discuss. ...
# Your name is Chandu!from langchain.agents import create_agent
from langchain_ollama import ChatOllama
from langgraph.checkpoint.memory import InMemorySaver
agent = create_agent(
ChatOllama(model="llama3.1", temperature=0),
tools=[],
system_prompt="You are a friendly assistant.",
checkpointer=InMemorySaver(),
)
config = {"configurable": {"thread_id": "user-123"}}
for question in ["Hi, my name is Chandu.", "What's my name?"]:
result = agent.invoke({"messages": [{"role": "user", "content": question}]}, config)
print(result["messages"][-1].content)
# Nice to meet you, Chandu! I'm happy to chat with you. How's your day going so far?
# I remember! Your name is Chandu!no checkpointer, two calls "I don't actually know your name."
thread user-123, turn 2 "I remember! Your name is Chandu!" input tokens 29 -> 68
thread user-456 "I don't actually know your name."
new InMemorySaver, thread user-123 "I don't actually know your name."
invoke() with no thread_id ValueError: Checkpointer requires ... thread_id
agent.get_state(config).values["messages"]
human Hi, my name is Chandu.
ai Nice to meet you, Chandu! I'm happy to c...
human What's my name?
ai I remember! Your name is Chandu!
<- no system message: system_prompt is added on every call# pip install langchain-classic (from langchain.memory import ... -> ModuleNotFoundError in 1.x)
from langchain_classic.memory import ConversationBufferMemory, ConversationBufferWindowMemory
memory = ConversationBufferMemory(return_messages=True)
memory.save_context({"input": "My name is Chandu."}, {"output": "Nice to meet you, Chandu!"})
print(memory.load_memory_variables({}))
# {'history': [HumanMessage(content='My name is Chandu.'), AIMessage(content='Nice to meet you, Chandu!')]}
window = ConversationBufferWindowMemory(k=1, return_messages=True)
for i in range(3):
window.save_context({"input": f"q{i}"}, {"output": f"a{i}"})
print(window.load_memory_variables({}))
# {'history': [HumanMessage(content='q2'), AIMessage(content='a2')]}
# LangChainDeprecationWarning: The class ConversationBufferMemory was deprecated in LangChain 0.3.1
# and will be removed in 2.0.0. Use langchain.agents.create_agent instead. For agents that need to
# remember prior interactions, use create_agent with checkpointing or the Store API.from langchain_core.messages import AIMessage, HumanMessage, SystemMessage, trim_messages
history = [SystemMessage("You are a friendly assistant.")]
for i in range(1, 6):
history += [HumanMessage(f"user {i}"), AIMessage(f"assistant {i}")]
kept = trim_messages(
history,
strategy="last", # keep the most recent
token_counter=len, # count messages, not tokens
max_tokens=6,
include_system=True, # never drop the system message
start_on="human", # start on a question - whole exchanges only
)
print([m.content for m in kept])
# with start_on="human": ['You are a friendly assistant.', 'user 4', 'assistant 4', 'user 5', 'assistant 5']
# without it: ['You are a friendly assistant.', 'assistant 3', 'user 4', 'assistant 4', 'user 5', 'assistant 5']from langchain.agents.middleware import before_model
from langchain_core.messages import RemoveMessage
from langgraph.graph.message import REMOVE_ALL_MESSAGES
llm = ChatOllama(model="llama3.1", temperature=0)
@before_model
def keep_recent(state, runtime):
recent = trim_messages(state["messages"], strategy="last", token_counter=len,
max_tokens=5, start_on="human")
return {"messages": [RemoveMessage(id=REMOVE_ALL_MESSAGES), *recent]}
agent = create_agent(llm, tools=[], system_prompt="You are a friendly assistant. Keep answers to one sentence.",
middleware=[keep_recent], checkpointer=InMemorySaver())Five turns: name, Python, city, then two questions.
input tokens per turn "What's my name?" "Where do I live?"
keep everything 35 57 88 120 141 Your name is Chandu! You live in Hyderabad!
trim to 5 messages 35 57 88 94 95 I don't know your name... You live in Hyderabad!
summarize (keep 2) - Your name is Chandu! I don't have that information...
The summary llama3 wrote:
"The user, Chandu, has introduced themselves and shared their interest in Python."
- the name survived, Hyderabad did not.from langchain.agents.middleware import SummarizationMiddleware
agent = create_agent(
llm,
tools=[],
middleware=[SummarizationMiddleware(model=llm, trigger=("messages", 6), keep=("messages", 2))],
checkpointer=InMemorySaver(),
)# pip install langgraph-checkpoint-sqlite
import sqlite3
from langgraph.checkpoint.sqlite import SqliteSaver
conn = sqlite3.connect("chat.db", check_same_thread=False)
agent = create_agent(llm, tools=[], system_prompt="You are a friendly assistant.",
checkpointer=SqliteSaver(conn))
# Run 1 (one process): "Hi, my name is Chandu." -> Nice to meet you, Chandu! ... 2 messages
# Run 2 (a new process): "What's my name?" -> I remember! Your name is Chandu! 4 messagesturn 1 human My order number is 12345.
ai Sure - I'll help you with order 12345.
turn 2 model received:
human My order number is 12345. <- loaded from the thread
ai Sure - I'll help you with order 12345.
human Where is my order?
ai get_order_status {'order_id': '12345'}
tool Shipped - arriving Friday
ai Order 12345 has shipped and arrives Friday.Tip: Print agent.get_state(config).values["messages"] whenever memory surprises you. It is the LangChain version of printing self.messages.
Memory at a glance
InMemorySaverConversation memory for the running process.
create_agent(..., checkpointer=InMemorySaver())
thread_idOne conversation; required with a checkpointer.
{"configurable": {"thread_id": "user-123"}}get_stateWhat the thread has stored.
agent.get_state(config).values["messages"]
trim_messagesKeep the recent part, by whole exchanges.
trim_messages(..., start_on="human")
before_modelHook to trim before each model call.
@before_model
SummarizationMiddlewareReplace old messages with a summary.
SummarizationMiddleware(model=llm, trigger=...)
SqliteSaverPersistent memory in a database file.
from langgraph.checkpoint.sqlite import SqliteSaver
ConversationBufferMemoryClassic and deprecated; in langchain-classic.
from langchain_classic.memory import ...
Try it yourself
The code does not change. Swap the content string and the program does something else entirely.
“Tell thread "a" your name and thread "b" a different one; ask both.”
“After three turns, print get_state(config).values["messages"] and count them.”
“Add the keep_recent hook with max_tokens=3 and find the first turn where the agent forgets your name.”
“Switch to SqliteSaver, tell it your name, stop the program, start it again and ask.”
What usually goes wrong
That module does not exist in LangChain 1.x. Use a checkpointer - or langchain-classic for old code.
✗ from langchain.memory import ConversationBufferMemory✓ create_agent(..., checkpointer=InMemorySaver())A fixed thread_id shares one conversation between all users. Use the user or session id.
✗ {"configurable": {"thread_id": "chat"}}✓ {"configurable": {"thread_id": session_id}}It lives in the process. A restart, or a new saver, starts every thread empty.
✗ checkpointer=InMemorySaver() # for production users✓ checkpointer=SqliteSaver(conn) # or another database-backed saverCounting messages can keep an answer without its question. Start the kept history on a human message.
✗ [messages[0]] + messages[-(n - 1):]✓ trim_messages(..., include_system=True, start_on="human")Our summary kept the name and dropped the city. Keep facts you need somewhere explicit.
Key points
- A model call is stateless; memory is the history your application sends back.
- ConversationBufferMemory and its relatives are deprecated; they live in langchain-classic.
- In LangChain 1.x: create_agent with a checkpointer, and a thread_id per conversation.
- The system prompt is added on every call, so trimming cannot remove it.
- History costs tokens on every turn: 35 to 141 in five short turns.
- Trimming forgets the oldest messages; summarizing forgets what the summary leaves out.
- InMemorySaver disappears with the process; SqliteSaver survives a restart.
Quick check before you move on
Quiz
- 1.
You call agent.invoke() with a checkpointer but no config. What happens?
- 2.
Same agent, thread user-123 knows the name. What does thread user-456 say?
- 3.
After trimming to 5 messages, the agent forgot the name but remembered the city. Why?
- 4.
Which classic class matches ConversationBufferWindowMemory(k=2) today?
Interview questions
Does an LLM have memory?
Not between calls. The application creates the appearance of memory by sending previous messages with each request.
What was ConversationBufferMemory, and what replaced it?
A classic LangChain class that stored the full conversation and injected it into the prompt. In LangChain 1.x, agents use a checkpointer keyed by thread_id; the old classes are deprecated in langchain-classic.
How would you manage long conversations?
Trim to recent whole exchanges, summarize older context, store important facts explicitly, and test what the agent still knows afterwards - each technique forgets something.
How do you implement persistent memory?
Use a database-backed checkpointer (SQLite, Postgres) for conversations, and a separate store for long-term facts, loaded when relevant.
How do you keep users’ conversations separate?
Give each conversation its own thread_id - never a shared global history.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...
AI
System Design
Backend
- GraphQL8 modules · 69 lessons planned
- Core Python13 modules · 75 lessons planned
- FastAPI5 sections · 20 lessons
- Node.js14 modules · 206 lessons planned
- Node.js Performance7 chapters · 36 topics
- Event Loop Lifecycle6 phases · 3 scenarios
- Docker & Containerization11 modules · 144 lessons planned
- AWS for Developers14 modules · 219 lessons planned
- CI/CD & DevOps Automation10 modules · 134 lessons planned