← Back to Agentic AI map
Lesson 2.8 · Building Agents with LangChain

Memory

The chat agent with memory, next to Module 1’s version: ConversationBufferMemory and its relatives, and the checkpointer that replaced them.

memory

What you will be able to do

  • Explain why memory is application-managed context, not the model remembering
  • Recognise ConversationBufferMemory and its relatives, and why they are deprecated
  • Give an agent memory with a checkpointer and a thread_id
  • Keep each user’s conversation separate, and inspect what is stored
  • Trim or summarize history - and predict what gets forgotten
  • Tell conversation memory from persistent memory, and make memory survive a restart

The idea, in plain English

A model call is stateless. Say "My name is Chandu" in one call and ask "What’s my name?" in the next, and the model does not know. Memory is your application keeping the conversation and sending it back each time - Lesson 1.8’s self.messages list.

Classic LangChain wrapped that list in memory classes: ConversationBufferMemory kept everything, ConversationBufferWindowMemory kept the last k exchanges, ConversationSummaryMemory kept a summary. They still run from the langchain-classic package, but every one prints a deprecation warning pointing to create_agent with checkpointing. In LangChain 1.x there is no langchain.memory module at all.

The replacement is a checkpointer. create_agent(..., checkpointer=InMemorySaver()) saves the agent’s messages after every run, under a thread_id you pass in the config. Call again with the same thread_id and the history is loaded and sent with your new message. Different thread_id, different conversation.

Every model output here is from llama3. A chat agent with no tools does not need tool calling, so these runs are the real thing.

Worked example: Chat agent with memory, next to Module 1’s manual version.

request flowTwo turns on one threadstep 1 / 4

1 - Turn 1 is saved

"Hi, my name is Chandu." The model replies, and the checkpointer stores both messages under thread user-123.

thread
user-123
stored
2 messages
input tokens
29
system prompt stored
no - added each call

create_agent with a checkpointer: the history is loaded before the model call and saved after it.

Memory is a message list

Without memory, every call is a blank slate. We ran create_agent with no checkpointer: turn 1 got "Nice to meet you, Chandu!", turn 2 got "I don’t actually know your name." Same agent object, no history between calls.

Lesson 1.8’s ChatAgent fixed it by appending each user message and reply to self.messages and sending the whole list: "I remember! Your name is Chandu!" The model did not remember anything - the application supplied the history. Every LangChain memory feature is that list, managed for you.

ConversationBufferMemory and its relatives

Classic LangChain had a family of memory classes. ConversationBufferMemory kept every message - a buffer, nothing dropped. ConversationBufferWindowMemory(k=1) kept the last k exchanges: after three, only q2/a2 remained. ConversationSummaryMemory replaced old messages with a model-written summary.

You will meet them in tutorials. In LangChain 1.4, from langchain.memory import ... fails with ModuleNotFoundError. From langchain-classic they still run, with a warning: "ConversationBufferMemory was deprecated in LangChain 0.3.1 and will be removed in 2.0.0. Use langchain.agents.create_agent instead. For agents that need to remember prior interactions, use create_agent with checkpointing or the Store API." Even InMemoryChatMessageHistory in langchain-core is deprecated as of 1.6.4.

Classic memory classes and today’s equivalents
ConversationBufferMemoryKeep everything -> a checkpointer with no trimming.
ConversationBufferWindowMemory(k)Keep the last k exchanges -> trim_messages in a before_model hook.
ConversationSummaryMemorySummarize old messages -> SummarizationMiddleware.
One memory object per userOne thread_id per conversation.

The checkpointer and thread_id

Pass checkpointer=InMemorySaver() to create_agent and pass {"configurable": {"thread_id": "user-123"}} with every invoke. You send only the new message; the agent loads the thread’s history, runs, and saves the result. Call it without a thread_id and it raises ValueError: Checkpointer requires one or more of the following ‘configurable’ keys: thread_id, ...

agent.get_state(config).values["messages"] shows exactly what is stored - the first thing to check when an agent "forgets". Notice the system prompt is not in it: create_agent adds system_prompt to every call, so trimming the history can never remove it. In Module 1 you had to protect messages[0] yourself.

One thread per conversation

A hundred users cannot share one list. thread_id is the session key: we told thread user-123 the name, then asked thread user-456 - "I don’t actually know your name." Same agent, separate histories.

Memory also lives in the checkpointer, not the agent. A new agent with a new InMemorySaver did not know the name, even with the same thread_id - the same as Lesson 1.8’s new ChatAgent() starting empty.

Memory costs tokens

Every turn resends the whole thread. Over five short turns the input tokens llama3 processed went 35, 57, 88, 120, 141 - and they keep growing with every exchange. Lesson 1.8 showed what happens past the context window: Ollama silently dropped the oldest messages.

So a buffer that keeps everything works for short chats and needs a strategy for long ones: trim, summarize, or move facts somewhere else.

Trimming - by whole exchanges

Module 1’s first trim kept messages[0] plus the last N-1 messages, and could start the kept history with an assistant reply whose question was gone. trim_messages does the same with start_on="human": with 5 exchanges and room for 6 messages, without start_on it kept "assistant 3, user 4 ... assistant 5"; with it, "user 4 ... assistant 5" - whole exchanges only.

In an agent, trim in a before_model hook. We kept the last 5 messages: after three facts, "What’s my name?" got "I don’t know your name" while "Where do I live?" got "You live in Hyderabad!" - the name had been trimmed, the city had not. Trimming saves tokens by forgetting, and it forgets the oldest things first.

Summarizing - and what a summary drops

SummarizationMiddleware asks a model to summarize older messages when a trigger is reached, and replaces them with the summary. With trigger=("messages", 6) and keep=("messages", 2), llama3’s summary read: "The user, Chandu, has introduced themselves and shared their interest in Python."

It kept the name and lost the city. "I live in Hyderabad" was summarized away, and "Where do I live?" got "I don’t have that information." A summary is the model deciding what matters - and it can decide wrong.

Watch out: Every way of shrinking memory forgets something. Trimming forgets the oldest messages; summarizing forgets whatever the summary leaves out. Test what your agent still knows after it happens.

Conversation memory vs persistent memory

InMemorySaver - like self.messages - lives in the running process. Stop the program and the conversation is gone. Persistent memory is stored outside: SqliteSaver (pip install langgraph-checkpoint-sqlite) writes threads to a database file. We ran "Hi, my name is Chandu." in one Python process and "What’s my name?" in a second one: "I remember! Your name is Chandu!" - 4 messages loaded from disk.

Not everything belongs in permanent memory. A thread is short-term: this conversation. Long-term facts - the user prefers Python examples - are worth storing separately and loading when relevant; LangGraph’s Store API, named in the deprecation warning, is built for that. Production systems combine recent messages, summaries and stored facts.

Memory and tools together

Lesson 2.7’s agent plus a checkpointer. Turn 1: "My order number is 12345." Turn 2: "Where is my order?" - on that call the model received both turns, so it could call get_order_status with order_id 12345, which it was never given in the second message. (Scripted model, as in Lesson 2.7; llama3 cannot call tools.)

When the agent forgets - a checklist

Before blaming the model, look at what it was given.

Debugging memory
Is it stored?agent.get_state(config).values["messages"]
Right thread?The same thread_id on every call of the conversation.
Trimmed or summarized?Check what the hook or the summary kept.
Too long?Watch input_tokens; past the window, Ollama drops the oldest text.
Reset?A new InMemorySaver - or a restart - starts empty.

Step-by-step code

Module 1: the manual ChatAgent
import ollama class ChatAgent: def __init__(self, model="llama3.1", max_history=10): self.model = model self.max_history = max_history self.messages = [{"role": "system", "content": "You are a friendly assistant."}] def ask(self, user_input): self.messages.append({"role": "user", "content": user_input}) if len(self.messages) > self.max_history: self.messages = [self.messages[0]] + self.messages[-(self.max_history - 1):] response = ollama.chat(model=self.model, messages=self.messages) reply = response["message"]["content"] self.messages.append({"role": "assistant", "content": reply}) return reply agent = ChatAgent() print(agent.ask("Hi, my name is Chandu.")) print(agent.ask("What's my name?")) # One run (the wording changes from run to run): # Nice to meet you, Chandu! I'm happy to help you with any questions or topics you'd like to discuss. ... # Your name is Chandu!
LangChain 1.x: the same agent with a checkpointer
from langchain.agents import create_agent from langchain_ollama import ChatOllama from langgraph.checkpoint.memory import InMemorySaver agent = create_agent( ChatOllama(model="llama3.1", temperature=0), tools=[], system_prompt="You are a friendly assistant.", checkpointer=InMemorySaver(), ) config = {"configurable": {"thread_id": "user-123"}} for question in ["Hi, my name is Chandu.", "What's my name?"]: result = agent.invoke({"messages": [{"role": "user", "content": question}]}, config) print(result["messages"][-1].content) # Nice to meet you, Chandu! I'm happy to chat with you. How's your day going so far? # I remember! Your name is Chandu!
Sessions, stored state, and resets - from real runs
no checkpointer, two calls "I don't actually know your name." thread user-123, turn 2 "I remember! Your name is Chandu!" input tokens 29 -> 68 thread user-456 "I don't actually know your name." new InMemorySaver, thread user-123 "I don't actually know your name." invoke() with no thread_id ValueError: Checkpointer requires ... thread_id agent.get_state(config).values["messages"] human Hi, my name is Chandu. ai Nice to meet you, Chandu! I'm happy to c... human What's my name? ai I remember! Your name is Chandu! <- no system message: system_prompt is added on every call
ConversationBufferMemory - the classic API
# pip install langchain-classic (from langchain.memory import ... -> ModuleNotFoundError in 1.x) from langchain_classic.memory import ConversationBufferMemory, ConversationBufferWindowMemory memory = ConversationBufferMemory(return_messages=True) memory.save_context({"input": "My name is Chandu."}, {"output": "Nice to meet you, Chandu!"}) print(memory.load_memory_variables({})) # {'history': [HumanMessage(content='My name is Chandu.'), AIMessage(content='Nice to meet you, Chandu!')]} window = ConversationBufferWindowMemory(k=1, return_messages=True) for i in range(3): window.save_context({"input": f"q{i}"}, {"output": f"a{i}"}) print(window.load_memory_variables({})) # {'history': [HumanMessage(content='q2'), AIMessage(content='a2')]} # LangChainDeprecationWarning: The class ConversationBufferMemory was deprecated in LangChain 0.3.1 # and will be removed in 2.0.0. Use langchain.agents.create_agent instead. For agents that need to # remember prior interactions, use create_agent with checkpointing or the Store API.
Trimming by whole exchanges
from langchain_core.messages import AIMessage, HumanMessage, SystemMessage, trim_messages history = [SystemMessage("You are a friendly assistant.")] for i in range(1, 6): history += [HumanMessage(f"user {i}"), AIMessage(f"assistant {i}")] kept = trim_messages( history, strategy="last", # keep the most recent token_counter=len, # count messages, not tokens max_tokens=6, include_system=True, # never drop the system message start_on="human", # start on a question - whole exchanges only ) print([m.content for m in kept]) # with start_on="human": ['You are a friendly assistant.', 'user 4', 'assistant 4', 'user 5', 'assistant 5'] # without it: ['You are a friendly assistant.', 'assistant 3', 'user 4', 'assistant 4', 'user 5', 'assistant 5']
Trimming inside the agent
from langchain.agents.middleware import before_model from langchain_core.messages import RemoveMessage from langgraph.graph.message import REMOVE_ALL_MESSAGES llm = ChatOllama(model="llama3.1", temperature=0) @before_model def keep_recent(state, runtime): recent = trim_messages(state["messages"], strategy="last", token_counter=len, max_tokens=5, start_on="human") return {"messages": [RemoveMessage(id=REMOVE_ALL_MESSAGES), *recent]} agent = create_agent(llm, tools=[], system_prompt="You are a friendly assistant. Keep answers to one sentence.", middleware=[keep_recent], checkpointer=InMemorySaver())
Output - what trimming and summarizing forgot
Five turns: name, Python, city, then two questions. input tokens per turn "What's my name?" "Where do I live?" keep everything 35 57 88 120 141 Your name is Chandu! You live in Hyderabad! trim to 5 messages 35 57 88 94 95 I don't know your name... You live in Hyderabad! summarize (keep 2) - Your name is Chandu! I don't have that information... The summary llama3 wrote: "The user, Chandu, has introduced themselves and shared their interest in Python." - the name survived, Hyderabad did not.
Summarizing old messages
from langchain.agents.middleware import SummarizationMiddleware agent = create_agent( llm, tools=[], middleware=[SummarizationMiddleware(model=llm, trigger=("messages", 6), keep=("messages", 2))], checkpointer=InMemorySaver(), )
Memory that survives a restart
# pip install langgraph-checkpoint-sqlite import sqlite3 from langgraph.checkpoint.sqlite import SqliteSaver conn = sqlite3.connect("chat.db", check_same_thread=False) agent = create_agent(llm, tools=[], system_prompt="You are a friendly assistant.", checkpointer=SqliteSaver(conn)) # Run 1 (one process): "Hi, my name is Chandu." -> Nice to meet you, Chandu! ... 2 messages # Run 2 (a new process): "What's my name?" -> I remember! Your name is Chandu! 4 messages
Memory plus a tool - what the model saw on turn 2
turn 1 human My order number is 12345. ai Sure - I'll help you with order 12345. turn 2 model received: human My order number is 12345. <- loaded from the thread ai Sure - I'll help you with order 12345. human Where is my order? ai get_order_status {'order_id': '12345'} tool Shipped - arriving Friday ai Order 12345 has shipped and arrives Friday.

Tip: Print agent.get_state(config).values["messages"] whenever memory surprises you. It is the LangChain version of printing self.messages.

Memory at a glance

InMemorySaver

Conversation memory for the running process.

create_agent(..., checkpointer=InMemorySaver())
thread_id

One conversation; required with a checkpointer.

{"configurable": {"thread_id": "user-123"}}
get_state

What the thread has stored.

agent.get_state(config).values["messages"]
trim_messages

Keep the recent part, by whole exchanges.

trim_messages(..., start_on="human")
before_model

Hook to trim before each model call.

@before_model
SummarizationMiddleware

Replace old messages with a summary.

SummarizationMiddleware(model=llm, trigger=...)
SqliteSaver

Persistent memory in a database file.

from langgraph.checkpoint.sqlite import SqliteSaver
ConversationBufferMemory

Classic and deprecated; in langchain-classic.

from langchain_classic.memory import ...

Try it yourself

The code does not change. Swap the content string and the program does something else entirely.

Two users

“Tell thread "a" your name and thread "b" a different one; ask both.”

Inspect

“After three turns, print get_state(config).values["messages"] and count them.”

Trim it

“Add the keep_recent hook with max_tokens=3 and find the first turn where the agent forgets your name.”

Restart

“Switch to SqliteSaver, tell it your name, stop the program, start it again and ask.”

What usually goes wrong

Importing memory classes from langchain.memory

That module does not exist in LangChain 1.x. Use a checkpointer - or langchain-classic for old code.

✗ from langchain.memory import ConversationBufferMemory
✓ create_agent(..., checkpointer=InMemorySaver())
One thread for everyone

A fixed thread_id shares one conversation between all users. Use the user or session id.

✗ {"configurable": {"thread_id": "chat"}}
✓ {"configurable": {"thread_id": session_id}}
Expecting InMemorySaver to persist

It lives in the process. A restart, or a new saver, starts every thread empty.

✗ checkpointer=InMemorySaver()   # for production users
✓ checkpointer=SqliteSaver(conn)  # or another database-backed saver
Trimming from the middle of an exchange

Counting messages can keep an answer without its question. Start the kept history on a human message.

✗ [messages[0]] + messages[-(n - 1):]
✓ trim_messages(..., include_system=True, start_on="human")
Trusting a summary to keep the facts

Our summary kept the name and dropped the city. Keep facts you need somewhere explicit.

Key points

  • A model call is stateless; memory is the history your application sends back.
  • ConversationBufferMemory and its relatives are deprecated; they live in langchain-classic.
  • In LangChain 1.x: create_agent with a checkpointer, and a thread_id per conversation.
  • The system prompt is added on every call, so trimming cannot remove it.
  • History costs tokens on every turn: 35 to 141 in five short turns.
  • Trimming forgets the oldest messages; summarizing forgets what the summary leaves out.
  • InMemorySaver disappears with the process; SqliteSaver survives a restart.

Quick check before you move on

Does an LLM automatically remember previous calls?
No. Each call is independent; the application has to send the earlier messages.
Where does basic conversation memory live?
In the application: a message list in Module 1, a checkpointer thread in LangChain.
Why trim conversation history?
Every turn resends it, so cost and latency grow, and past the context window the oldest text is dropped.
Why must the system message survive trimming?
It holds the agent’s instructions. With create_agent it is added on every call, so it cannot be trimmed away.
Conversation memory vs persistent memory?
Conversation memory lives in the running process (self.messages, InMemorySaver). Persistent memory is stored outside (SqliteSaver, a database) and survives a restart.
Why trim by whole exchanges?
A question and its answer belong together; message counting can keep an answer and drop the question.
What happens to the old conversation with a new ChatAgent() - or a new InMemorySaver?
Nothing carries over. The new one starts empty.

Quiz

  1. 1.

    You call agent.invoke() with a checkpointer but no config. What happens?

  2. 2.

    Same agent, thread user-123 knows the name. What does thread user-456 say?

  3. 3.

    After trimming to 5 messages, the agent forgot the name but remembered the city. Why?

  4. 4.

    Which classic class matches ConversationBufferWindowMemory(k=2) today?

Interview questions

Does an LLM have memory?

Not between calls. The application creates the appearance of memory by sending previous messages with each request.

What was ConversationBufferMemory, and what replaced it?

A classic LangChain class that stored the full conversation and injected it into the prompt. In LangChain 1.x, agents use a checkpointer keyed by thread_id; the old classes are deprecated in langchain-classic.

How would you manage long conversations?

Trim to recent whole exchanges, summarize older context, store important facts explicitly, and test what the agent still knows afterwards - each technique forgets something.

How do you implement persistent memory?

Use a database-backed checkpointer (SQLite, Postgres) for conversations, and a separate store for long-term facts, loaded when relevant.

How do you keep users’ conversations separate?

Give each conversation its own thread_id - never a shared global history.

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...