Conversation Memory
Memory is just a list of past messages you keep sending back - and trim before it gets too big.
What you will be able to do
- Explain why a model does not remember previous calls
- Give an agent memory by keeping and resending the message list
- Explain the context window, and what happens when a conversation outgrows it
- Trim history by whole exchanges while keeping the system message
- Predict what an agent forgets when history is trimmed
- Keep each conversation separate by giving each agent its own list
The idea, in plain English
Each call to the model starts from nothing. Say "My name is Chandu" in one call and ask "What’s my name?" in the next, and the model has no idea - nothing about the first call was kept anywhere. We tried exactly that: "I’m afraid I don’t know your name!"
So "memory" is less magical than it sounds. You keep a running list of every message so far, and send the whole list every time. The model appears to remember because the history is right there in front of it. The application remembers; the model just reads.
That list cannot grow forever. A model can only take in so much text at once - its context window - and every message is processed again on every call. So you trim: keep the system message, which holds the instructions, plus the most recent exchanges, and drop the oldest.
Trimming is forgetting. Whatever you drop is gone as far as the model is concerned - in our runs, after five small questions the agent no longer knew the name it was told at the start. Wrapping the list in a class keeps it tied to one conversation: two ChatAgent objects, two separate memories.
Worked example: A chat agent that remembers your name.
Call 1 - the introduction
Two messages go out: the system message and "Hi, my name is Chandu." The reply is appended to the list by your code.
A real run with max_history=10. Each step shows what is sent to the model on that call, and the prompt tokens it processed.
Memory is the message list
The messages list from Lesson 1.1 is the whole mechanism. Append the user’s message, send the list, append the reply. Next turn, the list is two messages longer and the model sees everything that came before.
Nothing is stored on the model side. Two independent calls know nothing about each other, and a brand-new ChatAgent - a new list - does not know your name either. Memory belongs to the list, and the list belongs to the agent instance.
Every call pays for the whole history
The model reads the entire list on every call. In our run the prompt tokens processed per call went 29, 67, 90, 111, 131 - each question more expensive than the last, even though every question was a single short line.
Locally that costs time; on a paid API it costs money per token. Either way, an untrimmed conversation gets steadily slower until it hits a hard limit.
The context window - and silent truncation
The context window is the most text a model can take in at once, measured in tokens. llama3 was trained for 8,192 - but Ollama ran it with 4,096 by default on our machine, as its own log shows.
Exceed the window and nothing errors. We sent a conversation of about 5,700 tokens: a name, then three long pasted texts, then "What’s my name?". Ollama silently dropped the oldest messages to make it fit - only 47 tokens were processed - and the model said it did not have that information.
Raising the window with options={"num_ctx": 8192} let all 5,668 tokens through. The model still could not find the name, buried under pages of filler. A bigger window is not better memory: the model has to notice what matters, and a long history makes that harder.
Watch out: Ollama does not raise an error when a conversation is too long - it drops the oldest messages without telling you. If an agent suddenly forgets something, check the length before you blame the model.
Trimming, done by whole exchanges
The course version keeps the system message plus the last max_history - 1 messages. With max_history=10 that is 9 messages - an odd number, so the oldest one kept is always an assistant reply whose question was dropped. We checked: after trimming, the roles ran system, assistant, user, assistant...
It rarely breaks anything, but it is easy to avoid. Count exchanges instead of messages - a question and its reply - and keep the last few whole ones: history[-2 * max_exchanges:].
Whichever way you count, keep the system message out of the trim. It holds the instructions; lose it and the agent forgets how it is supposed to behave partway through a long chat.
What happens when a call fails
ask() appends the user’s message before calling the model. If that call fails - Ollama not running, a timeout - the question stays in the list with no answer. On the next call the model receives two user messages in a row, the first of which never got a reply. We reproduced exactly that: the roles sent were system, user, user.
Remove the unanswered message when the call fails, and the history stays a clean record of questions and their answers.
This memory ends with the program
self.messages lives in memory. Close the program and the conversation is gone; start again and the agent is a blank slate. That is conversation memory.
Persistent memory - facts that survive a restart - needs a file or a database. POC 2 at the end of this module adds exactly that with a notes.json file, and Module 2 covers summarising old turns instead of simply dropping them.
Conversation memoryself.messages - gone when the program stops.Persistent memoryA file or database - survives restarts (POC 2).Step-by-step code
import ollama
class ChatAgent:
def __init__(self, model="llama3.1", max_history=10):
self.model = model
self.max_history = max_history
self.messages = [{"role": "system", "content": "You are a friendly assistant."}]
def ask(self, user_input):
self.messages.append({"role": "user", "content": user_input})
response = ollama.chat(model=self.model, messages=self.messages)
reply = response["message"]["content"]
self.messages.append({"role": "assistant", "content": reply})
# trim: keep system message + last N messages
if len(self.messages) > self.max_history:
self.messages = [self.messages[0]] + self.messages[-(self.max_history - 1):]
return reply
agent = ChatAgent()
print(agent.ask("Hi, my name is Chandu."))
print(agent.ask("What's my name?"))With the history:
Nice to meet you, Chandu! I'm happy to chat with you. ...
I remember! Your name is Chandu!
Two independent calls, no history:
I'm afraid I don't know your name! ...
A new ChatAgent():
I'm happy to help! However, I don't actually know your name. ...call 1 2 messages sent 29 prompt tokens "Hi, my name is Chandu."
call 2 4 messages sent 67 prompt tokens "Name a fruit."
call 3 6 messages sent 90 prompt tokens "Name a colour."
call 4 8 messages sent 111 prompt tokens "Name a planet."
call 5 10 messages sent 131 prompt tokens "Name a river."
call 6 11 messages sent 136 prompt tokens "Name a mountain." <- trimmed after this
call 7 11 messages sent 116 prompt tokens "What's my name?"
Reply: "I don't think you told me your name! ..."
Roles after trimming: system, assistant, user, assistant, ... <- an orphaned reply
Same conversation with max_history=40:
Reply: "I remember! Your name is Chandu!"import ollama
class ChatAgent:
def __init__(self, model="llama3.1", max_exchanges=4):
self.model = model
self.max_exchanges = max_exchanges # one exchange = a question and its reply
self.messages = [{"role": "system", "content": "You are a friendly assistant."}]
def ask(self, user_input):
self.messages.append({"role": "user", "content": user_input})
try:
response = ollama.chat(model=self.model, messages=self.messages)
except Exception:
self.messages.pop() # never leave a question without its answer
raise
reply = response["message"]["content"]
self.messages.append({"role": "assistant", "content": reply})
# Keep the system message and the last N whole exchanges.
system, history = self.messages[0], self.messages[1:]
self.messages = [system] + history[-2 * self.max_exchanges:]
return reply
agent = ChatAgent()
print(agent.ask("Hi, my name is Chandu."))
print(agent.ask("What's my name?"))Nice to meet you, Chandu! I'm happy to chat with you. ...
I remember! Your name is Chandu!
After six exchanges, roles: system, user, assistant, user, assistant, ... <- whole exchanges
After a failed call: the unanswered question is removed
The original class after a failed call sends: system, user, userresponse = ollama.chat(
model="llama3.1",
messages=messages, # about 5,700 tokens of history
options={"num_ctx": 8192}, # Ollama's default here was 4,096
)
print(response["prompt_eval_count"]) # tokens the model actually processed
# default window: 47 processed - the oldest messages were silently dropped
# num_ctx 8192: 5668 processed - everything went in, and the name was still missedTip: response["prompt_eval_count"] tells you how many tokens the model actually processed. Print it while you develop - it is the quickest way to see a history growing, or being cut.
Watch out: Trimming is forgetting. Anything the agent must always know - the user’s name, a goal, a constraint - belongs in the system message or in persistent memory, not in an old turn that will be dropped.
Memory at a glance
self.messagesThe conversation - the only memory there is.
self.messages.append({...})System messageInstructions; never trimmed.
self.messages[0]
Context windowThe most tokens a model takes in at once.
options={"num_ctx": 8192}prompt_eval_countTokens actually processed on a call.
response["prompt_eval_count"]
Trim by exchangesKeep the last N question-and-answer pairs.
history[-2 * max_exchanges:]
Try it yourself
The code does not change. Swap the content string and the program does something else entirely.
“Hi, my name is Chandu. -> What's my name?”
“Tell it your name, ask five unrelated questions, then ask your name again.”
“Tell agent1 your name, then ask agent2 what it is.”
“Print response["prompt_eval_count"] on every call and watch it grow.”
What usually goes wrong
Every call is independent. Without the history in messages, the model knows nothing about earlier turns.
Slicing the whole list drops the instructions along with the oldest turns.
✗ self.messages = self.messages[-10:]✓ self.messages = [self.messages[0]] + self.messages[1:][-8:]The oldest message kept is a reply whose question was dropped. Count whole exchanges.
✗ self.messages[-(self.max_history - 1):]✓ history[-2 * self.max_exchanges:]If the call fails after the user message is appended, the next call sends two user messages in a row.
✗ self.messages.append(user_message)
response = ollama.chat(...)✓ try:
response = ollama.chat(...)
except Exception:
self.messages.pop()
raisePast the context window, Ollama drops the oldest messages silently. Check prompt_eval_count.
They are the first thing trimming removes. Put what must persist in the system message or a file.
Key points
- A model remembers nothing between calls; memory is the history your code resends.
- The message list is the memory, and each ChatAgent has its own.
- Every call reprocesses the whole list - in our run 29, 67, 90, 111, 131 tokens.
- Past the context window, Ollama silently drops the oldest messages.
- A bigger window is not better memory - facts can be missed in a long history.
- Trim by whole exchanges, and never trim the system message.
- Trimming is forgetting: after five small questions, the agent lost the name.
- This memory ends with the program; persistent memory needs a file (POC 2).
Quick check before you move on
Quiz
- 1.
Where does "memory" actually live in this code?
- 2.
Why do we trim self.messages instead of letting it grow forever?
- 3.
What would happen if you created a new ChatAgent() before asking "What’s my name?"
- 4.
With max_history=10, why is the oldest kept message an assistant reply?
- 5.
An agent was told a fact early on and later cannot recall it. Name two possible causes.
Interview questions
Does an LLM really have memory?
Not between requests. In a basic chat application the model is stateless: the application keeps the conversation history and sends it with each new request. As the history grows it has to be trimmed or otherwise managed to stay within the context window.
What happens when a conversation exceeds the context window?
Depending on the runtime it either errors or truncates. Ollama drops the oldest messages silently, so the model loses early context without any signal. You detect it by counting tokens, and manage it by trimming, summarising, or moving durable facts into a system message or external store.
How would you give an agent memory that survives restarts?
Store it outside the process - a file or database - and load the relevant parts into the prompt on each call. Keep short-term conversation history separate from long-term facts, which should be saved deliberately rather than kept as old turns.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...
AI
System Design
Backend
- GraphQL8 modules · 69 lessons planned
- Core Python13 modules · 75 lessons planned
- FastAPI5 sections · 20 lessons
- Node.js14 modules · 206 lessons planned
- Node.js Performance7 chapters · 36 topics
- Event Loop Lifecycle6 phases · 3 scenarios
- Docker & Containerization11 modules · 144 lessons planned
- AWS for Developers14 modules · 219 lessons planned
- CI/CD & DevOps Automation10 modules · 134 lessons planned