← Back to Agentic AI map
Lesson 1.8 · Agents From Scratch

Conversation Memory

Memory is just a list of past messages you keep sending back - and trim before it gets too big.

memory

What you will be able to do

  • Explain why a model does not remember previous calls
  • Give an agent memory by keeping and resending the message list
  • Explain the context window, and what happens when a conversation outgrows it
  • Trim history by whole exchanges while keeping the system message
  • Predict what an agent forgets when history is trimmed
  • Keep each conversation separate by giving each agent its own list

The idea, in plain English

Each call to the model starts from nothing. Say "My name is Chandu" in one call and ask "What’s my name?" in the next, and the model has no idea - nothing about the first call was kept anywhere. We tried exactly that: "I’m afraid I don’t know your name!"

So "memory" is less magical than it sounds. You keep a running list of every message so far, and send the whole list every time. The model appears to remember because the history is right there in front of it. The application remembers; the model just reads.

That list cannot grow forever. A model can only take in so much text at once - its context window - and every message is processed again on every call. So you trim: keep the system message, which holds the instructions, plus the most recent exchanges, and drop the oldest.

Trimming is forgetting. Whatever you drop is gone as far as the model is concerned - in our runs, after five small questions the agent no longer knew the name it was told at the start. Wrapping the list in a class keeps it tied to one conversation: two ChatAgent objects, two separate memories.

Worked example: A chat agent that remembers your name.

workflowHow a conversation grows - and what trimming dropsstep 1 / 4

Call 1 - the introduction

Two messages go out: the system message and "Hi, my name is Chandu." The reply is appended to the list by your code.

messages sent
2
prompt tokens
29
list after
3 messages
knows the name
from now on

A real run with max_history=10. Each step shows what is sent to the model on that call, and the prompt tokens it processed.

Memory is the message list

The messages list from Lesson 1.1 is the whole mechanism. Append the user’s message, send the list, append the reply. Next turn, the list is two messages longer and the model sees everything that came before.

Nothing is stored on the model side. Two independent calls know nothing about each other, and a brand-new ChatAgent - a new list - does not know your name either. Memory belongs to the list, and the list belongs to the agent instance.

Every call pays for the whole history

The model reads the entire list on every call. In our run the prompt tokens processed per call went 29, 67, 90, 111, 131 - each question more expensive than the last, even though every question was a single short line.

Locally that costs time; on a paid API it costs money per token. Either way, an untrimmed conversation gets steadily slower until it hits a hard limit.

The context window - and silent truncation

The context window is the most text a model can take in at once, measured in tokens. llama3 was trained for 8,192 - but Ollama ran it with 4,096 by default on our machine, as its own log shows.

Exceed the window and nothing errors. We sent a conversation of about 5,700 tokens: a name, then three long pasted texts, then "What’s my name?". Ollama silently dropped the oldest messages to make it fit - only 47 tokens were processed - and the model said it did not have that information.

Raising the window with options={"num_ctx": 8192} let all 5,668 tokens through. The model still could not find the name, buried under pages of filler. A bigger window is not better memory: the model has to notice what matters, and a long history makes that harder.

Watch out: Ollama does not raise an error when a conversation is too long - it drops the oldest messages without telling you. If an agent suddenly forgets something, check the length before you blame the model.

Trimming, done by whole exchanges

The course version keeps the system message plus the last max_history - 1 messages. With max_history=10 that is 9 messages - an odd number, so the oldest one kept is always an assistant reply whose question was dropped. We checked: after trimming, the roles ran system, assistant, user, assistant...

It rarely breaks anything, but it is easy to avoid. Count exchanges instead of messages - a question and its reply - and keep the last few whole ones: history[-2 * max_exchanges:].

Whichever way you count, keep the system message out of the trim. It holds the instructions; lose it and the agent forgets how it is supposed to behave partway through a long chat.

What happens when a call fails

ask() appends the user’s message before calling the model. If that call fails - Ollama not running, a timeout - the question stays in the list with no answer. On the next call the model receives two user messages in a row, the first of which never got a reply. We reproduced exactly that: the roles sent were system, user, user.

Remove the unanswered message when the call fails, and the history stays a clean record of questions and their answers.

This memory ends with the program

self.messages lives in memory. Close the program and the conversation is gone; start again and the agent is a blank slate. That is conversation memory.

Persistent memory - facts that survive a restart - needs a file or a database. POC 2 at the end of this module adds exactly that with a notes.json file, and Module 2 covers summarising old turns instead of simply dropping them.

Two kinds of memory
Conversation memoryself.messages - gone when the program stops.
Persistent memoryA file or database - survives restarts (POC 2).

Step-by-step code

A chat agent with memory and trimming
import ollama class ChatAgent: def __init__(self, model="llama3.1", max_history=10): self.model = model self.max_history = max_history self.messages = [{"role": "system", "content": "You are a friendly assistant."}] def ask(self, user_input): self.messages.append({"role": "user", "content": user_input}) response = ollama.chat(model=self.model, messages=self.messages) reply = response["message"]["content"] self.messages.append({"role": "assistant", "content": reply}) # trim: keep system message + last N messages if len(self.messages) > self.max_history: self.messages = [self.messages[0]] + self.messages[-(self.max_history - 1):] return reply agent = ChatAgent() print(agent.ask("Hi, my name is Chandu.")) print(agent.ask("What's my name?"))
Output - from real runs
With the history: Nice to meet you, Chandu! I'm happy to chat with you. ... I remember! Your name is Chandu! Two independent calls, no history: I'm afraid I don't know your name! ... A new ChatAgent(): I'm happy to help! However, I don't actually know your name. ...
What memory costs, and what trimming forgets - max_history=10
call 1 2 messages sent 29 prompt tokens "Hi, my name is Chandu." call 2 4 messages sent 67 prompt tokens "Name a fruit." call 3 6 messages sent 90 prompt tokens "Name a colour." call 4 8 messages sent 111 prompt tokens "Name a planet." call 5 10 messages sent 131 prompt tokens "Name a river." call 6 11 messages sent 136 prompt tokens "Name a mountain." <- trimmed after this call 7 11 messages sent 116 prompt tokens "What's my name?" Reply: "I don't think you told me your name! ..." Roles after trimming: system, assistant, user, assistant, ... <- an orphaned reply Same conversation with max_history=40: Reply: "I remember! Your name is Chandu!"
Better: trim whole exchanges, and clean up after a failed call
import ollama class ChatAgent: def __init__(self, model="llama3.1", max_exchanges=4): self.model = model self.max_exchanges = max_exchanges # one exchange = a question and its reply self.messages = [{"role": "system", "content": "You are a friendly assistant."}] def ask(self, user_input): self.messages.append({"role": "user", "content": user_input}) try: response = ollama.chat(model=self.model, messages=self.messages) except Exception: self.messages.pop() # never leave a question without its answer raise reply = response["message"]["content"] self.messages.append({"role": "assistant", "content": reply}) # Keep the system message and the last N whole exchanges. system, history = self.messages[0], self.messages[1:] self.messages = [system] + history[-2 * self.max_exchanges:] return reply agent = ChatAgent() print(agent.ask("Hi, my name is Chandu.")) print(agent.ask("What's my name?"))
Output
Nice to meet you, Chandu! I'm happy to chat with you. ... I remember! Your name is Chandu! After six exchanges, roles: system, user, assistant, user, assistant, ... <- whole exchanges After a failed call: the unanswered question is removed The original class after a failed call sends: system, user, user
When the conversation outgrows the window
response = ollama.chat( model="llama3.1", messages=messages, # about 5,700 tokens of history options={"num_ctx": 8192}, # Ollama's default here was 4,096 ) print(response["prompt_eval_count"]) # tokens the model actually processed # default window: 47 processed - the oldest messages were silently dropped # num_ctx 8192: 5668 processed - everything went in, and the name was still missed

Tip: response["prompt_eval_count"] tells you how many tokens the model actually processed. Print it while you develop - it is the quickest way to see a history growing, or being cut.

Watch out: Trimming is forgetting. Anything the agent must always know - the user’s name, a goal, a constraint - belongs in the system message or in persistent memory, not in an old turn that will be dropped.

Memory at a glance

self.messages

The conversation - the only memory there is.

self.messages.append({...})
System message

Instructions; never trimmed.

self.messages[0]
Context window

The most tokens a model takes in at once.

options={"num_ctx": 8192}
prompt_eval_count

Tokens actually processed on a call.

response["prompt_eval_count"]
Trim by exchanges

Keep the last N question-and-answer pairs.

history[-2 * max_exchanges:]

Try it yourself

The code does not change. Swap the content string and the program does something else entirely.

The basic test

“Hi, my name is Chandu. -> What's my name?”

Make it forget

“Tell it your name, ask five unrelated questions, then ask your name again.”

Two agents

“Tell agent1 your name, then ask agent2 what it is.”

Watch the cost

“Print response["prompt_eval_count"] on every call and watch it grow.”

What usually goes wrong

Expecting the model to remember

Every call is independent. Without the history in messages, the model knows nothing about earlier turns.

Trimming away the system message

Slicing the whole list drops the instructions along with the oldest turns.

✗ self.messages = self.messages[-10:]
✓ self.messages = [self.messages[0]] + self.messages[1:][-8:]
Trimming an odd number of messages

The oldest message kept is a reply whose question was dropped. Count whole exchanges.

✗ self.messages[-(self.max_history - 1):]
✓ history[-2 * self.max_exchanges:]
Leaving a question without an answer

If the call fails after the user message is appended, the next call sends two user messages in a row.

✗ self.messages.append(user_message)
response = ollama.chat(...)
✓ try:
    response = ollama.chat(...)
except Exception:
    self.messages.pop()
    raise
Assuming a long history is all being read

Past the context window, Ollama drops the oldest messages silently. Check prompt_eval_count.

Keeping important facts only in old turns

They are the first thing trimming removes. Put what must persist in the system message or a file.

Key points

  • A model remembers nothing between calls; memory is the history your code resends.
  • The message list is the memory, and each ChatAgent has its own.
  • Every call reprocesses the whole list - in our run 29, 67, 90, 111, 131 tokens.
  • Past the context window, Ollama silently drops the oldest messages.
  • A bigger window is not better memory - facts can be missed in a long history.
  • Trim by whole exchanges, and never trim the system message.
  • Trimming is forgetting: after five small questions, the agent lost the name.
  • This memory ends with the program; persistent memory needs a file (POC 2).

Quick check before you move on

Where does memory actually live?
In self.messages - the list sent to the model on every call.
Why trim the history?
Every call reprocesses it, so a long history gets slower and more expensive, and eventually exceeds the context window.
What does a new ChatAgent() know about a previous one?
Nothing - it starts with a fresh list.
What happens when a conversation is longer than the context window?
Ollama silently drops the oldest messages so the rest fits. Nothing errors.

Quiz

  1. 1.

    Where does "memory" actually live in this code?

  2. 2.

    Why do we trim self.messages instead of letting it grow forever?

  3. 3.

    What would happen if you created a new ChatAgent() before asking "What’s my name?"

  4. 4.

    With max_history=10, why is the oldest kept message an assistant reply?

  5. 5.

    An agent was told a fact early on and later cannot recall it. Name two possible causes.

Interview questions

Does an LLM really have memory?

Not between requests. In a basic chat application the model is stateless: the application keeps the conversation history and sends it with each new request. As the history grows it has to be trimmed or otherwise managed to stay within the context window.

What happens when a conversation exceeds the context window?

Depending on the runtime it either errors or truncates. Ollama drops the oldest messages silently, so the model loses early context without any signal. You detect it by counting tokens, and manage it by trimming, summarising, or moving durable facts into a system message or external store.

How would you give an agent memory that survives restarts?

Store it outside the process - a file or database - and load the relevant parts into the prompt on each call. Keep short-term conversation history separate from long-term facts, which should be saved deliberately rather than kept as old turns.

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...