← Back to Agentic AI map
Lesson 1.12 · Agents From Scratch

Module 1 Projects

Three small projects that put the whole module together - built, run, and debugged.

build

What you will be able to do

  • Assemble a complete, runnable agent from the pieces built in Lessons 1.5 to 1.10
  • Give an agent memory that survives a restart, and keep that memory clean
  • Ground an agent in a tool, so its facts come from somewhere you control
  • Debug a real agent by reading its transcript, not by guessing

How to use this

Nothing new is introduced here. Each project combines lessons you have already worked through - the robust loop, real tools, memory, the hand-off between agents - into something you can run and use.

Every project below was built and run against a local model, and every one needed fixing before it worked. That is the most realistic part of the chapter: the first version of the notes agent saved its facts correctly and then never looked at them again. Reading the transcript showed why, and the fix taught something none of the earlier lessons did.

Build them in order. The first is mostly assembly, the second adds memory that outlives the program, and the third is the shape a real web-search agent takes. Do not reach for LangChain yet - the point of Module 1 is to know what is underneath it.

What each project proves

POC 1 answers "can I build and run an agent end to end?" - a terminal, a loop, a tool, and an answer. POC 2 answers "can it remember across restarts?" - and, as it turned out, "what happens when it remembers something wrong?". POC 3 answers "can agents cooperate on facts I can trust?" - and shows how far a stub search tool gets you.

Three projects, three lessons
POC 1 - CalculatorLLM + tools + the robust loop, in a terminal.
POC 2 - NotesMemory in a file that outlives the program.
POC 3 - ResearchTwo agents, with the facts from a tool.

Recall should not be optional

The notes agent was first built as described: two tools, remember(fact) and recall(). Saving worked - "I am learning Python" landed in notes.json. Then a fresh process was asked "What am I learning?" and answered "I’m not sure what you’re learning yet!" - word for word what it said with no notes file at all. The memory was on disk. The model never chose to call recall().

The fix was to stop asking. Every saved note is loaded into the system prompt when the agent starts, so the model always sees what it knows. remember stays a tool, because deciding what is worth saving is the model’s job; reading the notes is not a decision, so it is no longer left to one. After the change, a brand-new process answered "You are learning Python!" - and recommended VS Code for a Python course because that was the saved favourite editor.

Persistent memory persists mistakes

Asked "What do you know about me?", the model used remember as if it were recall, and twice saved the placeholder word "fact" - remember(fact), copied straight from the tool description. Conversation memory forgets that kind of slip when the program ends. A notes file keeps it forever, and loads it into every future conversation.

So the tool guards the door: remember refuses anything under three words, and tells the model what a fact should look like. A model will sometimes write junk; the file should not have to keep it.

Watch out: Anything an agent writes to persistent storage, it will read back in every later session. Validate in the tool before saving - and give yourself a way to inspect and edit the file.

Source tags make made-up facts visible

The research agent is told to use only the search results and to end every bullet with its [source]. One bullet came back as "You have control over the processing and data remains local. [no specific source mentioned]" - the model labelled its own unsupported claim.

That label is what makes it removable. cited_only() keeps only bullets that end with the tag of a real result, so an unsourced claim never reaches the writer. And when the search returns nothing, the code says so instead of letting the model fill the gap from memory.

The limit: the writer still embellished the clean notes - "protected from unauthorized access and potential data breaches", "financial or medical data". Grounding one agent does not ground the next. Each stage that writes new sentences needs the same treatment, or a person reading the result.

POC 1 - Terminal Calculator Agent

Lessons 1.5 to 1.7 and 1.10

A ReAct-loop agent with a calculate tool, run from the terminal so you can keep asking it questions until you type quit. It is assembly: the safe calculator from 1.6, the robust loop from 1.10, the "say so when no tool helps" line, and a small input loop around them. Every answer below came from the calculator - including 4783 * 2917, which the model got wrong every time without a tool in Lesson 1.4.

calculator_agent.py
"""POC 1 - Terminal Calculator Agent: Lessons 1.5 to 1.7 and 1.10 in one file.""" import ast import operator import re import ollama MODEL = "llama3.1" # --- The tool (Lesson 1.6's safe calculator: arithmetic only, no eval) --- OPERATORS = { ast.Add: operator.add, ast.Sub: operator.sub, ast.Mult: operator.mul, ast.Div: operator.truediv, ast.USub: operator.neg, } def calculate(expression): def evaluate(node): if isinstance(node, ast.Constant) and isinstance(node.value, (int, float)): return node.value if isinstance(node, ast.BinOp) and type(node.op) in OPERATORS: return OPERATORS[type(node.op)](evaluate(node.left), evaluate(node.right)) if isinstance(node, ast.UnaryOp) and type(node.op) in OPERATORS: return OPERATORS[type(node.op)](evaluate(node.operand)) raise ValueError("only numbers and + - * / are allowed") result = evaluate(ast.parse(expression, mode="eval").body) return str(int(result) if result == int(result) else result) TOOLS = {"calculate": calculate} SYSTEM_PROMPT = """You solve problems step by step using this format: Thought: what you're thinking Action: tool_name(argument) Observation: (this will be filled in for you) ... repeat Thought/Action/Observation as needed ... Final Answer: your final answer to the user Available tools: - calculate(expression): evaluates an arithmetic expression such as 47 * 89 Only output ONE Thought/Action pair at a time, then stop and wait for the Observation. If the question needs no tool, or the tool cannot help, answer straight away with a Final Answer that says so.""" ACTION = re.compile(r"Action:\s*(\w+)\((.*?)\)") # --- The loop (Lesson 1.10's run_agent_safe) --- def run_agent_safe(user_question, max_steps=5, max_retries_per_step=2): messages = [ {"role": "system", "content": SYSTEM_PROMPT}, {"role": "user", "content": user_question}, ] for step in range(max_steps): text = None for retry in range(max_retries_per_step): response = ollama.chat(model=MODEL, messages=messages, options={"temperature": 0}) candidate = response["message"]["content"] if "Final Answer:" in candidate or ACTION.search(candidate): text = candidate break messages.append({"role": "user", "content": "That wasn't in the right format. Use Thought/Action or Final Answer."}) if text is None: return "Agent failed to produce a valid step after retries." messages.append({"role": "assistant", "content": text}) if "Final Answer:" in text: return text.split("Final Answer:")[-1].strip() match = ACTION.search(text) tool_name, arg = match.group(1), match.group(2).strip("\"' ") if tool_name not in TOOLS: observation = f"error: '{tool_name}' is not a valid tool. Available tools: {list(TOOLS)}" else: try: observation = TOOLS[tool_name](arg) except Exception as e: observation = f"error running {tool_name}: {e}" messages.append({"role": "user", "content": f"Observation: {observation}"}) return "Max steps reached without a final answer." # --- The terminal --- if __name__ == "__main__": print("Calculator Agent - type 'quit' to exit") while True: q = input("\nYou: ") if q.strip().lower() == "quit": break print("Agent:", run_agent_safe(q))
A real session
Calculator Agent - type 'quit' to exit You: What is 47 * 89? Agent: The final answer is 4183. calculate("47 * 89") You: What is 1000 / 25? Agent: The result of 1000 divided by 25 is 40. calculate("1000 / 25") You: What is 4783 * 2917? Agent: The product of 4783 and 2917 is 13952011. calculate("4783 * 2917") You: What is 15% of 2400? Agent: The answer is 360. calculate("2400 * 0.15") You: What is the capital of France? Agent: Paris no tool needed You: quit

POC 2 - Personal Notes Agent

Lesson 1.8, extended

An agent with a notebook that outlives the program: a remember(fact) tool that saves to notes.json, and every saved note loaded into the system prompt at start-up. The first version left recall to the model - it saved facts and then never read them. The version below is what worked, run as five separate processes.

notes_agent.py
"""POC 2 - Personal Notes Agent: memory that survives a restart.""" import json import re import sys from pathlib import Path import ollama MODEL = "llama3.1" NOTES_FILE = Path("notes.json") # --- The tools: the only code that touches the file --- def load_notes(): if not NOTES_FILE.exists(): return [] return json.loads(NOTES_FILE.read_text(encoding="utf-8")) def remember(fact): fact = fact.strip() # Saved notes come back in every future session, so refuse junk at the door. if len(fact.split()) < 3: return f"error: '{fact}' is not a fact. Save a short sentence, such as: I am learning Python" notes = load_notes() if fact not in notes: notes.append(fact) NOTES_FILE.write_text(json.dumps(notes, indent=2), encoding="utf-8") return f"saved: {fact}" def recall(_=""): notes = load_notes() return "\n".join(f"- {note}" for note in notes) if notes else "no notes saved yet" TOOLS = {"remember": remember} PROMPT_TEMPLATE = """You are a personal assistant with a notebook that lasts between sessions. What you already know about the user, from earlier sessions: {notes} Use this format: Thought: what you're thinking Action: tool_name(argument) Observation: (this will be filled in for you) Final Answer: your answer to the user Available tools: - remember(fact): saves one short fact about the user, such as remember(I am learning Python) When the user tells you something new about themselves, remember it, then give a Final Answer. When the user asks a question, answer it with a Final Answer, using what you already know. Only output ONE Thought/Action pair at a time, then stop and wait for the Observation.""" def system_prompt(): # Recall is not left to the model: every saved note goes into every conversation. return PROMPT_TEMPLATE.format(notes=recall()) ACTION = re.compile(r"Action:\s*(\w+)\((.*?)\)") def run_agent(user_question, max_steps=5): messages = [ {"role": "system", "content": system_prompt()}, {"role": "user", "content": user_question}, ] for step in range(max_steps): response = ollama.chat( model=MODEL, messages=messages, options={"temperature": 0, "stop": ["Observation:"]}, ) text = response["message"]["content"] messages.append({"role": "assistant", "content": text}) match = ACTION.search(text) if match: # run any Action before accepting a Final Answer tool_name = match.group(1) # accept remember(fact="...") as well as remember(...) arg = re.sub(r"^\w+\s*=\s*", "", match.group(2)).strip("\"' ") if tool_name in TOOLS: observation = TOOLS[tool_name](arg) else: observation = f"error: '{tool_name}' is not a valid tool. Available tools: {list(TOOLS)}" messages.append({"role": "user", "content": f"Observation: {observation}"}) continue if "Final Answer:" in text: return text.split("Final Answer:")[-1].strip() # Neither a tool call nor an answer (often "Action: None"): nudge, as in Lesson 1.10. messages.append({"role": "user", "content": "No tool is needed now. Reply with Final Answer: and your answer."}) return "Max steps reached without a final answer." if __name__ == "__main__": print("Notes Agent:", run_agent(" ".join(sys.argv[1:])))
Five separate runs - the memory lives in the file
$ python notes_agent.py "Remember that I am learning Python." Notes Agent: It's great that you're learning Python! What specific aspects ... $ python notes_agent.py "Also remember that my favourite editor is VS Code." Notes Agent: It seems that you are learning Python and have a preference for VS Code ... $ cat notes.json ["I am learning Python", "favourite editor=VS Code"] $ python notes_agent.py "What am I learning?" Notes Agent: You are learning Python! $ python notes_agent.py "What do you know about me?" Notes Agent: So, you're learning Python and you use VS Code as your favorite editor. $ python notes_agent.py "Which editor should I use for my Python course?" Notes Agent: Based on what I know, I would recommend using VS Code for your Python course ... With notes.json removed, "What am I learning?" gets: Notes Agent: You didn't mention what you're learning, so I don't have an answer yet!
What the first version did
With recall() as a tool the model had to choose to call: $ python notes_agent.py "What am I learning?" # notes.json had the fact Notes Agent: Thought: I'm not sure what you're learning yet! Asked "What do you know about me?": Action: remember(fact) <- saved the word "fact" Action: remember(I am learning Python) <- "recalled" by saving again Action: remember(fact) -> notes.json: ["I am learning Python", "favourite editor=VS Code", "fact"] Fixes: load the notes into the prompt; refuse facts under three words.

POC 3 - Research-and-Summarize Agent

Lesson 1.11, extended

The made-up researcher from Lesson 1.11, replaced with one that calls a search(query) tool. A stub with fixed results stands in for a real search API - what matters is that the notes come from the tool. Every bullet must cite its source, unsourced bullets are dropped in code, and no results means no summary. Swap the stub for a real search and this is the shape of a web-search agent.

research_agent.py
"""POC 3 - Research-and-Summarize Agent: the notes come from a tool, not from memory.""" import ollama MODEL = "llama3.1" # --- A stub search tool: fixed results stand in for a real search API --- SEARCH_INDEX = { "local llm privacy": [ {"source": "ollama-docs", "text": "Ollama runs language models on your own computer and serves them at http://localhost:11434."}, {"source": "privacy-guide", "text": "With a local model, prompts and documents are processed on your machine and are not sent to a third-party API."}, {"source": "hardware-notes", "text": "The llama3 8B model is a 4.7 GB download and runs on a laptop with 8 GB of RAM or more."}, ], } def search(query): words = set(query.lower().split()) for key, results in SEARCH_INDEX.items(): if set(key.split()) <= words: return results return [] def cited_only(notes, results): """Keep only bullets that end with the tag of a real search result.""" tags = {f"[{r['source']}]" for r in results} return "\n".join( line for line in notes.splitlines() if any(line.rstrip().endswith(tag) for tag in tags) ) def researcher_agent(topic): results = search(topic) if not results: return "NO RESULTS" # say so - never let the model fill the gap from memory sources = "\n".join(f"[{r['source']}] {r['text']}" for r in results) response = ollama.chat( model=MODEL, messages=[ {"role": "system", "content": ( "You are a researcher. Turn the search results into short bullet-point notes. " "Use ONLY facts stated in the results, and end each bullet with its [source]. " "Do not add anything the results do not say." )}, {"role": "user", "content": f"Topic: {topic}\n\nSearch results:\n{sources}"}, ], options={"temperature": 0}, ) notes = response["message"]["content"] return cited_only(notes, results) # an unsourced claim never reaches the writer def writer_agent(topic, research_notes): response = ollama.chat( model=MODEL, messages=[ {"role": "system", "content": ( "You are a writer. Turn research notes into a short, friendly 2-paragraph summary. " "Use only the facts in the notes." )}, {"role": "user", "content": f"Topic: {topic}\n\nResearch notes:\n{research_notes}"}, ], options={"temperature": 0}, ) return response["message"]["content"] def research_and_summarize(topic): notes = researcher_agent(topic) if notes == "NO RESULTS": return f"I found no search results for '{topic}', so there is nothing to summarise." print("--- Research notes ---\n" + notes) return writer_agent(topic, notes) if __name__ == "__main__": print("\n--- Summary ---\n" + research_and_summarize("Benefits of a local LLM for privacy")) print("\n--- Summary ---\n" + research_and_summarize("Quantum computing in banking"))
Output - from a real run
What the researcher wrote, before filtering: Here are the short bullet-point notes on the benefits of a local LLM for privacy: • Processing occurs on your machine, keeping prompts and documents private. [privacy-guide] • No prompts or documents are sent to a third-party API. [privacy-guide] • You have control over the processing and data remains local. [no specific source mentioned] Note: These notes only summarize the information provided in the search results ... What reached the writer, after cited_only(): • Processing occurs on your machine, keeping prompts and documents private. [privacy-guide] • No prompts or documents are sent to a third-party API. [privacy-guide] --- Summary --- When you use a local LLM, you can rest assured that your prompts and documents are kept private and secure. ... protected from unauthorized access and potential data breaches. <- the writer still adds things --- Summary --- I found no search results for 'Quantum computing in banking', so there is nothing to summarise.

Key points

  • POC 1 is assembly: the safe calculator, the robust loop, and a terminal - and every answer came from the tool.
  • POC 2 moves memory into a file, so it survives a restart.
  • Do not make recall a decision: load saved notes into the prompt every time.
  • Persistent memory persists mistakes - validate before saving.
  • POC 3 grounds the researcher in a search tool; source tags make unsupported claims visible and removable.
  • Grounding one agent does not ground the next - the writer still embellished.
  • Debug by reading transcripts; every fix here came from one.

Quick check before you move on

What does POC 1 build?
A terminal ReAct agent with a calculate tool, built on the robust loop from Lesson 1.10.
What tools does the notes agent end up with?
Only remember(fact). Recall is not a tool any more - saved notes go into the system prompt on every start.
Where is persistent memory stored?
In a local notes.json file.
Why replace the made-up researcher with search(query)?
So the notes come from a tool you control rather than from the model’s memory.
Which project extends Lesson 1.11?
POC 3, the research-and-summarize agent.

Quiz

  1. 1.

    The notes agent saved "I am learning Python" but, in a new session, said it did not know what you were learning. Why?

  2. 2.

    Why is a bad entry in notes.json worse than a bad message in conversation memory?

  3. 3.

    How does POC 3 stop an unsupported claim reaching the writer?

  4. 4.

    POC 3’s researcher is grounded. Is the final summary therefore fully grounded?

Interview questions

How would you give an agent long-term memory?

Store facts outside the process - a file or database - and load the relevant ones into the prompt on each call rather than relying on the model to fetch them. Validate before writing, because anything saved will be read back in every later session.

How do you keep a research agent from inventing facts?

Give it a retrieval tool, tell it to use only retrieved content, require a source on every claim, drop claims without a valid source in code, and refuse to answer when retrieval returns nothing. Apply the same discipline to any later stage that writes prose.

How do you debug an agent that misbehaves?

Log every model reply and every tool call, read the transcript, and find the first step that went wrong. Most failures - a skipped tool, a placeholder argument, an invented observation - are obvious once you see the exact text.

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...