← Back to Agentic AI map
Lesson 1.6 · Agents From Scratch

The ReAct Loop by Hand

Thought, Action, Observation, repeat - the whole agent loop as a plain Python for loop.

core

What you will be able to do

  • Explain the ReAct pattern - Thought, Action, Observation, Final Answer
  • Build a tool registry and a system prompt that teaches the format
  • Write the loop that asks the model, runs the requested tool, and feeds the result back
  • Explain why the model must never write its own Observation - and stop it with a stop sequence
  • Explain what max_steps protects against
  • Say why eval() is unsafe, and replace it with a calculator that only does arithmetic

The idea, in plain English

ReAct is short for Reason and Act. The model writes down what it is thinking (Thought) and picks a tool (Action). Your code runs that tool and hands back the result (Observation). Round it goes until the model has what it needs and writes a Final Answer.

This is the heart of the module. Everything before it was preparation - talking to a model, getting structured output, asking for a tool - and everything after adds real tools, memory, planning, and error handling on top. Read the loop closely: the model is never in control of the program. Your for loop is. It calls the model, looks at the text, runs a tool if one was asked for, and appends the result to the conversation.

Two details carry most of the weight. The model must not write its own Observations - otherwise it will happily invent a weather report instead of waiting for yours. And max_steps caps the rounds, so a model that never reaches a Final Answer cannot loop forever.

The loop below was run many times against a local model, and the transcripts shown are real. They show it working - and failing in instructive ways a small model will, which is exactly what the rest of this module fixes.

Worked example: Build a two-tool ReAct loop from scratch.

workflowOne question, two tools, one loopstep 1 / 4

Turn 1 - Thought and Action

The model reads the question and writes: Thought: I need the weather in Mumbai. Action: get_weather(Mumbai). Generation stops there - a stop sequence cuts it off before it can write an Observation.

turn
1
model wrote
Thought + Action
tools run
0
Observation
not yet

A real run of "Should I bring an umbrella in Mumbai, and what’s 12 * 8?". Step through who does what on each turn.

The four labels

Thought is the model reasoning about what to do next. Action is the tool it wants, with its argument. Observation is the result - written by your code, never by the model. Final Answer is the model saying it is done.

The labels matter because your code reads them. Final Answer: tells the loop to stop; Action: tells it which function to run. Everything else is for you, reading the transcript.

Who writes what
Thought:Model - what it needs next.
Action:Model - the tool and argument.
Observation:Your code - the real result.
Final Answer:Model - the reply; your loop stops.

The loop, one pass at a time

Each pass does the same five things. Ask the model with the whole conversation so far. Append its reply as an assistant message. If there is an Action, run the tool through the TOOLS dictionary and append the result as "Observation: ...". If there is a Final Answer, return it. If there is neither, the agent is stuck.

The Observation goes back as a user message. The model has no other way to receive information - the conversation is its only input - which is why every result has to be written into it.

Never let the model write the Observation

Asked to use this format, a model will often write the whole exchange itself: Thought, Action, and then an Observation it made up. We checked the very first reply six times: five of them contained an Observation the model had written itself, before any tool had run.

The prompt asks it not to, and that is not enough. The reliable fix is a stop sequence: options={"stop": ["Observation:"]} tells Ollama to end generation the moment the model starts to write "Observation:". With it, none of six first replies contained one.

There is a second half to the same rule. If a reply contains both an Action and a Final Answer, the original loop checks for Final Answer first - and returns an answer built on a result that never existed, with no tool run at all. Checking for an Action first closes that gap: if the model asked for a tool, the tool runs before anything else.

Watch out: An Observation is only real if your code wrote it. Anything the model writes after "Observation:" is fiction, however plausible.

What the loop cannot fix

The loop is correct; the model is still a small model. Across our runs it sometimes wrote two Actions in one reply (the regex takes the first, and the second is never run), wrote "Action: (no action needed, just a thought)", or answered the maths half from memory instead of calling the calculator.

We ran the umbrella question eight times with the original loop and eight with both fixes. Original: both tools ran 4 times, 3 runs got stuck, 1 answered without the calculator. With the fixes: 5, 1, and 2. The fixes guarantee there are no invented Observations; they do not make an 8-billion-parameter model reliable. Lesson 1.10 handles the stuck cases by feeding the problem back to the model instead of giving up.

max_steps: the safety limit

A model that never writes Final Answer would keep the loop going forever - calling tools, costing time, possibly repeating the same Action. for step in range(max_steps) makes that impossible: after five passes the function returns "Max steps reached without a final answer".

Pick the limit from the task. Two tools and a two-part question need three passes; five leaves room for one detour. A limit is also a signal - if runs regularly hit it, the prompt or the tools need work.

The regex is deliberately simple

r"Action:\s*(\w+)\((.*?)\)" finds a tool name and everything up to the first closing bracket. That covers get_weather(Mumbai) and calculate(12 * 8), and breaks on two things worth knowing.

Nested brackets: calculate(12 * (3 + 4)) stops at the first ")" and passes "12 * (3 + 4" - an error. Keyword arguments: get_weather(city="Mumbai") passes city="Mumbai, and the weather lookup returns "unknown city". Lesson 1.5 taught the keyword form, so match what your prompt asks for - or, as Module 2 does, let structured tool calls replace the regex entirely.

Step-by-step code

The complete agent - two tools, one loop
import ollama import re def get_weather(city): fake_data = {"mumbai": "rainy, 27°C", "delhi": "sunny, 34°C"} return fake_data.get(city.lower(), "unknown city") def calculate(expression): try: return str(eval(expression)) except Exception as e: return f"error: {e}" TOOLS = {"get_weather": get_weather, "calculate": calculate} SYSTEM_PROMPT = """You solve problems step by step using this format: Thought: what you're thinking Action: tool_name(argument) Observation: (this will be filled in for you) ... repeat Thought/Action/Observation as needed ... Final Answer: your final answer to the user Available tools: - get_weather(city) - calculate(expression) Only output ONE Thought/Action pair at a time, then stop and wait for the Observation.""" def run_agent(user_question, max_steps=5): messages = [ {"role": "system", "content": SYSTEM_PROMPT}, {"role": "user", "content": user_question} ] for step in range(max_steps): response = ollama.chat(model="llama3.1", messages=messages) text = response["message"]["content"] print(text) messages.append({"role": "assistant", "content": text}) if "Final Answer:" in text: return text.split("Final Answer:")[-1].strip() match = re.search(r'Action:\s*(\w+)\((.*?)\)', text) if not match: return "Agent got stuck — no action found." tool_name, arg = match.group(1), match.group(2).strip('"\' ') if tool_name in TOOLS: observation = TOOLS[tool_name](arg) else: observation = f"unknown tool: {tool_name}" messages.append({"role": "user", "content": f"Observation: {observation}"}) return "Max steps reached without a final answer." print(run_agent("Should I bring an umbrella in Mumbai, and what's 12 * 8?"))
A real transcript - the model skipping a step
Thought: I'll check the weather in Mumbai and calculate the product of 12 and 8. Action: get_weather(Mumbai) Action: calculate(12*8) Please provide the Observation for the weather check and the result for the multiplication. Thought: With the weather observation, I think it's likely that I should bring an umbrella to stay dry in Mumbai. Action: No more action needed, the answer is clear! Final Answer: Yes, it's a good idea to bring an umbrella in Mumbai, and the answer to 12 * 8 is 96. Tools actually run: get_weather only. The "96" came from the model, not the calculator.
Two fixes to the loop
def run_agent(user_question, max_steps=5): messages = [ {"role": "system", "content": SYSTEM_PROMPT}, {"role": "user", "content": user_question} ] for step in range(max_steps): response = ollama.chat( model="llama3.1", messages=messages, options={"stop": ["Observation:"]}, # fix 1: the model cannot write its own results ) text = response["message"]["content"] print(text) messages.append({"role": "assistant", "content": text}) # fix 2: look for an Action first - never accept a Final Answer before the tool has run match = re.search(r'Action:\s*(\w+)\((.*?)\)', text) if match: tool_name, arg = match.group(1), match.group(2).strip('"\' ') if tool_name in TOOLS: observation = TOOLS[tool_name](arg) else: observation = f"unknown tool: {tool_name}" messages.append({"role": "user", "content": f"Observation: {observation}"}) continue if "Final Answer:" in text: return text.split("Final Answer:")[-1].strip() return "Agent got stuck — no action found." return "Max steps reached without a final answer."
What the fixes changed - eight runs each
First reply contains an Observation the model wrote itself: without stop sequence 5 of 6 with stop sequence 0 of 6 "Should I bring an umbrella in Mumbai, and what's 12 * 8?" both tools ran stuck answered without calculator original loop 4 3 1 with both fixes 5 1 2
A calculator that cannot run arbitrary code
import ast import operator OPERATORS = { ast.Add: operator.add, ast.Sub: operator.sub, ast.Mult: operator.mul, ast.Div: operator.truediv, ast.USub: operator.neg, } def safe_calculate(expression): """Arithmetic only: numbers, + - * /, brackets. Anything else is refused.""" def evaluate(node): if isinstance(node, ast.Constant) and isinstance(node.value, (int, float)): return node.value if isinstance(node, ast.BinOp) and type(node.op) in OPERATORS: return OPERATORS[type(node.op)](evaluate(node.left), evaluate(node.right)) if isinstance(node, ast.UnaryOp) and type(node.op) in OPERATORS: return OPERATORS[type(node.op)](evaluate(node.operand)) raise ValueError("only numbers and + - * / are allowed") try: return str(evaluate(ast.parse(expression, mode="eval").body)) except (SyntaxError, ValueError, ZeroDivisionError) as e: return f"error: {e}" print(safe_calculate("12 * 8")) print(safe_calculate("12 * (3 + 4)")) print(safe_calculate("-5 / 2")) print(safe_calculate('__import__("os").getcwd()')) print(safe_calculate("10 / 0"))
Output
96 84 -2.5 error: only numbers and + - * / are allowed error: division by zero

Watch out: calculate() uses eval(), which runs whatever string it is given as Python - calculate('__import__("os").getcwd()') returns your real working directory. Fine on your own machine with your own questions; dangerous anywhere a stranger can reach it. safe_calculate above accepts arithmetic and nothing else.

Tip: The print(text) inside the loop is not decoration. Reading the Thoughts go by is how you find out why an agent went the wrong way.

The loop at a glance

TOOLS

The tool registry - names the model may use, functions your code runs.

TOOLS = {"get_weather": get_weather, "calculate": calculate}
SYSTEM_PROMPT

Teaches the format and lists the tools.

Thought / Action / Observation / Final Answer
Action regex

Pulls the tool name and argument out of the reply.

re.search(r'Action:\s*(\w+)\((.*?)\)', text)
Observation

The tool result, appended as a user message.

{"role": "user", "content": f"Observation: {observation}"}
stop sequence

Ends generation before the model can invent a result.

options={"stop": ["Observation:"]}
max_steps

Caps the number of passes.

for step in range(max_steps):

Try it yourself

The code does not change. Swap the content string and the program does something else entirely.

The lesson question

“Should I bring an umbrella in Mumbai, and what's 12 * 8?”

Compare two cities

“Is it warmer in Mumbai or Delhi today?”

Needs two calculations

“What is 12 * 8, and what is that plus 100?”

No tool needed

“What is the capital of India?”

What usually goes wrong

Letting the model write Observations

It will invent them. Stop generation at "Observation:" and write every Observation from your code.

✗ ollama.chat(model="llama3.1", messages=messages)
✓ ollama.chat(model="llama3.1", messages=messages,
            options={"stop": ["Observation:"]})
Accepting a Final Answer that skipped the tool

If a reply has both an Action and a Final Answer, the answer was written before any result existed. Run the Action first.

✗ if "Final Answer:" in text:
    return ...
match = re.search(...)
✓ match = re.search(...)
if match:
    ...run the tool...
    continue
if "Final Answer:" in text:
    return ...
Forgetting to append the reply

Without the assistant message, the model does not see its own Thought and Action on the next turn, and the Observation arrives with no context.

✗ text = response["message"]["content"]
# straight on to the regex
✓ messages.append({"role": "assistant", "content": text})
No max_steps

A model that never writes Final Answer keeps the loop running forever.

✗ while True:
✓ for step in range(max_steps):
eval() on model output

The argument comes from the model, which can be steered by whoever wrote the question. eval runs it as Python.

✗ return str(eval(expression))
✓ return safe_calculate(expression)
Calling any name the model writes

Look the name up in TOOLS. An unknown name becomes an Observation the model can learn from, not a crash or a call to something you never registered.

Key points

  • ReAct = Reason + Act: Thought, Action, Observation, repeat, Final Answer.
  • The model decides; your for loop controls everything - including which tools run.
  • Observations are written by your code and sent back as user messages.
  • A stop sequence on "Observation:" stops the model inventing tool results.
  • Check for an Action before accepting a Final Answer.
  • max_steps caps the loop so it cannot run forever.
  • The simple regex breaks on nested brackets and keyword arguments.
  • eval() runs any Python; use a calculator that only accepts arithmetic.

Quick check before you move on

What does ReAct stand for?
Reason and Act.
What is the Thought for?
The model reasoning about what it needs next - it also makes the transcript readable when you debug.
Who executes the Action?
Your code, through the TOOLS dictionary. The model only names it.
What is an Observation?
The real result of a tool, written by your code and sent back to the model.
Why must the model not write its own Observation?
It would be invented. A stop sequence on "Observation:" prevents it.
What does max_steps protect against?
A loop that never ends because the model never writes a Final Answer.
What happens when the model writes Final Answer:?
If there is no Action in the same reply, the loop returns the text after it.
What is the role of the for loop?
It is the agent: it calls the model, runs tools, feeds back results, and decides when to stop.

Quiz

  1. 1.

    What do the four labels Thought, Action, Observation and Final Answer each represent?

  2. 2.

    Why does the loop stop once "Final Answer:" appears?

  3. 3.

    What is max_steps protecting against?

  4. 4.

    What does options={"stop": ["Observation:"]} do?

  5. 5.

    The model writes Action: calculate(12 * (3 + 4)). What does the regex pass to calculate?

Interview questions

What is the ReAct pattern?

Reason and Act: the LLM decides the next action, the application executes the tool, feeds the result back as an observation, and repeats until the model produces a final answer. The control flow stays in the application, not the LLM.

How do you stop a model inventing tool results?

Use a stop sequence so generation ends before the model writes an observation, write every observation from your own code, and never accept a final answer from a reply that also requested a tool.

Why cap an agent loop with a step limit?

A model can fail to converge - repeating actions or never producing a final answer. A step limit bounds cost and latency, and hitting it is a useful signal that the prompt or tools need work.

What is dangerous about an eval-based calculator tool?

The argument comes from model output, which can be influenced by user input, and eval executes arbitrary Python. Parse the expression and allow only arithmetic nodes instead.

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...