← Back to Agentic AI map
Lesson 1.10 · Agents From Scratch

Error Handling & Retries

Catch bad output, unknown tools, and tools that blow up - and keep the agent running.

robustness

What you will be able to do

  • Name the three failures a hand-built agent meets, and who caused each one
  • Validate model output, and retry it with corrective feedback
  • Turn an unknown tool and a failing tool into Observations the model can act on
  • Explain why errors should neither crash the agent nor disappear silently
  • Count what retries cost, and bound them
  • Give the model a way out when no tool can help

The idea, in plain English

Everything that can go wrong in Lesson 1.6 eventually does. The model writes something that does not match the format. It asks for a tool that does not exist. It passes an argument the tool chokes on.

A sturdy agent expects all three and keeps going, and the interesting part is that each one gets a different fix - because each one was made by someone different. Unreadable output is the model’s mistake, so you tell it so and ask again. An unknown tool is the model reaching for something you never gave it, so you reply with the list of tools that exist - an Observation is the natural place, because the model reads Observations. A tool that raises is your code failing, so you catch it and hand back the error text instead of letting the exception end the run.

None of these crash the program, and none of them silently swallow the problem. The agent always finds out what went wrong and gets a chance to fix it. Two limits keep that from turning into a loop that never ends: max_steps for the whole run, and max_retries_per_step for each reply.

We ran the course code against a local model to see what each mechanism actually buys - and found one failure the three fixes do not cover.

Worked example: Add retry logic to the ReAct loop from Lesson 1.6.

workflowThree failures, three fixesstep 1 / 4

Failure 1 - the reply cannot be read

No Action and no Final Answer - "Action: (no action needed, just a thought)", say. The original loop gave up here. Now it appends a nudge and asks again, up to max_retries_per_step times.

made by
the model
fix
nudge + retry
costs
one more call
limit
max_retries_per_step

One agent step, and every way it can go wrong. Each failure ends up as text the model reads on its next turn - never as a crash.

Two loops

The outer loop - for step in range(max_steps) - is the agent: Thought, Action, Observation, repeat. The inner loop - for retry in range(max_retries_per_step) - only decides whether this one reply is usable. A reply is usable if it contains a Final Answer or an Action the regex can read.

Despite the name, max_retries_per_step=2 means two attempts in total: the first try and one retry. If both are unreadable the function returns "Agent failed to produce a valid step after retries" rather than trying forever.

Failure 1: output that cannot be read

In Lesson 1.6 an unreadable reply ended the run: "Agent got stuck". Now the loop appends "That wasn’t in the right format. Use Thought/Action or Final Answer." and asks again. The conversation has changed, so the next reply can too - unlike retrying the identical request at temperature 0, which Lesson 1.3 showed repeats the same failure.

It works. On "Should I bring an umbrella in Mumbai, and what’s 12 * 8?", eight runs each: the original loop got stuck twice and used both tools in six; run_agent_safe never got stuck or failed, and used both tools in seven. The price was about a quarter more model calls - 3.6 per question on average instead of 2.9.

Failure 2: a tool that does not exist

The model can only ask for tools it has heard of, so this usually means a mismatch - a tool in the prompt that never made it into TOOLS, the exact mistake Lesson 1.7 warned about. We listed a search(query) tool in the prompt but did not register it, and asked for the population of Hyderabad.

The error Observation did its job: the model read "search is not a valid tool" and the list of real tools. What it did next is the lesson. It tried read_file on a Google URL, then calculate on the words "population of Hyderabad", then get_weather on Hyderabad - every tool it had, none of which could help - until max_steps ran out. Three runs out of three.

Telling a model which tools exist also tells it, implicitly, that it should use one. It needs explicit permission to stop.

Failure 3: a tool that raises

try/except around the call turns any exception into text: "error running read_file: [Errno 21] Is a directory: ’.’". The agent keeps running, and the model knows that path did not work.

The opposite mistake is except: pass - no crash, but no one knows anything went wrong, least of all the model, which carries on as if the tool had returned nothing. An error the model never sees is an error it cannot recover from.

Watch out: Error text goes straight into the conversation. Keep it to what the model needs - a short reason - rather than full tracebacks with paths and internal details.

The failure none of the three cover

The course prompt allows exactly two kinds of reply: a Thought and Action, or a Final Answer after using tools. A question that needs no tool - "Hi! How are you today?" - fits neither comfortably, and in four of eight runs the agent failed even with the nudge. A question no tool can answer produced the tool-cycling above.

One line in the system prompt covers both: if the question needs no tool, or none of the tools can help, answer straight away with a Final Answer that says so. With it, greetings were answered four times out of four, and the Hyderabad question four out of four - three saying plainly that there is no search tool. The umbrella question still used both tools in three of four runs, as it had before.

The fourth Hyderabad answer is the reminder that robustness has limits: "According to Google, the population of Hyderabad is approximately 7.7 million" - a search result from a tool that does not exist. Error handling keeps the agent running. It cannot make the model honest.

Retries are not free

Every retry is another model call. With max_steps=5 and max_retries_per_step=2, one question can take up to ten calls - fine on a local model, worth counting anywhere you pay per call or users wait for a reply.

Keep both limits small and log when they are hit. An agent that regularly needs retries is telling you the prompt, the tools, or the model is the real problem.

Step-by-step code

The loop from 1.6, made robust
# Reuses the tools, TOOLS and SYSTEM_PROMPT from Lesson 1.7. def run_agent_safe(user_question, max_steps=5, max_retries_per_step=2): messages = [ {"role": "system", "content": SYSTEM_PROMPT}, {"role": "user", "content": user_question} ] for step in range(max_steps): text = None for retry in range(max_retries_per_step): response = ollama.chat(model="llama3.1", messages=messages) candidate = response["message"]["content"] if "Final Answer:" in candidate or re.search(r'Action:\s*(\w+)\((.*?)\)', candidate): text = candidate break # bad/unparsable output — nudge and retry messages.append({"role": "user", "content": "That wasn't in the right format. Use Thought/Action or Final Answer."}) if text is None: return "Agent failed to produce a valid step after retries." messages.append({"role": "assistant", "content": text}) if "Final Answer:" in text: return text.split("Final Answer:")[-1].strip() match = re.search(r'Action:\s*(\w+)\((.*?)\)', text) tool_name, arg = match.group(1), match.group(2).strip('"\' ') if tool_name not in TOOLS: observation = f"error: '{tool_name}' is not a valid tool. Available tools: {list(TOOLS.keys())}" else: try: observation = TOOLS[tool_name](arg) except Exception as e: observation = f"error running {tool_name}: {e}" messages.append({"role": "user", "content": f"Observation: {observation}"}) return "Max steps reached without a final answer."
What it bought - from real runs
"Should I bring an umbrella in Mumbai, and what's 12 * 8?" - eight runs each both tools used stuck or failed model calls (avg) run_agent (1.6) 6 2 2.9 run_agent_safe 7 0 3.6 A tool that raises - read_file on a folder: Observation: error running read_file: [Errno 21] Is a directory: '.' (in Lesson 1.7 the same call crashed the whole agent)
An unknown tool - and what the model did next
Prompt lists search(query); TOOLS does not have it. "Search for the population of Hyderabad and tell me what you find." Action: search("population of Hyderabad") Observation: error: 'search' is not a valid tool. Available tools: ['get_weather', 'calculate', 'read_file'] "...I can use the read_file tool to find the population of Hyderabad..." "...the calculate tool is not able to process the string 'population of Hyderabad'..." "...the get_weather tool is not able to find the city 'Hyderabad'..." -> Max steps reached without a final answer. (3 of 3 runs)
Give the model a way out
SYSTEM_PROMPT += ( "\n\nIf the question needs no tool, or none of these tools can help, " "answer straight away with a Final Answer that says so." )
With and without that line - four runs each
without with "Hi! How are you today?" 4 of 8 ok 4 of 4 ok "Search for the population ..." 0 of 4 ok 4 of 4 ok "Umbrella ... and 12 * 8?" 3 of 4 both 3 of 4 both tools With the line, the Hyderabad question: "Unfortunately, the available tools do not include a search function." "Unfortunately, I'm unable to search for the population of Hyderabad ..." "Unfortunately, I don't have a tool to search for information about ..." "According to Google, the population of Hyderabad is approximately 7.7 million" <- invented

Tip: Feeding the error back as an Observation works because models are good at correcting themselves when told plainly what was wrong and what the valid options are - as long as "none of them" is also a valid option.

Watch out: Retries cost a call each. With max_steps at 5 and max_retries_per_step at 2, one question can mean ten calls to the model - fine locally, worth counting anywhere you pay per call.

Three failures, three fixes

Unreadable output

Model mistake: nudge and retry, within max_retries_per_step.

"That wasn't in the right format. Use Thought/Action or Final Answer."
Unknown tool

Model reached for something absent: list the real tools.

f"error: '{tool_name}' is not a valid tool. Available tools: {list(TOOLS.keys())}"
Tool raises

Your code failed: catch it and report it.

except Exception as e:
    observation = f"error running {tool_name}: {e}"
No tool fits

Let the model say so in a Final Answer.

answer straight away with a Final Answer that says so
Limits

Bound the whole run and every reply.

max_steps=5, max_retries_per_step=2

Try it yourself

The code does not change. Swap the content string and the program does something else entirely.

Needs both tools

“Should I bring an umbrella in Mumbai, and what's 12 * 8?”

Needs no tool

“Hi! How are you today?”

A tool that raises

“Use read_file on the path . and tell me what it says.”

A tool you do not have

“Search for the population of Hyderabad and tell me what you find.”

What usually goes wrong

Retrying without saying why

The identical request can produce the identical failure. Add the nudge so the next attempt has something new to go on.

Swallowing the error

The agent keeps running, and nobody - including the model - knows anything went wrong.

✗ try:
    observation = TOOLS[tool_name](arg)
except:
    pass
✓ try:
    observation = TOOLS[tool_name](arg)
except Exception as e:
    observation = f"error running {tool_name}: {e}"
Crashing on an unknown tool

TOOLS[tool_name] raises KeyError for a name that is not there. Check first and answer with the valid names.

✗ observation = TOOLS[tool_name](arg)
✓ if tool_name not in TOOLS:
    observation = f"error: ... Available tools: {list(TOOLS)}"
No limit on retries

A model that never produces a readable reply would retry forever. Bound it, and return a clear failure.

✗ while not valid(candidate):
✓ for retry in range(max_retries_per_step):
No way to say "I can’t"

Told only that a tool is invalid, the model tries every other tool instead of giving up - until max_steps runs out.

✗ Available tools:
- get_weather(city)
- calculate(expression)
- read_file(path)
✓ If none of these tools can help, answer straight away with a Final Answer that says so.
Believing error handling makes answers true

It keeps the agent running. One of our runs still reported a Google search result from a search tool that did not exist.

Key points

  • Three failures: unreadable output, an unknown tool, a tool that raises.
  • Fix each where it came from: nudge the model, list the real tools, catch the exception.
  • Every error becomes an Observation - no crash, and nothing silently swallowed.
  • The nudge stopped every stuck run in our tests, for about a quarter more model calls.
  • max_retries_per_step=2 means two attempts; worst case is max_steps × max_retries_per_step calls.
  • Give the model permission to answer without a tool - or to say no tool can help.
  • Error handling keeps the agent running; it cannot make the model honest.

Quick check before you move on

What three failures does this lesson handle?
Output the parser cannot read, a request for a tool that does not exist, and a tool that raises an exception.
What should happen when the model asks for an unknown tool?
No crash - return an error Observation that names the valid tools.
Why wrap tool execution in try/except?
So one failing tool does not end the run. The error goes back to the model as an Observation.
Why is except: pass a bad fix?
Nothing crashes, but the model never learns the tool failed, so it cannot correct course.

Quiz

  1. 1.

    Name two different kinds of failure this code guards against.

  2. 2.

    What does the code do when the model calls a tool that is not in TOOLS?

  3. 3.

    Why wrap the actual tool call in try/except?

  4. 4.

    With max_steps=5 and max_retries_per_step=2, what is the most model calls one question can take?

  5. 5.

    The model is told "search is not a valid tool" and then tries every other tool on a question none of them can answer. What is missing?

Interview questions

How would you make an LLM agent more robust?

Validate the model’s output before acting on it, retry malformed responses with corrective feedback, check requested tool names against an allowlist, wrap tool execution in exception handling, and return recoverable failures to the model as observations. Put limits on both steps and retries so the agent cannot run indefinitely.

Why return errors to the model instead of raising them?

An agent can often recover from a failed step - try another path, fix an argument, pick a different tool - but only if it learns the step failed. Raising ends the run; swallowing hides the problem. An error observation gives the model the information and the chance.

What are the costs of retries?

Every retry is another model call, so latency and cost rise - up to max_steps times max_retries_per_step calls per question. Frequent retries usually point at a prompt, tool, or model problem worth fixing directly.

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...