← Back to Agentic AI map
Lesson 1.5 · Agents From Scratch

Manual Tool Calling

Teach the model to ask for a tool in a fixed format - then parse that request and run the function yourself.

tools

What you will be able to do

  • Describe a tool to a model in a system message
  • Define an exact request format, and explain why it has to be exact
  • Parse the request with a regular expression and run the matching function
  • Handle the case where no tool is needed
  • Recognise the ways a model breaks the format - including pretending it already used the tool
  • Explain what "manual" means here, and what built-in tool calling automates later

The idea, in plain English

Lesson 1.4 said an agent is an LLM, tools, and a loop. Before the loop, one question needs answering: how does the model tell your program it wants a tool? The model can only write text, so the answer is a text format both sides agree on.

You describe the tools in the system message - their names, their arguments, and what they do - and tell the model exactly what to write when it wants one: a line reading TOOL: get_weather(city="Mumbai"). Your program looks for that line, pulls out the tool name and the argument, and calls the real function.

The format has to be exact, because code has to recognise it. "Could you check the weather in Mumbai?" is obvious to a person and a guessing game for a program. TOOL: get_weather(city="Mumbai") has a keyword, a name, and an argument in fixed places, so one regular expression finds all three.

The model is choosing; your code is doing. That split is the whole point of this lesson - and the outputs below, from real runs, show why your code also has to assume the model will not always keep its side of the deal.

Worked example: A pretend get_weather(city) tool.

workflowFrom a question to a function callstep 1 / 5

1 - You describe the tool

The system message lists get_weather(city) and the exact line to write when it is needed. This is the only way the model knows the tool exists.

tools
1
format
TOOL: name(arg="x")
written by
you
model has seen code
never

Step through one request. The model produces a line of text; everything after that is your code.

Describing a tool

The model has never seen your code. Everything it knows about a tool comes from the system message: its name, its arguments, and one line on what it returns. get_weather(city): returns the current weather for a city is enough for one tool; with several, the descriptions are what the model uses to choose between them.

The same message sets the format and the fallback: write TOOL: ... when a tool is needed, "otherwise, just answer normally". That fallback is what keeps "What is Python?" from turning into a pointless tool call - in our runs it never asked for the weather tool for that question.

Who does what

Keep the split sharp. The model understands the request, decides whether a tool is needed, picks the tool, and later writes the final answer. Your application parses the request, checks it, runs the function, touches the filesystem or the network, and hands results back.

Model and application
Understand the requestModel
Decide whether a tool is neededModel
Choose the tool and its argumentModel
Parse and check the requestYour code
Run the function, touch files or APIsYour code
Return the result to the modelYour code (Lesson 1.6)
Write the final answerModel

What the model actually sends back

We ran the umbrella question and a similar one about Delhi five times each at the default temperature. The umbrella question produced the exact TOOL line four times out of five. The Delhi question managed it only twice.

The failures fall into three groups. Some wrap the request in chat - "Let me see... TOOL: get_weather(city="Delhi")" - which a search-anywhere parser still catches. Some break the format - TOOLS: instead of TOOL:, or get_weather("Mumbai") with no keyword - which the parser rightly ignores. And some are dangerous: "Delhi is experiencing warm weather with a temperature of 32°C" - an answer that reads as if the tool ran, with numbers the model invented.

The last kind is why the model must never be trusted to report a tool result. Only your code knows whether a function was really called.

Watch out: A model can describe a tool result it never received. If your code did not run the tool, there is no result - however confident the text sounds.

Temperature 0 steadies the format

The same three questions at temperature 0 behaved perfectly and identically across five runs each: TOOL lines for Mumbai and Delhi, a plain answer for "What is Python?". The decision to call a tool is a formatting job as much as a reasoning one, and it benefits from the same setting as Lesson 1.3.

Temperature 0 makes the model consistent, not correct. A question it gets wrong at temperature 0 it will get wrong every time - which is why parsing and checking still stay.

Why this is called manual

Here you write every piece yourself: the tool description, the request format, the parser, and the call. Modern model APIs, including Ollama, also offer built-in tool calling: you pass a list of tool schemas, and the reply contains structured tool calls instead of text you parse.

It depends on the model. llama3.1 supports it; llama3 does not - Ollama rejects the request with "does not support tools". Building it by hand first shows you what the built-in version automates, and works with any model. Module 2 uses the built-in route through LangChain.

Step-by-step code

Describe a tool and let the model ask for it
import ollama SYSTEM_PROMPT = """You have access to one tool: - get_weather(city): returns the current weather for a city If you need to use it, respond with EXACTLY this format on its own line: TOOL: get_weather(city="CityName") Otherwise, just answer normally.""" response = ollama.chat( model="llama3.1", messages=[ {"role": "system", "content": SYSTEM_PROMPT}, {"role": "user", "content": "Should I bring an umbrella in Mumbai today?"} ] ) output = response["message"]["content"] print(output) # Expect something like: TOOL: get_weather(city="Mumbai")
What came back - five runs of each question, default temperature
"Should I bring an umbrella in Mumbai today?" 4 of 5: TOOL: get_weather(city="Mumbai") 1 of 5: I'd recommend checking the weather first! According to the current weather forecast, Mumbai ... <- invented "Is it hot in Delhi right now?" - some of the misses TOOL: get_weather(city="Delhi") <- correct I'd recommend checking the current weather in Delhi. Let me see... TOOL: get_weather(city="Delhi") <- wrapped in chat I don't have that information. TOOLS: get_weather(city="Delhi") <- TOOLS, not TOOL According to the current weather, Delhi is experiencing warm weather with a temperature of 32°C ... <- invented, no tool "What is Python?" 5 of 5: Python is a high-level, interpreted programming language ...
Parse the request and run the tool
import re import ollama SYSTEM_PROMPT = """You have access to one tool: - get_weather(city): returns the current weather for a city If you need to use it, respond with EXACTLY this format on its own line: TOOL: get_weather(city="CityName") Otherwise, just answer normally.""" def get_weather(city): fake_data = {"mumbai": "rainy, 27°C", "delhi": "sunny, 34°C"} return fake_data.get(city.lower(), "unknown city") TOOLS = {"get_weather": get_weather} # TOOL: then a name, then city="..." in brackets TOOL_PATTERN = re.compile(r'TOOL:\s*(\w+)\(\s*city\s*=\s*"([^"]*)"\s*\)') def handle(question): response = ollama.chat( model="llama3.1", messages=[ {"role": "system", "content": SYSTEM_PROMPT}, {"role": "user", "content": question}, ], options={"temperature": 0}, ) output = response["message"]["content"].strip() match = TOOL_PATTERN.search(output) if match is None: return f"No tool requested. The model said: {output[:60]}..." tool_name, city = match.groups() if tool_name not in TOOLS: return f"The model asked for a tool that does not exist: {tool_name}" result = TOOLS[tool_name](city) # the model asked; this line runs it return f"Ran {tool_name}({city!r}) -> {result}" for question in [ "Should I bring an umbrella in Mumbai today?", "Is it hot in Delhi right now?", "What is Python?", ]: print(handle(question))
Output - from a real run
Ran get_weather('Mumbai') -> rainy, 27°C Ran get_weather('Delhi') -> sunny, 34°C No tool requested. The model said: Python is a high-level, interpreted programming language tha...
What the parser catches
TOOL: get_weather(city="Mumbai") match (get_weather, Mumbai) ... Let me see... TOOL: get_weather(city="Delhi") match (get_weather, Delhi) I don't have that information. TOOLS: get_weather(...) no match I think you should use the weather tool: get_weather("Mumbai") no match Delhi is experiencing warm weather ... 32°C no match - and no tool ran

Tip: Search for the pattern anywhere in the reply rather than demanding it on a line of its own. It rescues requests the model wrapped in a sentence, without accepting anything that is not the real format.

Watch out: Check the tool name against your own dictionary before calling anything. The model can name a tool that does not exist - and in later lessons, arguments you never meant to allow.

Manual tool calling, piece by piece

Tool description

Tells the model the tool exists and what it does.

- get_weather(city): returns the weather
Request format

The exact text the model writes to ask.

TOOL: get_weather(city="Mumbai")
Fallback

What to do when no tool is needed.

Otherwise, just answer normally.
Parser

Finds the tool name and argument in the reply.

re.search(TOOL_PATTERN, output)
Tool registry

The functions your code is willing to run.

TOOLS = {"get_weather": get_weather}
Execution

Your code calls the function - never the model.

TOOLS[name](city)

Try it yourself

The code does not change. Swap the content string and the program does something else entirely.

Needs the tool

“Is it hot in Delhi right now?”

Does not need it

“What is Python?”

Unknown city

“What is the weather like in Reykjavik?”

Two cities

“Is it warmer in Mumbai or Delhi today?”

What usually goes wrong

A format your code cannot parse

Free-form requests are easy for people and unreliable for code. Give the model one exact line to write.

✗ If you need the weather, just ask for it.
✓ Respond with EXACTLY: TOOL: get_weather(city="CityName")
Believing the model ran the tool

A reply that reads like a weather report is not a tool result. Only your code knows whether a function was called.

✗ print(output)  # "Delhi is 32°C" - invented
✓ result = TOOLS[name](city)  # the only real weather
Calling whatever name the model wrote

Look the name up in your own registry. Anything else is a request for a tool you do not have.

✗ globals()[tool_name](city)
✓ if tool_name in TOOLS:
    result = TOOLS[tool_name](city)
Forgetting the fallback

Without "otherwise, just answer normally", every question looks like a reason to call a tool.

Leaving temperature at the default

At the default temperature the Delhi question followed the format only two times in five. At temperature 0 it followed it every time.

✗ ollama.chat(model=..., messages=...)
✓ ollama.chat(model=..., messages=..., options={"temperature": 0})

Key points

  • The model can only write text, so tool calls start as a text format both sides agree on.
  • Describe each tool in the system message; the model knows nothing else about it.
  • An exact format - TOOL: name(arg="x") - lets one regular expression find the name and argument.
  • The model asks; your code parses, checks the name, and runs the function.
  • No match means no tool call - treat the text as an ordinary answer.
  • Models break the format, wrap it in chat, and sometimes invent a tool result outright.
  • Temperature 0 made the format reliable in our runs; checking the request is still your job.
  • "Manual" means you wrote every piece - built-in tool calling automates it for models that support it.

Quick check before you move on

Who actually executes the tool?
Your Python code. The model only asks for it.
Why use an exact format like TOOL: get_weather(city="Mumbai")?
So code can reliably detect the request and extract the tool name and argument.
What happens when the question does not need a tool?
The model answers normally, the parser finds no TOOL line, and nothing runs.
The reply says "It is 32°C in Delhi", but no TOOL line was found. What happened?
The model invented the weather. No tool ran, so that number came from nowhere.

Quiz

  1. 1.

    At this stage, who actually executes the tool - the model or your Python code?

  2. 2.

    Why do we ask for an exact, fixed format like TOOL: get_weather(city="Mumbai")?

  3. 3.

    What would happen if the user's question did not need a tool at all?

  4. 4.

    Why check the tool name against a dictionary before calling it?

  5. 5.

    What does built-in tool calling automate, compared with this lesson?

Interview questions

How does an LLM call a tool?

It does not call anything. It produces a request - as formatted text, or as a structured tool call in APIs that support one - and the application parses it, validates it, runs the function, and returns the result.

What can go wrong with text-based tool calling, and how do you guard against it?

The model may break the format, wrap it in extra text, name a tool that does not exist, or describe a result it never received. Use an exact format, parse with a pattern, check the name against a registry, lower the temperature, and never trust a result your code did not produce.

Why learn manual tool calling when APIs have it built in?

It shows exactly what built-in tool calling automates, it works with models that do not support it, and it makes the boundary between the model deciding and the application executing impossible to miss.

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...