Manual Tool Calling
Teach the model to ask for a tool in a fixed format - then parse that request and run the function yourself.
What you will be able to do
- Describe a tool to a model in a system message
- Define an exact request format, and explain why it has to be exact
- Parse the request with a regular expression and run the matching function
- Handle the case where no tool is needed
- Recognise the ways a model breaks the format - including pretending it already used the tool
- Explain what "manual" means here, and what built-in tool calling automates later
The idea, in plain English
Lesson 1.4 said an agent is an LLM, tools, and a loop. Before the loop, one question needs answering: how does the model tell your program it wants a tool? The model can only write text, so the answer is a text format both sides agree on.
You describe the tools in the system message - their names, their arguments, and what they do - and tell the model exactly what to write when it wants one: a line reading TOOL: get_weather(city="Mumbai"). Your program looks for that line, pulls out the tool name and the argument, and calls the real function.
The format has to be exact, because code has to recognise it. "Could you check the weather in Mumbai?" is obvious to a person and a guessing game for a program. TOOL: get_weather(city="Mumbai") has a keyword, a name, and an argument in fixed places, so one regular expression finds all three.
The model is choosing; your code is doing. That split is the whole point of this lesson - and the outputs below, from real runs, show why your code also has to assume the model will not always keep its side of the deal.
Worked example: A pretend get_weather(city) tool.
1 - You describe the tool
The system message lists get_weather(city) and the exact line to write when it is needed. This is the only way the model knows the tool exists.
Step through one request. The model produces a line of text; everything after that is your code.
Describing a tool
The model has never seen your code. Everything it knows about a tool comes from the system message: its name, its arguments, and one line on what it returns. get_weather(city): returns the current weather for a city is enough for one tool; with several, the descriptions are what the model uses to choose between them.
The same message sets the format and the fallback: write TOOL: ... when a tool is needed, "otherwise, just answer normally". That fallback is what keeps "What is Python?" from turning into a pointless tool call - in our runs it never asked for the weather tool for that question.
Who does what
Keep the split sharp. The model understands the request, decides whether a tool is needed, picks the tool, and later writes the final answer. Your application parses the request, checks it, runs the function, touches the filesystem or the network, and hands results back.
Understand the requestModelDecide whether a tool is neededModelChoose the tool and its argumentModelParse and check the requestYour codeRun the function, touch files or APIsYour codeReturn the result to the modelYour code (Lesson 1.6)Write the final answerModelWhat the model actually sends back
We ran the umbrella question and a similar one about Delhi five times each at the default temperature. The umbrella question produced the exact TOOL line four times out of five. The Delhi question managed it only twice.
The failures fall into three groups. Some wrap the request in chat - "Let me see... TOOL: get_weather(city="Delhi")" - which a search-anywhere parser still catches. Some break the format - TOOLS: instead of TOOL:, or get_weather("Mumbai") with no keyword - which the parser rightly ignores. And some are dangerous: "Delhi is experiencing warm weather with a temperature of 32°C" - an answer that reads as if the tool ran, with numbers the model invented.
The last kind is why the model must never be trusted to report a tool result. Only your code knows whether a function was really called.
Watch out: A model can describe a tool result it never received. If your code did not run the tool, there is no result - however confident the text sounds.
Temperature 0 steadies the format
The same three questions at temperature 0 behaved perfectly and identically across five runs each: TOOL lines for Mumbai and Delhi, a plain answer for "What is Python?". The decision to call a tool is a formatting job as much as a reasoning one, and it benefits from the same setting as Lesson 1.3.
Temperature 0 makes the model consistent, not correct. A question it gets wrong at temperature 0 it will get wrong every time - which is why parsing and checking still stay.
Why this is called manual
Here you write every piece yourself: the tool description, the request format, the parser, and the call. Modern model APIs, including Ollama, also offer built-in tool calling: you pass a list of tool schemas, and the reply contains structured tool calls instead of text you parse.
It depends on the model. llama3.1 supports it; llama3 does not - Ollama rejects the request with "does not support tools". Building it by hand first shows you what the built-in version automates, and works with any model. Module 2 uses the built-in route through LangChain.
Step-by-step code
import ollama
SYSTEM_PROMPT = """You have access to one tool:
- get_weather(city): returns the current weather for a city
If you need to use it, respond with EXACTLY this format on its own line:
TOOL: get_weather(city="CityName")
Otherwise, just answer normally."""
response = ollama.chat(
model="llama3.1",
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": "Should I bring an umbrella in Mumbai today?"}
]
)
output = response["message"]["content"]
print(output)
# Expect something like: TOOL: get_weather(city="Mumbai")"Should I bring an umbrella in Mumbai today?"
4 of 5: TOOL: get_weather(city="Mumbai")
1 of 5: I'd recommend checking the weather first! According to the
current weather forecast, Mumbai ... <- invented
"Is it hot in Delhi right now?" - some of the misses
TOOL: get_weather(city="Delhi") <- correct
I'd recommend checking the current weather in Delhi. Let me see...
TOOL: get_weather(city="Delhi") <- wrapped in chat
I don't have that information. TOOLS: get_weather(city="Delhi") <- TOOLS, not TOOL
According to the current weather, Delhi is experiencing warm
weather with a temperature of 32°C ... <- invented, no tool
"What is Python?"
5 of 5: Python is a high-level, interpreted programming language ...import re
import ollama
SYSTEM_PROMPT = """You have access to one tool:
- get_weather(city): returns the current weather for a city
If you need to use it, respond with EXACTLY this format on its own line:
TOOL: get_weather(city="CityName")
Otherwise, just answer normally."""
def get_weather(city):
fake_data = {"mumbai": "rainy, 27°C", "delhi": "sunny, 34°C"}
return fake_data.get(city.lower(), "unknown city")
TOOLS = {"get_weather": get_weather}
# TOOL: then a name, then city="..." in brackets
TOOL_PATTERN = re.compile(r'TOOL:\s*(\w+)\(\s*city\s*=\s*"([^"]*)"\s*\)')
def handle(question):
response = ollama.chat(
model="llama3.1",
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": question},
],
options={"temperature": 0},
)
output = response["message"]["content"].strip()
match = TOOL_PATTERN.search(output)
if match is None:
return f"No tool requested. The model said: {output[:60]}..."
tool_name, city = match.groups()
if tool_name not in TOOLS:
return f"The model asked for a tool that does not exist: {tool_name}"
result = TOOLS[tool_name](city) # the model asked; this line runs it
return f"Ran {tool_name}({city!r}) -> {result}"
for question in [
"Should I bring an umbrella in Mumbai today?",
"Is it hot in Delhi right now?",
"What is Python?",
]:
print(handle(question))Ran get_weather('Mumbai') -> rainy, 27°C
Ran get_weather('Delhi') -> sunny, 34°C
No tool requested. The model said: Python is a high-level, interpreted programming language tha...TOOL: get_weather(city="Mumbai") match (get_weather, Mumbai)
... Let me see... TOOL: get_weather(city="Delhi") match (get_weather, Delhi)
I don't have that information. TOOLS: get_weather(...) no match
I think you should use the weather tool: get_weather("Mumbai") no match
Delhi is experiencing warm weather ... 32°C no match - and no tool ranTip: Search for the pattern anywhere in the reply rather than demanding it on a line of its own. It rescues requests the model wrapped in a sentence, without accepting anything that is not the real format.
Watch out: Check the tool name against your own dictionary before calling anything. The model can name a tool that does not exist - and in later lessons, arguments you never meant to allow.
Manual tool calling, piece by piece
Tool descriptionTells the model the tool exists and what it does.
- get_weather(city): returns the weather
Request formatThe exact text the model writes to ask.
TOOL: get_weather(city="Mumbai")
FallbackWhat to do when no tool is needed.
Otherwise, just answer normally.
ParserFinds the tool name and argument in the reply.
re.search(TOOL_PATTERN, output)
Tool registryThe functions your code is willing to run.
TOOLS = {"get_weather": get_weather}ExecutionYour code calls the function - never the model.
TOOLS[name](city)
Try it yourself
The code does not change. Swap the content string and the program does something else entirely.
“Is it hot in Delhi right now?”
“What is Python?”
“What is the weather like in Reykjavik?”
“Is it warmer in Mumbai or Delhi today?”
What usually goes wrong
Free-form requests are easy for people and unreliable for code. Give the model one exact line to write.
✗ If you need the weather, just ask for it.✓ Respond with EXACTLY: TOOL: get_weather(city="CityName")A reply that reads like a weather report is not a tool result. Only your code knows whether a function was called.
✗ print(output) # "Delhi is 32°C" - invented✓ result = TOOLS[name](city) # the only real weatherLook the name up in your own registry. Anything else is a request for a tool you do not have.
✗ globals()[tool_name](city)✓ if tool_name in TOOLS:
result = TOOLS[tool_name](city)Without "otherwise, just answer normally", every question looks like a reason to call a tool.
At the default temperature the Delhi question followed the format only two times in five. At temperature 0 it followed it every time.
✗ ollama.chat(model=..., messages=...)✓ ollama.chat(model=..., messages=..., options={"temperature": 0})Key points
- The model can only write text, so tool calls start as a text format both sides agree on.
- Describe each tool in the system message; the model knows nothing else about it.
- An exact format - TOOL: name(arg="x") - lets one regular expression find the name and argument.
- The model asks; your code parses, checks the name, and runs the function.
- No match means no tool call - treat the text as an ordinary answer.
- Models break the format, wrap it in chat, and sometimes invent a tool result outright.
- Temperature 0 made the format reliable in our runs; checking the request is still your job.
- "Manual" means you wrote every piece - built-in tool calling automates it for models that support it.
Quick check before you move on
Quiz
- 1.
At this stage, who actually executes the tool - the model or your Python code?
- 2.
Why do we ask for an exact, fixed format like TOOL: get_weather(city="Mumbai")?
- 3.
What would happen if the user's question did not need a tool at all?
- 4.
Why check the tool name against a dictionary before calling it?
- 5.
What does built-in tool calling automate, compared with this lesson?
Interview questions
How does an LLM call a tool?
It does not call anything. It produces a request - as formatted text, or as a structured tool call in APIs that support one - and the application parses it, validates it, runs the function, and returns the result.
What can go wrong with text-based tool calling, and how do you guard against it?
The model may break the format, wrap it in extra text, name a tool that does not exist, or describe a result it never received. Use an exact format, parse with a pattern, check the name against a registry, lower the temperature, and never trust a result your code did not produce.
Why learn manual tool calling when APIs have it built in?
It shows exactly what built-in tool calling automates, it works with models that do not support it, and it makes the boundary between the model deciding and the application executing impossible to miss.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...
AI
System Design
Backend
- GraphQL8 modules · 69 lessons planned
- Core Python13 modules · 75 lessons planned
- FastAPI5 sections · 20 lessons
- Node.js14 modules · 206 lessons planned
- Node.js Performance7 chapters · 36 topics
- Event Loop Lifecycle6 phases · 3 scenarios
- Docker & Containerization11 modules · 144 lessons planned
- AWS for Developers14 modules · 219 lessons planned
- CI/CD & DevOps Automation10 modules · 134 lessons planned