Custom Tools & Multi-Tool Agents
Give one agent a calculator, a search tool and RAG as a tool - and let the model choose, per question, which ones to use.
What you will be able to do
- Wrap a retriever as a tool, so RAG becomes one capability among several
- Build an agent with a calculator, a search tool and a document tool
- Follow how one question can use one tool, several, or none
- Write tool names and descriptions a model can route between
- Handle unknown tools, bad arguments and tool exceptions
- Keep control of what tools may do, and who they act for
The idea, in plain English
Lesson 2.12’s RAG chain retrieves for every question - even "What is 125 * 48?". An agent turns it around: retrieval becomes a tool, search_documents, next to calculate and search, and the model decides per question which to call. Calculation goes to the calculator, policy questions to the documents, current events to search - and a question that needs two gets two.
The model chooses from what it can see: each tool’s name, description and argument schema (Lesson 2.6). Your application runs the chosen tool and returns the result (Lesson 2.7). Nothing about the loop is new; what is new is that the choice between capabilities now matters, and so do the descriptions that drive it.
Our llama3 cannot call tools, so the agent runs below use a scripted model for the choices, as in Lesson 2.7 - with the real tools, the real create_agent loop and real retrieval. To test the choosing itself, we asked llama3 to pick a tool from the descriptions with structured output: 8 of 8 right with good descriptions, 6 of 8 with vague ones.
Worked example: Multi-tool assistant.
1 - Three tools on offer
The model receives the question and three schemas: calculate(expression), search(query), search_documents(question), each with its description.
"According to our documents, how many annual leave days do employees get, and what is 20 * 12?" - the run as create_agent executed it.
RAG chain vs RAG tool
A RAG chain is a fixed path: every question is retrieved for, then answered. A RAG tool is optional: the agent retrieves only when the question needs the documents. "What is 20 * 12?" goes straight to the calculator; "how many leave days?" goes to the documents.
The cost is predictability. A chain always does the same thing and is easy to test. An agent can choose wrongly, skip a tool it needed, or call one it did not - which is why the rest of this lesson is about making the choice easy and the consequences safe.
FlowChain: fixed. Agent: chosen per question.RetrievalChain: always. Agent: when the model calls the RAG tool.TestingChain: one path. Agent: every route the model might take.Good forChain: known workflows. Agent: open-ended questions across capabilities.Three tools
calculate is Lesson 1.6’s safe evaluator as a @tool - numbers and + - * / only. Not eval(): with eval the model’s argument is Python on your machine. search is a placeholder that returns "Search results for: ..."; in a real system it calls a search API. search_documents wraps the retriever: call retriever.invoke(question), join the chunks, return the text.
LangChain has a helper for the last one: create_retriever_tool(retriever, "search_documents", description) from langchain_core.tools returns the same kind of tool, with an argument named query. Writing it yourself lets you choose the argument name and how chunks are formatted - for example with page numbers, as in Lesson 2.12.
The model routes by description
The model never sees your code. It sees name, description and parameters - the JSON schema from Lesson 2.6 - for every tool, on every call. "Search." tells it nothing about which search; "Search the company’s internal documents. Use this for questions about company policies, procedures or internal documentation." tells it when.
We tested this with llama3. It cannot call tools, but it can choose one with structured output, so we gave it the tool names and descriptions and eight questions. With our descriptions it chose correctly 8 of 8 times. With every description cut to "Calculate." or "Search.", 6 of 8: "How many days a week can I work from home?" and "Can I claim my taxi fare from last month’s client visit?" both went to "none" - policy questions that never say "documents", so nothing told the model the documents were the place to look.
What is 125 * 48?good: calculate. vague: calculate.According to our documents, how many annual leave days...?good: search_documents. vague: search_documents.How many days a week can I work from home?good: search_documents. vague: none.Can I claim my taxi fare from last month’s client visit?good: search_documents. vague: none.What’s the latest news about Ollama?good: search. vague: search.If I take 3 days of leave every month, how many in a year?good: calculate. vague: calculate.One question, one tool - or several
Each question took its own route through the same agent. "What is 125 * 48?" - calculate, 6000. "According to our documents, how many annual leave days...?" - search_documents, which returned the leave chunk and, because k=2, the travel-expenses chunk too. "Search for information about Python 3.13." - search.
A question that needs two capabilities can get them in sequence - search_documents, then calculate, then the answer - or in a single reply with two tool calls, which the tools node runs in the same step. Either way the order is the model’s decision. A fixed chain would have to retrieve for "What is 20 * 12?" too.
When tools go wrong
Three failures, three behaviours. A tool that does not exist - the model asked for delete_file - is answered for you: "Error: delete_file is not a valid tool, try one of [calculate, search, search_documents]." Wrong arguments are too: calculate called with {"expr": ...} came back as "Error invoking tool ‘calculate’ ... expression: Field required. Please fix the error and try again." The model sees the error and can correct itself.
An exception inside a tool is different: a search tool that raised TimeoutError ended the whole run (as in Lesson 2.1). Either catch it inside the tool and return the error as text (Lesson 1.10), or add ToolErrorMiddleware(on_error) - your function decides which exceptions become a message for the model ("search timed out - try again later or answer without it.") and which still stop the run. It is opt-in on purpose: unhandled internal errors are never shown to the model unless you choose.
The model decides; your application controls
The tool list is the agent’s whole reach. It cannot delete files without a delete tool, and it cannot run code through a calculator that only does arithmetic - the safe calculate answered __import__(‘os’).getcwd() with "Calculation error: only numbers and + - * / are allowed."
Inside each tool, validate as if the arguments came from a stranger - they did. Allowlist paths and operations, never pass model text to SQL or a shell, add timeouts, log every call. And keep identity out of the model’s hands: on a multi-tenant platform, search_documents should filter by the logged-in user’s tenant_id from your session, not by a tenant_id argument the model fills in.
Watch out: Never make permissions a tool argument. A tenant_id, user_id or role that the model supplies is a value the model - or a cleverly worded document - can change.
How many tools?
Every tool’s schema is sent on every model call, and every extra tool is another option to choose wrongly between. Start with the few the task needs. Give each one job and a name that says it - search_documents, not do_search - and descriptions that say when to use it, not only what it does.
Hard-coded routing (if "document" in question: ...) works for two tools and breaks for twenty, and our vague-description test shows why: real questions rarely contain the keyword. An agent routes by meaning - as well as its descriptions allow.
Multi-tool agent vs multi-agent system
This lesson is one agent with several tools. A multi-agent system has several agents - a researcher, a writer, a reviewer - each with its own prompt and tools, coordinated by code or by a supervisor agent. Lesson 1.11 had two agents talking; LangGraph (Module 3) is where those systems are built properly.
Step-by-step code
import ast
import operator
from langchain_core.tools import tool
OPERATORS = {
ast.Add: operator.add,
ast.Sub: operator.sub,
ast.Mult: operator.mul,
ast.Div: operator.truediv,
ast.USub: operator.neg,
}
def evaluate(node):
if isinstance(node, ast.Constant) and isinstance(node.value, (int, float)):
return node.value
if isinstance(node, ast.BinOp) and type(node.op) in OPERATORS:
return OPERATORS[type(node.op)](evaluate(node.left), evaluate(node.right))
if isinstance(node, ast.UnaryOp) and type(node.op) in OPERATORS:
return OPERATORS[type(node.op)](evaluate(node.operand))
raise ValueError("only numbers and + - * / are allowed")
@tool
def calculate(expression: str) -> str:
"""Calculate an arithmetic expression using numbers, + - * / and brackets, e.g. '125 * 48'."""
try:
return str(evaluate(ast.parse(expression, mode="eval").body))
except (SyntaxError, ValueError, ZeroDivisionError) as error:
return f"Calculation error: {error}"
@tool
def search(query: str) -> str:
"""Search the web for current or external information, such as news or software releases."""
return f"Search results for: {query}" # placeholder - call a search API here
@tool
def search_documents(question: str) -> str:
"""Search the company's internal documents. Use this for questions about company policies, procedures or internal documentation."""
documents = retriever.invoke(question) # the retriever from Lessons 2.11-2.12
return "\n\n".join(document.page_content for document in documents)from langchain.agents import create_agent
from langchain_ollama import ChatOllama
agent = create_agent(
ChatOllama(model="llama3.1", temperature=0),
tools=[calculate, search, search_documents],
)
question = "According to our documents, how many annual leave days do employees get, and what is 20 * 12?"
result = agent.invoke(
{"messages": [{"role": "user", "content": question}]},
{"recursion_limit": 10},
)
for message in result["messages"]:
print(message.type, message.content or message.tool_calls)"What is 125 * 48?"
ai calculate {'expression': '125 * 48'}
tool 6000
"According to our documents, how many annual leave days do employees receive?"
ai search_documents {'question': 'annual leave days'}
tool Employees receive 20 days of annual leave per year.
Travel expenses must be submitted within 30 days.
"Search for information about Python 3.13."
ai search {'query': 'Python 3.13'}
tool Search results for: Python 3.13
"... how many annual leave days do employees get, and what is 20 * 12?"
ai search_documents {'question': 'annual leave days'}
tool Employees receive 20 days of annual leave per year. ...
ai calculate {'expression': '20 * 12'}
tool 240
ai Employees get 20 days of annual leave, and 20 * 12 = 240.
The same question, both calls in one reply:
ai search_documents {...} + calculate {...}
tool Employees receive 20 days ...
tool 240Unknown tool
ai delete_file {'path': '/var/log'}
tool Error: delete_file is not a valid tool, try one of [calculate, search, search_documents].
Wrong argument name
ai calculate {'expr': '2 * 3'}
tool Error invoking tool 'calculate' with kwargs {'expr': '2 * 3'} with error:
expression: Field required
Please fix the error and try again.
Code instead of arithmetic
ai calculate {'expression': "__import__('os').getcwd()"}
tool Calculation error: only numbers and + - * / are allowed
A tool raises TimeoutError
TimeoutError: search API timed out - the whole run stopsfrom langchain.agents.middleware import ToolErrorMiddleware
def on_error(exc, request):
if isinstance(exc, TimeoutError):
return f"{request.tool_call['name']} timed out - try again later or answer without it."
return None # anything else still stops the run
agent = create_agent(llm, tools=[calculate, search, search_documents],
middleware=[ToolErrorMiddleware(on_error)])
# tool search timed out - try again later or answer without it.
# ai Search is unavailable right now.from typing import Literal
from pydantic import BaseModel
class Route(BaseModel):
tool: Literal["calculate", "search", "search_documents", "none"]
descriptions = {
"calculate": calculate.description,
"search": search.description,
"search_documents": search_documents.description,
"none": "No tool is needed.",
}
menu = "\n".join(f"- {name}: {text}" for name, text in descriptions.items())
router = ChatOllama(model="llama3.1", temperature=0).with_structured_output(Route)
choice = router.invoke(
f"Choose the one tool that should handle the user's message.\nTools:\n{menu}\n\n"
"User message: How many days a week can I work from home?"
)
print(choice.tool)
# good descriptions: search_documents - 8 of 8 questions routed correctly
# "Search." etc.: none - 6 of 8Tip: Write down ten real questions and the tool each should use. Run them after every change to a tool’s name or description - it is the cheapest test an agent has.
Multi-tool agents at a glance
RAG as a toolWrap the retriever; retrieval becomes optional.
@tool def search_documents(question: str)
create_retriever_toolThe same, built for you (argument: query).
create_retriever_tool(retriever, name, description)
create_agentOne model, several tools, one loop.
create_agent(llm, tools=[calculate, search, search_documents])
Unknown toolAnswered with the list of valid tools.
Error: ... is not a valid tool
Bad argumentsAnswered with the validation error.
Please fix the error and try again.
ToolErrorMiddlewareTurn chosen exceptions into messages.
ToolErrorMiddleware(on_error)
Try it yourself
The code does not change. Swap the content string and the program does something else entirely.
“Script tool calls for five questions and check each reaches the tool you expect.”
“Run the llama3 router with "Search." as both search descriptions. Which questions go wrong?”
“Ask a question that needs the documents and the calculator; script both calls in one reply.”
“Make search raise TimeoutError, run it, then add ToolErrorMiddleware and run again.”
What usually goes wrong
The model’s argument becomes code on your machine. Use an arithmetic-only evaluator.
✗ return str(eval(expression))✓ return str(evaluate(ast.parse(expression, mode="eval").body))With "Search." llama3 sent policy questions to no tool at all.
✗ """Search."""✓ """Search the company's internal documents. Use this for questions about company policies..."""One timeout ends the whole run. Catch inside the tool, or use ToolErrorMiddleware.
Take tenant and user identity from your session, never from the model.
✗ def search_documents(question: str, tenant_id: str)✓ filter={"tenant_id": current_user.tenant_id} # inside the toolEvery schema is sent on every call, and every tool is another wrong choice available. Add tools when the task needs them.
Key points
- A multi-tool agent chooses, per question, which tools to use.
- RAG as a tool makes retrieval optional instead of automatic.
- The model routes by name, description and schema - descriptions decided 2 of 8 choices in our test.
- One question can use several tools, in sequence or in one step.
- Unknown tools and bad arguments come back as error messages; exceptions stop the run unless handled.
- The tool list and the code inside each tool are your control.
- Never let the model supply identity or permissions.
Quick check before you move on
Quiz
- 1.
With vague descriptions, where did llama3 send "How many days a week can I work from home?"
- 2.
The model calls calculate with {"expr": "2 * 3"}. What happens?
- 3.
A tool raises TimeoutError. What happens by default, and how do you change it?
- 4.
Why should tenant_id not be an argument of search_documents?
Interview questions
What is a multi-tool agent?
An agent with access to several tools that selects and executes the appropriate tool, or sequence of tools, for each task.
How do you improve tool selection?
Fewer, single-purpose tools; clear names; descriptions that say when to use each; typed arguments - and a set of test questions with expected tools to measure changes.
Why make RAG a tool instead of a chain?
So retrieval happens only when it helps, and can be combined with other tools - calculation, search - in one conversation.
How do you secure a multi-tool agent?
Limit the tool list, validate every argument, never execute model text as code or SQL, apply permissions from the session inside tools, add timeouts and logging.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...
AI
System Design
Backend
- GraphQL8 modules · 69 lessons planned
- Core Python13 modules · 75 lessons planned
- FastAPI5 sections · 20 lessons
- Node.js14 modules · 206 lessons planned
- Node.js Performance7 chapters · 36 topics
- Event Loop Lifecycle6 phases · 3 scenarios
- Docker & Containerization11 modules · 144 lessons planned
- AWS for Developers14 modules · 219 lessons planned
- CI/CD & DevOps Automation10 modules · 134 lessons planned