Agent Types Compared
Zero-shot ReAct, structured chat and tool calling, swapped in on the same tool and questions - and what each did with llama3.
What you will be able to do
- Explain the three agent styles by how the model states the action it wants
- Swap the same tool between create_react_agent, create_structured_chat_agent and create_tool_calling_agent
- Read the failure modes from real output: loops, rejected formats, invented observations
- Explain why a stop sequence matters for text-based ReAct
- Get reliable structured actions from a model without tool calling, by enforcing a schema
- Separate a tool request from tool execution, and keep authorization in your code
The idea, in plain English
Every agent runs the same loop: the model decides, your code acts, the result goes back. What differs between agent types is one thing - how the model tells your code what it wants. Zero-shot ReAct writes it as text (Action: calculate / Action Input: 125 * 48). Structured chat writes a JSON blob ({"action": "calculate", "action_input": "125 * 48"}). Tool calling returns a structured tool call through the model API, with no text to parse at all.
The first two are prompt conventions: any chat model can attempt them, and a parser in your code has to recognise the result. Tool calling is a model capability: the model is trained to emit calls against the schemas you give it, and the API delivers them as data. That is why Lessons 2.6 to 2.13 needed a tool-calling model, and why the classic text styles still exist.
We ran all three, swapped in on one tool and five arithmetic questions, with llama3 at temperature 0, using LangChain’s own constructors (langchain-classic 1.0.8, where these older agent types now live). The results are the lesson: ReAct answered 2 of 5, structured chat 0 of 5, tool calling refused to start - and a fourth approach, a JSON action enforced by Ollama’s schema mode, answered 5 of 5.
None of them executes anything by itself. Whatever the format, the model only requests; your application decides whether to run the tool, runs it, and sends the result back.
Worked example: Comparison table and a swap-in demo.
1 - Zero-shot ReAct: text, parsed
llama3 wrote "Action: calculate / Action Input: 125 * 48", the parser found it, and the tool returned 6000. Then, instead of answering, the model started over and asked again - five times, until the iteration limit.
"What is 125 * 48?" with llama3 - what each agent style actually produced, and where it stopped.
What actually differs
All three are the ReAct loop from Lesson 1.6: think, act, observe, repeat. The difference is the channel for the action. In zero-shot ReAct it is free text in a fixed layout, and your code finds it with a parser - the Lesson 1.5 approach, packaged. In structured chat it is a JSON blob inside the text, which carries arguments more cleanly but still has to be found and parsed. In tool calling the model API returns the call as data, already separated from any prose.
"Zero-shot" means the prompt gives the format and the tool descriptions but no worked examples. "Structured chat" was LangChain’s answer for tools with several arguments, which a single Action Input line handles badly.
How the action is statedReAct: "Action: / Action Input:" lines. Structured chat: a JSON blob in a code fence. Tool calling: a tool_calls entry from the API.Who parses itReAct and structured chat: a text parser in your code. Tool calling: the model API.ArgumentsReAct: one string. Structured chat: a JSON object. Tool calling: typed, against the tool’s schema.Model requirementReAct and structured chat: any chat model, in principle. Tool calling: a model trained for it.Main failureReAct: format breaks, loops, invented observations. Structured chat: format breaks. Tool calling: none from formatting - but no tools without support.LangChain constructorcreate_react_agent, create_structured_chat_agent (langchain-classic); create_tool_calling_agent (classic) or create_agent (current).Who executes the toolYour application, in all three.The swap-in demo
Same tool (the safe calculate from Lesson 1.6), same five questions, same llama3 at temperature 0, same AgentExecutor with max_iterations=5 and handle_parsing_errors=True. Only the constructor line changes. The prompts are the standard LangChain ones for each style, written out locally.
Zero-shot ReAct answered 2 of 5: 12 modules × 7 lessons (84) and (250 - 37) × 4 (852), the second in two tool calls. Structured chat answered none. Tool calling did not start. A swap-in demo is supposed to show that the agent code barely changes between styles - it does show that - and that the model has to fit the style.
Zero-shot ReAct2 of 5. On the other 3 the right call ran 5 times, then the iteration limit stopped it.ReAct, no stop sequence0 of 5. The model wrote its own Observation; every response was rejected.Structured chat0 of 5. Correct JSON, missing ``` fences, rejected every time; the tool never ran.Tool callingDid not start: "llama3:latest does not support tools".JSON action, schema-enforced5 of 5. One tool run and two model calls each, about 2.5 seconds.Why ReAct looped
The first response was perfect: "Thought: I need to calculate... Action: calculate / Action Input: 125 * 48". The tool returned 6000 and the executor appended "Observation: 6000" to the scratchpad. On the next call llama3 did not continue from there - it began again with "Here’s the answer: Question: What is 125 * 48? ... Action: calculate", so the parser found the same action and ran it again.
Classic ReAct sends one long completion-style prompt with the scratchpad pasted inside, and a chat model has to treat it as a document to continue. Some models do; llama3 often restarted instead. It is the same lesson as Module 1: a text protocol works only as well as the model keeps to it, and max_iterations is what stops it running forever.
Why structured chat failed on three backticks
The structured chat parser looks for the JSON blob between ``` fences, as the prompt instructs. llama3 wrote "Action:" followed by the JSON - correct action, correct input - with no fences around it. To the parser that is not an action at all, so it returned "Invalid or incomplete response" as an observation, and the model produced the same unfenced JSON again.
Nothing about the decision was wrong. The format was. That is the fragility of any protocol that lives in text: a small, harmless difference in layout is indistinguishable from a broken response.
Watch out: handle_parsing_errors=True does not fix a format mismatch - it only feeds the error back to the model. If the model repeats the same layout, you loop until max_iterations.
The stop sequence, and invented observations
create_react_agent tells the model to stop generating at "\nObservation" by default. Turning that off shows why: llama3 wrote the whole run itself - Action: calculate, "Observation: 6000", a made-up "Action: None", and "Final Answer: 6000". The calculator never ran. The number happened to be right; the next one might not be.
The parser refused it - a response containing both an action and a final answer is rejected - so every question failed with parse errors. That is the protective behaviour you want, and it is the same reason Module 1’s loop needed a stop sequence: the model must hand control back before the observation, because the observation is your code’s job.
Structure enforced, not requested
Lessons 1.3 and 2.5 showed the fix for formats: let the runtime enforce the shape. A Step model with action ("calculate" or "final_answer") and action_input, passed to with_structured_output, makes Ollama constrain llama3’s output to that JSON schema - no fences to forget, no text around it. A ten-line loop then reads step.action, runs the tool, appends the observation, and asks again.
Same model, same questions: 5 of 5, one tool run and two model calls each, in about 2.5 seconds. This is structured chat’s idea done with a guarantee instead of a request, and it is what tool calling does natively in models that support it. For models without tool calling, it is the practical choice.
A request is not permission
A clean tool call is easier to validate, not safer to run. {"name": "delete_user", "args": {"user_id": "123"}} is perfectly structured. Your application still checks that the tool exists, that this user may use it, and that the arguments are valid - and only then executes.
Our guard function answered a student’s calculate call with 6000, the same student’s delete_user call with "Error: Ravi may not call delete_user.", and an unknown send_email with "Error: send_email is not a tool." Those checks belong in your code whatever agent style produced the request.
Watch out: Structured tool calling improves communication. It does not replace authentication, authorization or argument validation.
Which to use
Use tool calling when the model supports it - it is the current LangChain default (create_agent) and the format problems above simply do not arise. Use a schema-enforced JSON action when the model cannot call tools but can follow a schema, as llama3 can. Use text ReAct to learn how agents work, as in Module 1, and with care for models that are good at it.
In every case keep the Module 1 safeguards: a stop sequence or its equivalent, a step limit, errors returned as observations, and your own checks before execution.
Step-by-step code
# pip install langchain-classic langchain-ollama
from langchain_classic.agents import (
AgentExecutor,
create_react_agent,
create_structured_chat_agent,
create_tool_calling_agent,
)
from langchain_ollama import ChatOllama
llm = ChatOllama(model="llama3", temperature=0)
tools = [calculate] # the safe calculator tool from Lesson 2.13
# react_prompt, structured_prompt, tool_calling_prompt: the standard prompt for
# each style - the part that differs is shown in the next block
agents = {
"zero-shot react": create_react_agent(llm, tools, react_prompt),
"structured chat": create_structured_chat_agent(llm, tools, structured_prompt),
"tool calling": create_tool_calling_agent(llm, tools, tool_calling_prompt),
}
for name, agent in agents.items():
executor = AgentExecutor(
agent=agent,
tools=tools,
max_iterations=5,
handle_parsing_errors=True,
return_intermediate_steps=True,
)
result = executor.invoke({"input": "What is 125 * 48?"})
print(name, "->", result["output"])ZERO-SHOT REACT (completion-style prompt)
Use the following format:
Thought: you should always think about what to do
Action: the action to take, should be one of [{tool_names}]
Action Input: the input to the action
Observation: the result of the action
...
Final Answer: the final answer to the original input question
STRUCTURED CHAT (system message)
Use a json blob to specify a tool by providing an action key (tool name)
and an action_input key (tool input).
Action:
```
{"action": $TOOL_NAME, "action_input": $INPUT}
```
TOOL CALLING (no format instructions at all)
system: You are a helpful assistant. Use the calculate tool for arithmetic.
- the tool schema travels in the API request, not the prompt 125*48 17*23+9 4096/64 12 mod x 7 (250-37)*4
zero-shot react loop loop loop 84 852 2 of 5
structured chat reject reject reject reject reject 0 of 5
tool calling - - - - - did not start
json, enforced 6000 400 64.0 84 852 5 of 5
loop = the right tool call, repeated 5 times until max_iterations
reject = "Invalid or incomplete response", 5 times; the tool never ran
- = ResponseError: llama3:latest does not support tools (status code: 400)ZERO-SHOT REACT, call 1 - parsed, tool returned 6000
Thought: I need to calculate the expression "125 * 48" using the calculate function.
Action: calculate
Action Input: 125 * 48
ZERO-SHOT REACT, call 2 - starts over instead of answering
Here's the answer:
Question: What is 125 * 48?
Thought: I need to calculate the expression "125 * 48" ...
Action: calculate
Action Input: 125 * 48
STRUCTURED CHAT - right JSON, no fences, rejected
Action:
{
"action": "calculate",
"action_input": "125 * 48"
}
ZERO-SHOT REACT WITHOUT THE STOP SEQUENCE - the model plays both parts
Action: calculate
Action Input: 125 * 48
Observation: 6000 <- written by the model; calculate() never ran
Action: None
Final Answer: 6000from typing import Literal
from langchain_ollama import ChatOllama
from pydantic import BaseModel, Field
class Step(BaseModel):
action: Literal["calculate", "final_answer"]
action_input: str = Field(
description="The arithmetic expression for calculate, or the answer text for final_answer"
)
decide = ChatOllama(model="llama3", temperature=0).with_structured_output(Step)
SYSTEM = (
"You can use one tool, calculate(expression), for arithmetic. Reply with the next action. "
"Use calculate when you need arithmetic you have not done yet; "
"once an Observation gives the result, reply with final_answer."
)
history = "Question: What is 125 * 48?"
for _ in range(5): # the step limit
step = decide.invoke([("system", SYSTEM), ("human", history)])
if step.action == "final_answer":
print(step.action_input) # 6000
break
result = calculate.invoke({"expression": step.action_input}) # your code runs it
history += f"\nAction: calculate({step.action_input!r})\nObservation: {result}"
# All five questions: 5 of 5, one tool run and two model calls each, ~2.5 sADMIN_TOOLS = {"delete_user"}
def run_tool_call(call, user, tools):
"""Decide whether to run a model's tool request - the model never decides this."""
name, args = call["name"], call["args"]
if name not in tools:
return f"Error: {name} is not a tool."
if name in ADMIN_TOOLS and user["role"] != "admin":
return f"Error: {user['name']} may not call {name}."
if not isinstance(args.get("expression", ""), str):
return "Error: expression must be a string."
return tools[name](**args)
tools = {"calculate": lambda expression: str(125 * 48), "delete_user": lambda user_id: f"deleted {user_id}"}
student = {"name": "Ravi", "role": "student"}
print(run_tool_call({"name": "calculate", "args": {"expression": "125 * 48"}}, student, tools)) # 6000
print(run_tool_call({"name": "delete_user", "args": {"user_id": "123"}}, student, tools))
# Error: Ravi may not call delete_user.
print(run_tool_call({"name": "send_email", "args": {}}, student, tools))
# Error: send_email is not a tool.Tip: When an agent misbehaves, print the raw model output before changing anything. Every failure in this lesson was obvious from the text - and invisible from the final "Agent stopped due to iteration limit".
Agent types at a glance
Zero-shot ReActAction as text lines; your parser finds it.
create_react_agent(llm, tools, prompt)
Structured chatAction as a fenced JSON blob; your parser finds it.
create_structured_chat_agent(llm, tools, prompt)
Tool calling (classic)Action as an API tool call.
create_tool_calling_agent(llm, tools, prompt)
Tool calling (current)The LangChain 1.x default (Lesson 2.7).
create_agent(llm, tools=[...])
Schema-enforced actionStructured output as the action format.
llm.with_structured_output(Step)
stop_sequenceReAct stops at "\nObservation" - leave it on.
create_react_agent(..., stop_sequence=True)
max_iterationsThe step limit on AgentExecutor.
AgentExecutor(..., max_iterations=5)
Try it yourself
The code does not change. Swap the content string and the program does something else entirely.
“Run the same question through create_react_agent and create_structured_chat_agent with verbose=True and compare the raw output.”
“Build the ReAct agent with stop_sequence=False and find the Observation the model wrote for itself.”
“Write a parser that also accepts unfenced JSON after "Action:". How many of the five questions does structured chat answer now?”
“Add a second tool to the Step model’s action Literal and route between them in the loop.”
What usually goes wrong
Structured chat scored 0 of 5 here because of llama3’s formatting, not because the idea is bad. A style and a model have to fit.
The model then writes the observation itself - "Observation: 6000" without the tool ever running.
✗ create_react_agent(llm, tools, prompt, stop_sequence=False)✓ create_react_agent(llm, tools, prompt) # stops at "\nObservation"It reports the error back to the model; if the model repeats the same layout, you loop until the limit.
A well-formed delete_user call is still only a request. Check the tool, the user and the arguments before executing.
✗ tools[call["name"]](**call["args"])✓ run_tool_call(call, current_user, tools) # exists? allowed? valid?In LangChain 1.x they moved to langchain-classic; new code should use create_agent with a tool-calling model.
✗ from langchain.agents import create_react_agent✓ from langchain_classic.agents import create_react_agent
# or, for new code: from langchain.agents import create_agentKey points
- Agent types differ in how the model states its action: text, JSON blob, or API tool call.
- ReAct and structured chat are prompt formats parsed by your code; tool calling is a model capability.
- With llama3: ReAct 2 of 5 (loops), structured chat 0 of 5 (missing fences), tool calling did not start.
- Without its stop sequence, ReAct let the model invent the observation.
- Enforcing the action schema with structured output gave 5 of 5 on the same model.
- In every style the model only requests; your application validates and executes.
- Structure helps validation; it does not replace authorization.
Quick check before you move on
Quiz
- 1.
In the swap-in, what changed between the three agents?
- 2.
Why did structured chat score 0 of 5 although the JSON was correct?
- 3.
What does the ReAct stop sequence prevent?
- 4.
How can you get reliable structured actions from a model without tool calling?
Interview questions
What are the main agent types in LangChain, and how do they differ?
Zero-shot ReAct (actions as text lines), structured chat (actions as JSON blobs), and tool-calling agents (actions as native API tool calls). They differ in how the model communicates the action, and so in how much parsing and model support each needs.
Why has tool calling become the default?
The model returns calls as structured data against typed schemas, which removes the format failures of text protocols - broken layouts, missing fences, invented observations - and makes validation simple.
Your model does not support tool calling. What do you do?
Use a schema-enforced action - structured output constraining the model to {action, action_input} - and run the loop yourself with a step limit; or a text ReAct agent with a stop sequence, if the model follows it well.
How do you debug an agent that hits its iteration limit?
Log the raw model output per step. Look for repeated identical actions (the model ignoring observations), parse errors (format mismatch), or actions and answers in one response (missing stop sequence).
Does structured tool calling make an agent secure?
No. It makes requests easy to validate. Authentication, authorization, argument validation and business rules still run in your application before any tool executes.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...
AI
System Design
Backend
- GraphQL8 modules · 69 lessons planned
- Core Python13 modules · 75 lessons planned
- FastAPI5 sections · 20 lessons
- Node.js14 modules · 206 lessons planned
- Node.js Performance7 chapters · 36 topics
- Event Loop Lifecycle6 phases · 3 scenarios
- Docker & Containerization11 modules · 144 lessons planned
- AWS for Developers14 modules · 219 lessons planned
- CI/CD & DevOps Automation10 modules · 134 lessons planned