← Back to Agentic AI map
Lesson 2.2 · Building Agents with LangChain

LangChain + Ollama Setup

Connect LangChain to the same local Ollama model with ChatOllama, and run a hello-world chain.

setup

What you will be able to do

  • Install the LangChain Ollama integration and check that Ollama is ready
  • Create a ChatOllama model and configure it
  • Call it with invoke(), and read the text and metadata it returns
  • Send system and user messages the LangChain way
  • Build a first chain with prompt | llm
  • Diagnose the three setup errors you are most likely to hit

The idea, in plain English

In Module 1 your Python code called Ollama directly with ollama.chat(). Now LangChain sits in between: your code talks to a LangChain model object called ChatOllama, which talks to the same Ollama server, which runs the same model.

Nothing about the model changes. To prove it, we sent the same prompt at temperature 0 through ollama.chat() and through ChatOllama - and got the same reply, word for word. What changes is the interface: instead of building a dict and digging the text out of response["message"]["content"], you call llm.invoke() and read response.content.

That interface is the point. Every LangChain component - prompts, models, parsers, whole chains - is run with invoke(), so they can be joined together. The first such join, prompt | llm, appears at the end of this lesson; Lesson 2.4 explains it properly.

Everything here was run against a local model, and the outputs shown are real. They were produced with llama3; the code says llama3.1, the model this course uses - swap in whichever name ollama list shows.

Worked example: Hello-world chain.

ArchitectureSame model, one more layer

Module 1 on the left, Module 2 on the right. Ollama and the model are identical - LangChain adds an interface on top.

Getting set up

You need two Python packages - langchain and the Ollama integration, langchain-ollama - plus everything from Module 1: the Ollama server running, and a model pulled. ChatOllama lives in langchain_ollama, not in langchain itself.

Ollama is still not the model. Ollama is the local runner; llama3.1 is one model it runs. ChatOllama connects to the runner and names the model.

ChatOllama and invoke()

ChatOllama(model="llama3.1") creates a model object - it does not contact Ollama yet. The request happens when you call llm.invoke("Hello"): the text goes to Ollama, the model generates a reply, and invoke returns it.

The reply is an AIMessage, not a dict. Its .content is the text. It also carries details you had to dig for in Module 1: response.usage_metadata gives the token counts - {"input_tokens": 23, "output_tokens": 44, "total_tokens": 67} in our run - and response.response_metadata holds the model name, timings, and why generation stopped.

Module 1 and LangChain, side by side
Make the callModule 1: ollama.chat(model=..., messages=[...]). LangChain: llm.invoke(...).
Read the textModule 1: response["message"]["content"]. LangChain: response.content.
Token countsModule 1: response["prompt_eval_count"]. LangChain: usage_metadata.
SettingsModule 1: options={...} on every call. LangChain: set once on the model.

Configuration

Settings go on the model object once, instead of in an options dict on every call: ChatOllama(model="llama3.1", temperature=0). Temperature 0 behaves exactly as it did in Module 1 - three identical runs in our test - and is the right choice whenever you will parse the output.

Leave a setting out and ChatOllama sends nothing for it, so Ollama’s own default applies. That includes num_ctx, the context window from Lesson 1.8 - set it on the model when a conversation is long.

Messages, the LangChain way

invoke() accepts a plain string, which becomes a single user message. For anything more, pass a list of message objects: SystemMessage for standing instructions, HumanMessage for the user, AIMessage for an earlier reply. They are the system, user, and assistant roles from Lesson 1.1 with class names.

A conversation is just that list in order - system, human, AI, human - which is what memory and agents will build on later in this module.

A first chain

ChatPromptTemplate.from_template("Say hello to {name} in a friendly way.") is a prompt with a gap. Invoked with {"name": "Chandu"}, it produces the finished message. prompt | llm joins the two, so the template’s output becomes the model’s input: chain.invoke({"name": "Chandu"}) returned "Hello Chandu! It’s great to meet you!"

Leave the variable out and the chain stops before reaching the model, with a clear error: "Input to ChatPromptTemplate is missing variables {’name’}". Templates are Lesson 2.3, and the | operator is Lesson 2.4.

The three setup errors

Ollama is not running: ConnectError: [Errno 61] Connection refused. Start ollama serve and leave it running.

The model is not pulled, or the name is wrong: ResponseError: model ’llama3.1’ not found (status code: 404). The name must match ollama list, tag included - in our test llama3 and llama3:latest both worked, but llama3:8b was "not found" even though it names the same model, because only the latest tag had been pulled.

The package is missing: ChatOllama comes from langchain-ollama. pip install langchain-ollama, then from langchain_ollama import ChatOllama.

Tip: A wrong model name only fails when you first call invoke(), which can be far from where you set it. Pass validate_model_on_init=True to check the name as soon as the ChatOllama object is created.

Step-by-step code

Install and check
pip install langchain langchain-ollama ollama serve # leave this running ollama pull llama3.1 # if you have not already ollama list # the exact name and tag to use
Module 1 and LangChain - the same call
# Module 1 import ollama response = ollama.chat( model="llama3.1", messages=[{"role": "user", "content": "Hello"}], ) print(response["message"]["content"]) # Module 2 from langchain_ollama import ChatOllama llm = ChatOllama(model="llama3.1") response = llm.invoke("Hello") print(response.content)
Hello world, and what comes back
from langchain_ollama import ChatOllama llm = ChatOllama(model="llama3.1") response = llm.invoke("Say hello and explain what an AI agent is in one sentence.") print(type(response).__name__) # AIMessage print(response.content) # the text print(response.usage_metadata) # token counts
Output - from a real run
AIMessage Hello! An AI agent is a computer program that perceives its environment and takes actions to achieve a specific goal or set of goals, making decisions and adapting to new situations based on its programming and learning from experience. {'input_tokens': 23, 'output_tokens': 44, 'total_tokens': 67}
System and user messages
from langchain_core.messages import HumanMessage, SystemMessage from langchain_ollama import ChatOllama llm = ChatOllama(model="llama3.1", temperature=0) messages = [ SystemMessage(content="You are a Python teacher. Explain concepts using simple English."), HumanMessage(content="What is a list comprehension?"), ] response = llm.invoke(messages) print(response.content)
Output - from a real run
As a Python teacher, I'm excited to explain list comprehensions in a way that's easy to understand! A list comprehension is a way to create a new list from an existing list or other iterable (like a string or a dictionary) using a concise and readable syntax. ...
The hello-world chain
from langchain_core.prompts import ChatPromptTemplate from langchain_ollama import ChatOllama llm = ChatOllama(model="llama3.1", temperature=0) prompt = ChatPromptTemplate.from_template("Say hello to {name} in a friendly way.") chain = prompt | llm # the prompt's output becomes the model's input response = chain.invoke({"name": "Chandu"}) print(response.content)
Output, and what each piece produced
Hello Chandu! It's great to meet you! How's your day going so far? prompt.invoke({"name": "Chandu"}) on its own: messages=[HumanMessage(content='Say hello to Chandu in a friendly way.')] chain.invoke({}) - the variable missing: KeyError: Input to ChatPromptTemplate is missing variables {'name'}. Expected: ['name'] Received: []
The setup errors - from real runs
Ollama not running: ConnectError: [Errno 61] Connection refused Model name not pulled: ChatOllama(model="llama3.1").invoke("Hi") ResponseError: model 'llama3.1' not found (status code: 404) Same model, different tag - only llama3:latest was pulled: ChatOllama(model="llama3") works ChatOllama(model="llama3:latest") works ChatOllama(model="llama3:8b") ResponseError: model 'llama3:8b' not found Checked at construction instead: ChatOllama(model="llama3.1", validate_model_on_init=True) ValidationError: Model llama3.1 not found in Ollama. Please pull the model ...

Tip: Same prompt, same model, temperature 0: ollama.chat() and ChatOllama returned identical text in our test. When LangChain output surprises you, try the same prompt through ollama.chat() - if it does the same thing, the model is the cause, not the framework.

ChatOllama at a glance

ChatOllama

LangChain’s chat model for Ollama.

from langchain_ollama import ChatOllama
model

The name from ollama list, tag included.

ChatOllama(model="llama3.1")
temperature

Variation; 0 for anything you will parse.

ChatOllama(..., temperature=0)
num_ctx

The context window - Ollama’s default if not set.

ChatOllama(..., num_ctx=8192)
invoke

Send input, get an AIMessage back.

llm.invoke("Hello")
.content

The generated text.

response.content
usage_metadata

Input, output, and total tokens.

response.usage_metadata
prompt | llm

A chain: the prompt feeds the model.

chain.invoke({"name": "Chandu"})

Try it yourself

The code does not change. Swap the content string and the program does something else entirely.

Compare

“Send the same prompt at temperature 0 through ollama.chat() and ChatOllama, and compare the text.”

Break the name

“Use a model name that ollama list does not show, with and without validate_model_on_init=True.”

Read the metadata

“Print response.usage_metadata for a short prompt and a long one.”

Change the chain

“Change the template to "Explain {topic} to a 10-year-old." and invoke it with a topic.”

What usually goes wrong

Ollama is not running

LangChain still needs the Ollama server. The error is a connection refusal, not a LangChain bug.

A model name that is not pulled

The name must match ollama list exactly, tag included - llama3:8b failed in our test while llama3 worked.

✗ ChatOllama(model="llama3.1")  # never pulled
✓ ollama pull llama3.1
ChatOllama(model="llama3.1")
Reading the response like a dict

invoke returns an AIMessage. The text is an attribute.

✗ response["message"]["content"]
✓ response.content
Importing from the wrong package

ChatOllama is in the separate langchain-ollama package.

✗ from langchain import ChatOllama
✓ from langchain_ollama import ChatOllama
Forgetting a template variable

The chain stops before the model with a KeyError naming the missing variable.

✗ chain.invoke({})
✓ chain.invoke({"name": "Chandu"})

Key points

  • ChatOllama connects LangChain to the same local Ollama server and model.
  • model= names the Ollama model - it must match ollama list, tag included.
  • llm.invoke() sends input; it returns an AIMessage.
  • response.content is the text; response.usage_metadata has the token counts.
  • Settings such as temperature go on the model once; unset ones use Ollama’s defaults.
  • SystemMessage, HumanMessage, and AIMessage are Module 1’s roles as classes.
  • prompt | llm is a chain: the prompt’s output becomes the model’s input.

Quick check before you move on

What is ChatOllama?
LangChain’s interface to a chat model running in Ollama.
What does llm.invoke("Hello") do?
Sends "Hello" to the configured model through Ollama and returns the reply as an AIMessage.
Where is the generated text?
In response.content.
What is the difference between ollama.chat(...) and llm.invoke(...)?
The first calls the Ollama Python client directly; the second goes through LangChain’s model interface. The model and the server are the same.
What does prompt | llm represent?
A chain: the filled-in prompt is passed to the model.

Quiz

  1. 1.

    You create ChatOllama(model="llama3.1") but never pulled that model. When does it fail?

  2. 2.

    ollama list shows llama3:latest. Will ChatOllama(model="llama3:8b") work?

  3. 3.

    What type does llm.invoke() return, and where are the token counts?

  4. 4.

    Same prompt, same model, temperature 0 - through ollama.chat() and through ChatOllama. What should you expect?

Interview questions

How would you connect a local Ollama model to LangChain?

Install langchain-ollama, create ChatOllama with the exact model name from ollama list, and call it with invoke(). The request is still served by the local Ollama server; LangChain only provides the interface.

What does LangChain’s common invoke() interface buy you?

Prompts, models, parsers, and whole chains all run the same way, so they can be composed with | and swapped - a different model provider, say - without rewriting the code around them.

An LLM app works locally but fails on a teammate’s machine with "model not found". What do you check?

That the model is pulled on their machine under exactly the name and tag the code uses, and that Ollama is running. validate_model_on_init=True surfaces the problem at start-up instead of on the first request.

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...