LangChain + Ollama Setup
Connect LangChain to the same local Ollama model with ChatOllama, and run a hello-world chain.
What you will be able to do
- Install the LangChain Ollama integration and check that Ollama is ready
- Create a ChatOllama model and configure it
- Call it with invoke(), and read the text and metadata it returns
- Send system and user messages the LangChain way
- Build a first chain with prompt | llm
- Diagnose the three setup errors you are most likely to hit
The idea, in plain English
In Module 1 your Python code called Ollama directly with ollama.chat(). Now LangChain sits in between: your code talks to a LangChain model object called ChatOllama, which talks to the same Ollama server, which runs the same model.
Nothing about the model changes. To prove it, we sent the same prompt at temperature 0 through ollama.chat() and through ChatOllama - and got the same reply, word for word. What changes is the interface: instead of building a dict and digging the text out of response["message"]["content"], you call llm.invoke() and read response.content.
That interface is the point. Every LangChain component - prompts, models, parsers, whole chains - is run with invoke(), so they can be joined together. The first such join, prompt | llm, appears at the end of this lesson; Lesson 2.4 explains it properly.
Everything here was run against a local model, and the outputs shown are real. They were produced with llama3; the code says llama3.1, the model this course uses - swap in whichever name ollama list shows.
Worked example: Hello-world chain.
Module 1 on the left, Module 2 on the right. Ollama and the model are identical - LangChain adds an interface on top.
Getting set up
You need two Python packages - langchain and the Ollama integration, langchain-ollama - plus everything from Module 1: the Ollama server running, and a model pulled. ChatOllama lives in langchain_ollama, not in langchain itself.
Ollama is still not the model. Ollama is the local runner; llama3.1 is one model it runs. ChatOllama connects to the runner and names the model.
ChatOllama and invoke()
ChatOllama(model="llama3.1") creates a model object - it does not contact Ollama yet. The request happens when you call llm.invoke("Hello"): the text goes to Ollama, the model generates a reply, and invoke returns it.
The reply is an AIMessage, not a dict. Its .content is the text. It also carries details you had to dig for in Module 1: response.usage_metadata gives the token counts - {"input_tokens": 23, "output_tokens": 44, "total_tokens": 67} in our run - and response.response_metadata holds the model name, timings, and why generation stopped.
Make the callModule 1: ollama.chat(model=..., messages=[...]). LangChain: llm.invoke(...).Read the textModule 1: response["message"]["content"]. LangChain: response.content.Token countsModule 1: response["prompt_eval_count"]. LangChain: usage_metadata.SettingsModule 1: options={...} on every call. LangChain: set once on the model.Configuration
Settings go on the model object once, instead of in an options dict on every call: ChatOllama(model="llama3.1", temperature=0). Temperature 0 behaves exactly as it did in Module 1 - three identical runs in our test - and is the right choice whenever you will parse the output.
Leave a setting out and ChatOllama sends nothing for it, so Ollama’s own default applies. That includes num_ctx, the context window from Lesson 1.8 - set it on the model when a conversation is long.
Messages, the LangChain way
invoke() accepts a plain string, which becomes a single user message. For anything more, pass a list of message objects: SystemMessage for standing instructions, HumanMessage for the user, AIMessage for an earlier reply. They are the system, user, and assistant roles from Lesson 1.1 with class names.
A conversation is just that list in order - system, human, AI, human - which is what memory and agents will build on later in this module.
A first chain
ChatPromptTemplate.from_template("Say hello to {name} in a friendly way.") is a prompt with a gap. Invoked with {"name": "Chandu"}, it produces the finished message. prompt | llm joins the two, so the template’s output becomes the model’s input: chain.invoke({"name": "Chandu"}) returned "Hello Chandu! It’s great to meet you!"
Leave the variable out and the chain stops before reaching the model, with a clear error: "Input to ChatPromptTemplate is missing variables {’name’}". Templates are Lesson 2.3, and the | operator is Lesson 2.4.
The three setup errors
Ollama is not running: ConnectError: [Errno 61] Connection refused. Start ollama serve and leave it running.
The model is not pulled, or the name is wrong: ResponseError: model ’llama3.1’ not found (status code: 404). The name must match ollama list, tag included - in our test llama3 and llama3:latest both worked, but llama3:8b was "not found" even though it names the same model, because only the latest tag had been pulled.
The package is missing: ChatOllama comes from langchain-ollama. pip install langchain-ollama, then from langchain_ollama import ChatOllama.
Tip: A wrong model name only fails when you first call invoke(), which can be far from where you set it. Pass validate_model_on_init=True to check the name as soon as the ChatOllama object is created.
Step-by-step code
pip install langchain langchain-ollama
ollama serve # leave this running
ollama pull llama3.1 # if you have not already
ollama list # the exact name and tag to use# Module 1
import ollama
response = ollama.chat(
model="llama3.1",
messages=[{"role": "user", "content": "Hello"}],
)
print(response["message"]["content"])
# Module 2
from langchain_ollama import ChatOllama
llm = ChatOllama(model="llama3.1")
response = llm.invoke("Hello")
print(response.content)from langchain_ollama import ChatOllama
llm = ChatOllama(model="llama3.1")
response = llm.invoke("Say hello and explain what an AI agent is in one sentence.")
print(type(response).__name__) # AIMessage
print(response.content) # the text
print(response.usage_metadata) # token countsAIMessage
Hello! An AI agent is a computer program that perceives its environment and takes
actions to achieve a specific goal or set of goals, making decisions and adapting to
new situations based on its programming and learning from experience.
{'input_tokens': 23, 'output_tokens': 44, 'total_tokens': 67}from langchain_core.messages import HumanMessage, SystemMessage
from langchain_ollama import ChatOllama
llm = ChatOllama(model="llama3.1", temperature=0)
messages = [
SystemMessage(content="You are a Python teacher. Explain concepts using simple English."),
HumanMessage(content="What is a list comprehension?"),
]
response = llm.invoke(messages)
print(response.content)As a Python teacher, I'm excited to explain list comprehensions in a way that's easy
to understand!
A list comprehension is a way to create a new list from an existing list or other
iterable (like a string or a dictionary) using a concise and readable syntax. ...from langchain_core.prompts import ChatPromptTemplate
from langchain_ollama import ChatOllama
llm = ChatOllama(model="llama3.1", temperature=0)
prompt = ChatPromptTemplate.from_template("Say hello to {name} in a friendly way.")
chain = prompt | llm # the prompt's output becomes the model's input
response = chain.invoke({"name": "Chandu"})
print(response.content)Hello Chandu! It's great to meet you! How's your day going so far?
prompt.invoke({"name": "Chandu"}) on its own:
messages=[HumanMessage(content='Say hello to Chandu in a friendly way.')]
chain.invoke({}) - the variable missing:
KeyError: Input to ChatPromptTemplate is missing variables {'name'}.
Expected: ['name'] Received: []Ollama not running:
ConnectError: [Errno 61] Connection refused
Model name not pulled:
ChatOllama(model="llama3.1").invoke("Hi")
ResponseError: model 'llama3.1' not found (status code: 404)
Same model, different tag - only llama3:latest was pulled:
ChatOllama(model="llama3") works
ChatOllama(model="llama3:latest") works
ChatOllama(model="llama3:8b") ResponseError: model 'llama3:8b' not found
Checked at construction instead:
ChatOllama(model="llama3.1", validate_model_on_init=True)
ValidationError: Model llama3.1 not found in Ollama. Please pull the model ...Tip: Same prompt, same model, temperature 0: ollama.chat() and ChatOllama returned identical text in our test. When LangChain output surprises you, try the same prompt through ollama.chat() - if it does the same thing, the model is the cause, not the framework.
ChatOllama at a glance
ChatOllamaLangChain’s chat model for Ollama.
from langchain_ollama import ChatOllama
modelThe name from ollama list, tag included.
ChatOllama(model="llama3.1")
temperatureVariation; 0 for anything you will parse.
ChatOllama(..., temperature=0)
num_ctxThe context window - Ollama’s default if not set.
ChatOllama(..., num_ctx=8192)
invokeSend input, get an AIMessage back.
llm.invoke("Hello").contentThe generated text.
response.content
usage_metadataInput, output, and total tokens.
response.usage_metadata
prompt | llmA chain: the prompt feeds the model.
chain.invoke({"name": "Chandu"})Try it yourself
The code does not change. Swap the content string and the program does something else entirely.
“Send the same prompt at temperature 0 through ollama.chat() and ChatOllama, and compare the text.”
“Use a model name that ollama list does not show, with and without validate_model_on_init=True.”
“Print response.usage_metadata for a short prompt and a long one.”
“Change the template to "Explain {topic} to a 10-year-old." and invoke it with a topic.”
What usually goes wrong
LangChain still needs the Ollama server. The error is a connection refusal, not a LangChain bug.
The name must match ollama list exactly, tag included - llama3:8b failed in our test while llama3 worked.
✗ ChatOllama(model="llama3.1") # never pulled✓ ollama pull llama3.1
ChatOllama(model="llama3.1")invoke returns an AIMessage. The text is an attribute.
✗ response["message"]["content"]✓ response.contentChatOllama is in the separate langchain-ollama package.
✗ from langchain import ChatOllama✓ from langchain_ollama import ChatOllamaThe chain stops before the model with a KeyError naming the missing variable.
✗ chain.invoke({})✓ chain.invoke({"name": "Chandu"})Key points
- ChatOllama connects LangChain to the same local Ollama server and model.
- model= names the Ollama model - it must match ollama list, tag included.
- llm.invoke() sends input; it returns an AIMessage.
- response.content is the text; response.usage_metadata has the token counts.
- Settings such as temperature go on the model once; unset ones use Ollama’s defaults.
- SystemMessage, HumanMessage, and AIMessage are Module 1’s roles as classes.
- prompt | llm is a chain: the prompt’s output becomes the model’s input.
Quick check before you move on
Quiz
- 1.
You create ChatOllama(model="llama3.1") but never pulled that model. When does it fail?
- 2.
ollama list shows llama3:latest. Will ChatOllama(model="llama3:8b") work?
- 3.
What type does llm.invoke() return, and where are the token counts?
- 4.
Same prompt, same model, temperature 0 - through ollama.chat() and through ChatOllama. What should you expect?
Interview questions
How would you connect a local Ollama model to LangChain?
Install langchain-ollama, create ChatOllama with the exact model name from ollama list, and call it with invoke(). The request is still served by the local Ollama server; LangChain only provides the interface.
What does LangChain’s common invoke() interface buy you?
Prompts, models, parsers, and whole chains all run the same way, so they can be composed with | and swapped - a different model provider, say - without rewriting the code around them.
An LLM app works locally but fails on a teammate’s machine with "model not found". What do you check?
That the model is pulled on their machine under exactly the name and tag the code uses, and that Ollama is running. validate_model_on_init=True surfaces the problem at start-up instead of on the first request.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...
AI
System Design
Backend
- GraphQL8 modules · 69 lessons planned
- Core Python13 modules · 75 lessons planned
- FastAPI5 sections · 20 lessons
- Node.js14 modules · 206 lessons planned
- Node.js Performance7 chapters · 36 topics
- Event Loop Lifecycle6 phases · 3 scenarios
- Docker & Containerization11 modules · 144 lessons planned
- AWS for Developers14 modules · 219 lessons planned
- CI/CD & DevOps Automation10 modules · 134 lessons planned