Talking to a Local Model
Send a prompt to a model running on your own machine and read the reply.
What you will be able to do
- Explain what Ollama does, and how it differs from the model itself
- Run a language model locally on your own computer
- Send a prompt to that model from Python
- Read the generated text out of the response
- Describe the request and response flow end to end
The idea, in plain English
A large language model - an LLM - is a program that takes some text and predicts the text that should come next. Ollama is a tool that runs one of these models on your own computer and makes it available over a local address, the way a small web server would.
That distinction matters and is worth getting right early: Ollama is not the model. Ollama is the thing that downloads, runs, and manages models; llama3.1 is a model it runs. Swapping the model is a one-word change, because the tool stays the same.
Your Python code sends it a prompt, which is just a piece of text, and gets a reply back. Nothing leaves your machine - no API key, no bill, no network round trip to someone else’s server. That single exchange is the foundation everything else in this course is built on.
Before you run anything here: start Ollama with `ollama serve`, and download a model with `ollama pull llama3.1`. Check what you have with `ollama list`. If you pulled a different model, swap its name into every example.
Worked example: Ask the model to write a haiku about rain.
The same application, the same prompt. The difference is who owns the machine doing the work - and, on the right, nothing leaves yours.
Step 1 - Write the prompt
Your prompt is ordinary text. It goes inside a message object with a role, and that message goes into a list.
What happens between calling ollama.chat() and printing the answer. Step through it before you run the code.
Step-by-step code
ollama serve # start the local server
ollama pull llama3.1 # download the model
ollama list # check what you havepip install ollamaimport ollama
response = ollama.chat(
model="llama3.1",
messages=[
{"role": "user", "content": "Write a haiku about rain."}
]
)
print(response["message"]["content"])Soft rain taps the leaves
Clouds whisper across the sky
Earth drinks quietlymessages = [
{"role": "user", "content": "What is a circuit breaker pattern?"},
{"role": "assistant", "content": "It stops repeated calls to a failing service."},
{"role": "user", "content": "Can you give me a real-world example?"},
]Tip: The application stays the same while the prompt changes what it does. That one idea is most of what makes LLMs useful - and most of what makes them hard to test.
Watch out: The same prompt will not give you the same answer twice. Nothing is broken; the model generates rather than looks up. Expect this whenever you compare two runs.
The three message roles
Every message is a dict with two keys: role says who is speaking, content says what they said. This lesson only needs user, but all three appear from the next lesson onward.
userSomething the person said. Your prompts go here.
{"role": "user", "content": "..."}assistantSomething the model said. You append these yourself to give the model memory of its own replies.
{"role": "assistant", "content": "..."}systemStanding instructions about how the model should behave. Usually first in the list, and not addressed to anyone.
{"role": "system", "content": "..."}Try it yourself
The code does not change. Swap the content string and the program does something else entirely.
“Explain what an API is in simple terms.”
“Write a Python function that reverses a string.”
“Explain the difference between SQL and NoSQL databases.”
“Create a 5-question quiz about Docker.”
What usually goes wrong
The Python package talks to a server. If that server is not up, the call fails before it reaches any model. Start it with ollama serve and leave it running in its own terminal.
Naming a model does not fetch it. Run ollama pull llama3.1 first, and ollama list to confirm what is actually on disk.
The string in your code has to match the name in ollama list exactly, tag included. llama3.1 and llama3.1:8b are different names.
messages is a list of dicts, not a piece of text. This is the single most common first error, and the reason for it is the whole point of the next lesson - a list can hold a conversation.
✗ messages="Write a haiku about rain."✓ messages=[{"role": "user", "content": "Write a haiku about rain."}]Key points
- Ollama runs models on your machine; the model, such as llama3.1, is a separate thing it runs.
- ollama.chat() needs two things: a model name and a messages list.
- Each message is a dict with a role and content.
- user is what the person said, assistant is what the model said, system is standing instructions.
- messages is a list because a conversation is a list - history is how the model gets context.
- The generated text is at response["message"]["content"].
- The whole pattern is: prompt → Ollama → local model → response.
- Nothing leaves your computer, so there is no API key and no cost per call.
Quick check before you move on
Quiz
- 1.
What two things does ollama.chat() need at minimum?
- 2.
Where does the model's reply live in the response object?
- 3.
True or false: Ollama sends your prompt to a cloud server.
- 4.
Why is messages a list rather than a single string?
- 5.
What is the difference between Ollama and llama3.1?
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...
AI
System Design
Backend
- GraphQL8 modules · 69 lessons planned
- Core Python13 modules · 75 lessons planned
- FastAPI5 sections · 20 lessons
- Node.js14 modules · 206 lessons planned
- Node.js Performance7 chapters · 36 topics
- Event Loop Lifecycle6 phases · 3 scenarios
- Docker & Containerization11 modules · 144 lessons planned
- AWS for Developers14 modules · 219 lessons planned
- CI/CD & DevOps Automation10 modules · 134 lessons planned