← Back to Agentic AI map
Lesson 1.1 · Agents From Scratch

Talking to a Local Model

Send a prompt to a model running on your own machine and read the reply.

basics

What you will be able to do

  • Explain what Ollama does, and how it differs from the model itself
  • Run a language model locally on your own computer
  • Send a prompt to that model from Python
  • Read the generated text out of the response
  • Describe the request and response flow end to end

The idea, in plain English

A large language model - an LLM - is a program that takes some text and predicts the text that should come next. Ollama is a tool that runs one of these models on your own computer and makes it available over a local address, the way a small web server would.

That distinction matters and is worth getting right early: Ollama is not the model. Ollama is the thing that downloads, runs, and manages models; llama3.1 is a model it runs. Swapping the model is a one-word change, because the tool stays the same.

Your Python code sends it a prompt, which is just a piece of text, and gets a reply back. Nothing leaves your machine - no API key, no bill, no network round trip to someone else’s server. That single exchange is the foundation everything else in this course is built on.

Before you run anything here: start Ollama with `ollama serve`, and download a model with `ollama pull llama3.1`. Check what you have with `ollama list`. If you pulled a different model, swap its name into every example.

Worked example: Ask the model to write a haiku about rain.

ArchitectureWhere the model actually runs

The same application, the same prompt. The difference is who owns the machine doing the work - and, on the right, nothing leaves yours.

request flowOne call, step by stepstep 1 / 5

Step 1 - Write the prompt

Your prompt is ordinary text. It goes inside a message object with a role, and that message goes into a list.

where
your code
role
user
network
none yet
cost
nothing

What happens between calling ollama.chat() and printing the answer. Step through it before you run the code.

Step-by-step code

Start Ollama and get a model
ollama serve # start the local server ollama pull llama3.1 # download the model ollama list # check what you have
Install the package
pip install ollama
Your first call
import ollama response = ollama.chat( model="llama3.1", messages=[ {"role": "user", "content": "Write a haiku about rain."} ] ) print(response["message"]["content"])
What you get back
Soft rain taps the leaves Clouds whisper across the sky Earth drinks quietly
Why messages is a list: a conversation
messages = [ {"role": "user", "content": "What is a circuit breaker pattern?"}, {"role": "assistant", "content": "It stops repeated calls to a failing service."}, {"role": "user", "content": "Can you give me a real-world example?"}, ]

Tip: The application stays the same while the prompt changes what it does. That one idea is most of what makes LLMs useful - and most of what makes them hard to test.

Watch out: The same prompt will not give you the same answer twice. Nothing is broken; the model generates rather than looks up. Expect this whenever you compare two runs.

The three message roles

Every message is a dict with two keys: role says who is speaking, content says what they said. This lesson only needs user, but all three appear from the next lesson onward.

user

Something the person said. Your prompts go here.

{"role": "user", "content": "..."}
assistant

Something the model said. You append these yourself to give the model memory of its own replies.

{"role": "assistant", "content": "..."}
system

Standing instructions about how the model should behave. Usually first in the list, and not addressed to anyone.

{"role": "system", "content": "..."}

Try it yourself

The code does not change. Swap the content string and the program does something else entirely.

Explain a concept

“Explain what an API is in simple terms.”

Write code

“Write a Python function that reverses a string.”

Compare two things

“Explain the difference between SQL and NoSQL databases.”

Generate structure

“Create a 5-question quiz about Docker.”

What usually goes wrong

Ollama is not running

The Python package talks to a server. If that server is not up, the call fails before it reaches any model. Start it with ollama serve and leave it running in its own terminal.

The model was never downloaded

Naming a model does not fetch it. Run ollama pull llama3.1 first, and ollama list to confirm what is actually on disk.

The model name does not match

The string in your code has to match the name in ollama list exactly, tag included. llama3.1 and llama3.1:8b are different names.

Passing messages as a string

messages is a list of dicts, not a piece of text. This is the single most common first error, and the reason for it is the whole point of the next lesson - a list can hold a conversation.

✗ messages="Write a haiku about rain."
✓ messages=[{"role": "user", "content": "Write a haiku about rain."}]

Key points

  • Ollama runs models on your machine; the model, such as llama3.1, is a separate thing it runs.
  • ollama.chat() needs two things: a model name and a messages list.
  • Each message is a dict with a role and content.
  • user is what the person said, assistant is what the model said, system is standing instructions.
  • messages is a list because a conversation is a list - history is how the model gets context.
  • The generated text is at response["message"]["content"].
  • The whole pattern is: prompt → Ollama → local model → response.
  • Nothing leaves your computer, so there is no API key and no cost per call.

Quick check before you move on

Where is the model running?
On your own machine, managed by Ollama.
What sends the request?
Your Python application, through the ollama package.
What identifies the model?
The model argument passed to ollama.chat().
Where is the conversation stored?
In the messages list.
Where is the generated text?
At response["message"]["content"].

Quiz

  1. 1.

    What two things does ollama.chat() need at minimum?

  2. 2.

    Where does the model's reply live in the response object?

  3. 3.

    True or false: Ollama sends your prompt to a cloud server.

  4. 4.

    Why is messages a list rather than a single string?

  5. 5.

    What is the difference between Ollama and llama3.1?

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...