← Back to Agentic AI map
Lesson 1.2 · Agents From Scratch

Prompting Basics

Set the model’s behaviour with a system message, and control randomness with temperature.

basics

The idea, in plain English

Every message in the list has a role. A "system" message sets how the model should behave overall - think of it as the job description you hand someone before they start work. A "user" message is what the person is actually asking. An "assistant" message is something the model said earlier.

The model reads the system message first and keeps following it for the whole conversation, so that is where instructions about tone, format, and what not to do belong.

The distinction worth holding onto: the system message does not answer anyone. It sets the rules. The user message supplies the task. The assistant message is the result. Or shorter - system says how to behave, user says what to do, assistant says what came back.

Worked example: Make the model always answer in bullet points.

workflowWhat the model actually receivesstep 1 / 4

1 - The system message sets the rules

Your application writes this, not the person using it. It says what role to play, how to answer, and what format to use. It is not a question and nothing answers it.

written by
your app
says
how to behave
applies to
whole chat
answers anything
no

Building a Python tutor. Watch what reaches the model when the user asks one short question.

workflowWhy the list grows: a follow-up questionstep 1 / 5

Turn 1 - two messages go out

The system message and the first question. This is exactly the call from the previous simulation.

messages
2
system
1
user
1
assistant
0

The user now asks "How do I add another number?" - a sentence that means nothing on its own. Step through what makes it work.

What is temperature?

Temperature controls how much randomness goes into picking each next word. Near 0 the model takes the safest option every time, so you get steady, repeatable answers. Higher values let it wander, which is what you want for brainstorming and not what you want for anything your code has to parse.

It is passed separately from the messages, in an options dict, because it is a setting on the generation rather than part of the conversation.

Temperature, roughly
0.0Very focused and predictable. The same prompt gives close to the same answer.
0.3Low variation. A good default for structured tasks and anything you will parse.
0.7More varied and creative. Good for drafting and brainstorming.
1.0+Greater variation and experimentation. Expect surprises.

Watch out: Temperature does not control intelligence and does not make the model think harder. It only changes how much variation is allowed while generating. Turning it up will not fix a wrong answer.

Step-by-step code

System message plus temperature
import ollama response = ollama.chat( model="llama3.1", messages=[ {"role": "system", "content": "You always answer using short bullet points. Never write paragraphs."}, {"role": "user", "content": "What are the benefits of exercise?"} ], options={"temperature": 0.3} ) print(response["message"]["content"])
What comes back
- Improves cardiovascular health - Builds muscle and strength - Helps maintain a healthy weight - Can improve mood and reduce stress - Supports better sleep
Keeping the conversation: append the reply, then ask again
messages = [ {"role": "system", "content": "You are a Python tutor. Answer in short bullet points."}, {"role": "user", "content": "What is a Python list?"}, ] response = ollama.chat(model="llama3.1", messages=messages, options={"temperature": 0.3}) print(response["message"]["content"]) # The model remembers nothing. Put its reply back into the list yourself. messages.append(response["message"]) # Now a follow-up that only makes sense with the history above. messages.append({"role": "user", "content": "How do I add another number?"}) response = ollama.chat(model="llama3.1", messages=messages, options={"temperature": 0.3}) print(response["message"]["content"])

Tip: Say what you want, not only what you do not want. "Answer in short bullet points" works better than "do not write paragraphs" on its own - the first one gives the model something to aim at.

Quiz

  1. 1.

    What is the difference between a "system" message and a "user" message?

  2. 2.

    If you want consistent, repeatable answers, should temperature be high or low?

  3. 3.

    What happens if you do not include a "system" message at all?

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...