Prompting Basics
Set the model’s behaviour with a system message, and control randomness with temperature.
The idea, in plain English
Every message in the list has a role. A "system" message sets how the model should behave overall - think of it as the job description you hand someone before they start work. A "user" message is what the person is actually asking. An "assistant" message is something the model said earlier.
The model reads the system message first and keeps following it for the whole conversation, so that is where instructions about tone, format, and what not to do belong.
The distinction worth holding onto: the system message does not answer anyone. It sets the rules. The user message supplies the task. The assistant message is the result. Or shorter - system says how to behave, user says what to do, assistant says what came back.
Worked example: Make the model always answer in bullet points.
1 - The system message sets the rules
Your application writes this, not the person using it. It says what role to play, how to answer, and what format to use. It is not a question and nothing answers it.
Building a Python tutor. Watch what reaches the model when the user asks one short question.
Turn 1 - two messages go out
The system message and the first question. This is exactly the call from the previous simulation.
The user now asks "How do I add another number?" - a sentence that means nothing on its own. Step through what makes it work.
What is temperature?
Temperature controls how much randomness goes into picking each next word. Near 0 the model takes the safest option every time, so you get steady, repeatable answers. Higher values let it wander, which is what you want for brainstorming and not what you want for anything your code has to parse.
It is passed separately from the messages, in an options dict, because it is a setting on the generation rather than part of the conversation.
0.0Very focused and predictable. The same prompt gives close to the same answer.0.3Low variation. A good default for structured tasks and anything you will parse.0.7More varied and creative. Good for drafting and brainstorming.1.0+Greater variation and experimentation. Expect surprises.Watch out: Temperature does not control intelligence and does not make the model think harder. It only changes how much variation is allowed while generating. Turning it up will not fix a wrong answer.
Step-by-step code
import ollama
response = ollama.chat(
model="llama3.1",
messages=[
{"role": "system", "content": "You always answer using short bullet points. Never write paragraphs."},
{"role": "user", "content": "What are the benefits of exercise?"}
],
options={"temperature": 0.3}
)
print(response["message"]["content"])- Improves cardiovascular health
- Builds muscle and strength
- Helps maintain a healthy weight
- Can improve mood and reduce stress
- Supports better sleepmessages = [
{"role": "system", "content": "You are a Python tutor. Answer in short bullet points."},
{"role": "user", "content": "What is a Python list?"},
]
response = ollama.chat(model="llama3.1", messages=messages, options={"temperature": 0.3})
print(response["message"]["content"])
# The model remembers nothing. Put its reply back into the list yourself.
messages.append(response["message"])
# Now a follow-up that only makes sense with the history above.
messages.append({"role": "user", "content": "How do I add another number?"})
response = ollama.chat(model="llama3.1", messages=messages, options={"temperature": 0.3})
print(response["message"]["content"])Tip: Say what you want, not only what you do not want. "Answer in short bullet points" works better than "do not write paragraphs" on its own - the first one gives the model something to aim at.
Quiz
- 1.
What is the difference between a "system" message and a "user" message?
- 2.
If you want consistent, repeatable answers, should temperature be high or low?
- 3.
What happens if you do not include a "system" message at all?
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...
AI
System Design
Backend
- GraphQL8 modules · 69 lessons planned
- Core Python13 modules · 75 lessons planned
- FastAPI5 sections · 20 lessons
- Node.js14 modules · 206 lessons planned
- Node.js Performance7 chapters · 36 topics
- Event Loop Lifecycle6 phases · 3 scenarios
- Docker & Containerization11 modules · 144 lessons planned
- AWS for Developers14 modules · 219 lessons planned
- CI/CD & DevOps Automation10 modules · 134 lessons planned