Cost and Latency Awareness
Measure what every model call costs in tokens and seconds: reading the prompt, writing the answer, loading the model. Then benchmark llama3 against gemma3 - and see a benchmark lie by 5 times because of how it was run.
What you will be able to do
- Read the token counts and timings Ollama reports for every call
- Explain where the time of a model call goes: load, read, write
- Make answers faster: shorter outputs, num_predict, a stable prompt start
- Estimate the cost of a call or an agent run from token counts
- Benchmark two models fairly - speed and quality together
- Spot measurements you cannot trust
The idea, in plain English
Every model call costs two things: time (latency) and money or compute (cost). With a hosted model you pay per token. With a local model like ours the tokens are free, but they still cost time - and on a busy server, time is money too.
Users feel latency. An agent that makes five model calls of three seconds each keeps a user waiting fifteen seconds. Lesson 5.1’s trace showed our agent making three calls where two would do. Knowing where the time goes is the first step to making it shorter.
In this lesson we measure llama3 and gemma3 on our machine using the numbers Ollama reports with every answer: how many tokens were read and written, and how long each part took. We find what really makes calls slow, discover that repeating the start of a prompt is almost free, and compare two models fairly. Ollama 0.34.1, llama3 (8B) and gemma3 (4.3B), both Q4_K_M.
Worked example: Real measurements with llama3 and gemma3 on Ollama: a cold start, long prompts, long answers, prompt caching, and a quality-and-speed benchmark on the Lesson 5.4 test set.
1 - Load: only the first time
If the model is not in memory, Ollama loads it first. Our cold starts took 1.1 and 4.1 seconds to load. After that, load time was 0.00 s.
Median of repeated real calls with llama3 on our machine.
Words you will see in this lesson
A few words about speed and cost.
TokenA small piece of text the model reads or writes - often part of a word. Models count and charge in tokens.Input tokensTokens the model reads: your prompt, instructions, documents, history.Output tokensTokens the model writes: the answer.LatencyHow long the user waits.ThroughputHow fast tokens come out - tokens per second.Cold startThe first call, when the model must first be loaded into memory.Prompt cachingReusing work for the part of a prompt that was seen before.MedianThe middle value of several measurements - not fooled by one strange run.An everyday example: a taxi ride
A taxi ride costs you in two ways: money and time. The time has parts. Waiting for the taxi to arrive (a cold start). Getting in and saying where you go (reading the prompt). The drive itself (writing the answer) - and the drive is usually the longest part.
To arrive sooner you do not ask the driver to talk faster when you get in. You choose a shorter trip. With models it is the same: the answer length matters most. And if you take the same route every day, the driver already knows the way - that is prompt caching.
What Ollama tells you after every call
Every answer from ChatOllama carries two kinds of numbers. usage_metadata has the token counts: input_tokens, output_tokens and total_tokens - the same names LangChain uses for hosted models, so code that counts cost works for both. response_metadata has Ollama’s timings, in nanoseconds (billionths of a second): load_duration, prompt_eval_duration (reading) and eval_duration (writing).
from langchain_ollama import ChatOllama
llm = ChatOllama(model="llama3", temperature=0)
reply = llm.invoke("In one sentence: what is a refund?")
print("usage_metadata:", reply.usage_metadata)
m = reply.response_metadata
for key in ["total_duration", "load_duration", "prompt_eval_count", "prompt_eval_duration", "eval_count", "eval_duration"]:
print(f"{key:21} {m.get(key)}")usage_metadata: {'input_tokens': 19, 'output_tokens': 42, 'total_tokens': 61}
total_duration 2276100083
load_duration 2970375
prompt_eval_count 19
prompt_eval_duration 230293000
eval_count 42
eval_duration 2036974000Our first measurements could not be trusted
We wrote a small script that runs a call and prints load, read and write times - and got two impossible results. One call said it spent 156 seconds writing, but the whole call took 2.5 seconds on our stopwatch. Another: 1134 seconds of writing inside a 19-second call.
We checked Ollama directly with curl. Most calls were normal, about 2.1 seconds. One took 607 seconds by the stopwatch, and Ollama reported 607 seconds of writing - yet its own total_duration said 2.51 seconds, and its log shows the request finishing about ten minutes after the previous one. Something outside the request paused it. We did not find out what; the computer may have gone idle.
Two rules came out of this. First, always measure with your own stopwatch next to the reported numbers, and distrust any number that is larger than the stopwatch. Second, measure several times and use the median. For the rest of the lesson we ran each test 2 or 3 times, dropped any run whose numbers did not add up, and kept the Mac awake with caffeinate. After that, no run had to be dropped.
1. cold start (model not loaded) wall 3.30s | load 1.08s | in 19 tok 0.19s | out 42 tok 2.02s (20.8 tok/s)
2. same call, warm wall 2.53s | load 0.00s | in 19 tok 0.05s | out 42 tok 156.67s ( 0.3 tok/s) <- impossible
3. long prompt (handbook x4) wall 5.38s | load 0.00s | in 736 tok 3.76s | out 33 tok 1.61s (20.4 tok/s)
4. long answer (300 words) wall 19.30s | load 0.00s | in 17 tok 0.19s | out 358 tok 1134.55s ( 0.3 tok/s) <- impossible
5. same, num_predict=60 wall 2.98s | load 0.00s | in 17 tok 0.05s | out 60 tok 2.90s (20.7 tok/s)wall 2.14s eval_count 42 eval_duration 2.03s total_duration 2.09s prompt_eval 0.06s
wall 607.63s eval_count 42 eval_duration 607.54s total_duration 2.51s prompt_eval 0.05s
wall 2.11s eval_count 42 eval_duration 2.03s total_duration 2.08s prompt_eval 0.05s
wall 2.11s eval_count 42 eval_duration 2.02s total_duration 2.08s prompt_eval 0.05sExample 2 - measuring properly
The measuring helper turns Ollama’s numbers into seconds and tokens, next to our own stopwatch. measure() runs a call several times, drops runs that do not add up, and prints medians.
import statistics, subprocess, time
from langchain_ollama import ChatOllama
from common import HANDBOOK # the support handbook from Lesson 2.15
def run(model, prompt, **options):
"""One call; return the numbers Ollama reports, in seconds and tokens."""
llm = ChatOllama(model=model, temperature=0, **options)
start = time.perf_counter()
reply = llm.invoke(prompt)
wall = time.perf_counter() - start # our own stopwatch
m = reply.response_metadata
return {
"wall_s": wall,
"load_s": m["load_duration"] / 1e9,
"in_tok": m["prompt_eval_count"], "read_s": m["prompt_eval_duration"] / 1e9,
"out_tok": m["eval_count"], "write_s": m["eval_duration"] / 1e9,
"text": reply.content,
}
def measure(label, model, prompt, repeat=3, **options):
"""Run the same call several times; report the median, and flag numbers we cannot trust."""
runs = [run(model, prompt, **options) for _ in range(repeat)]
bad = sum(r["write_s"] > r["wall_s"] for r in runs) # server time longer than the stopwatch?
good = [r for r in runs if r["write_s"] <= r["wall_s"]] or runs
med = lambda key: statistics.median(r[key] for r in good)
print(f"{label:30} wall {med('wall_s'):5.2f}s | in {med('in_tok'):4.0f} tok read {med('read_s'):4.2f}s | "
f"out {med('out_tok'):4.0f} tok write {med('write_s'):5.2f}s ({med('out_tok') / med('write_s'):4.1f} tok/s)"
+ (f" [{bad} odd run(s) dropped]" if bad else ""))
subprocess.run(["ollama", "stop", "llama3"], capture_output=True) # unload the model: a cold start
cold = run("llama3", "In one sentence: what is a refund?")
print(f"{'cold start, once':30} wall {cold['wall_s']:5.2f}s | load {cold['load_s']:4.2f}s (loading the model into memory)")
q = "In one sentence: what is a refund?"
measure("short prompt, short answer", "llama3", q)
measure("long prompt (handbook x4)", "llama3", HANDBOOK * 4 + "\n\n" + q)
measure("long answer (300 words)", "llama3", "Write 300 words about refunds.", repeat=2)
measure("same, num_predict=60", "llama3", "Write 300 words about refunds.", num_predict=60)cold start, once wall 6.60s | load 4.09s (loading the model into memory)
short prompt, short answer wall 2.37s | in 19 tok read 0.06s | out 42 tok write 2.31s (18.2 tok/s)
long prompt (handbook x4) wall 1.94s | in 736 tok read 0.06s | out 33 tok write 1.84s (18.0 tok/s)
long answer (300 words) wall 19.97s | in 17 tok read 0.12s | out 358 tok write 19.83s (18.1 tok/s)
same, num_predict=60 wall 3.34s | in 17 tok read 0.05s | out 60 tok write 3.28s (18.3 tok/s)Reading the numbers
Writing speed was the same in every test: about 18 tokens per second. So answer time is almost exactly answer length divided by 18. 42 tokens: 2.3 s. 358 tokens: 19.8 s. The single biggest thing you control is how much the model writes.
num_predict is a hard limit on output tokens. With num_predict=60, the "300 words" request stopped at 60 tokens and took 3.3 s instead of 20 s. Be careful: it cuts the answer in the middle of a sentence. Asking for a short answer in the prompt is gentler; num_predict is the safety limit.
The cold start - loading the model into memory - took 1.08 s in one run and 4.09 s in another. It happens on the first call, and again whenever the model was unloaded. Keep a model loaded if users are waiting.
And look at the long prompt: 736 tokens read in 0.06 s. In our first run it took 3.76 s. That difference is the next topic.
A repeated prompt start is almost free
Reading a new long prompt took about 4.2 seconds. Reading the SAME prompt again took 0.06 seconds - about 70 times faster. Ollama keeps the work it did for the start of the last prompt, and reuses it if the next prompt starts the same way. This is called prompt caching. Many hosted providers offer something similar, often at a lower price per cached token.
The most useful case is the last line: the same long start with a new ending read in 0.20 seconds. That is the shape of a real app: the same instructions and documents every time, a different question at the end.
So build prompts with the stable parts first - system instructions, tool descriptions, the documents - and the changing parts last - the user’s question, the latest tool result. Putting something that changes every time (a ticket number, a timestamp) at the very start throws the cache away: our "new" prompts started with a new ticket id, and every one took over 4 seconds.
import uuid
from measure import run
from common import HANDBOOK
q = "\n\nIn one sentence: what is a refund?"
for i in range(3):
fresh = f"Ticket {uuid.uuid4().hex}\n" + HANDBOOK * 4 + q # a new prompt each time
r = run("llama3", fresh)
print(f"new long prompt #{i + 1}: in {r['in_tok']} tok, read {r['read_s']:.2f}s, wall {r['wall_s']:.2f}s")
same = "Ticket 42\n" + HANDBOOK * 4 + q
for i in range(3):
r = run("llama3", same)
print(f"same long prompt again #{i + 1}: in {r['in_tok']} tok, read {r['read_s']:.2f}s, wall {r['wall_s']:.2f}s")
tail = HANDBOOK * 4 + f"\n\nTicket {uuid.uuid4().hex}" + q # same start, new end
r = run("llama3", tail)
print(f"same start, new end: in {r['in_tok']} tok, read {r['read_s']:.2f}s, wall {r['wall_s']:.2f}s")new long prompt #1: in 756 tok, read 4.28s, wall 6.16s
new long prompt #2: in 757 tok, read 4.22s, wall 6.08s
new long prompt #3: in 756 tok, read 4.26s, wall 6.11s
same long prompt again #1: in 740 tok, read 4.01s, wall 5.82s
same long prompt again #2: in 740 tok, read 0.06s, wall 1.87s
same long prompt again #3: in 740 tok, read 0.06s, wall 1.87s
same start, new end: in 754 tok, read 0.20s, wall 2.14sTip: This also explains Lesson 5.1’s trace: the agent’s prompt grows with each step, but its start - the system prompt and tool list - stays the same, so each new step should mostly pay to read only the new part.
Estimating cost
Hosted models charge per token, with one price for input tokens and a higher one for output tokens, usually quoted per million tokens. The cost of one call is: input tokens x input price + output tokens x output price. Look up the current prices for the model you use - they change often.
For an agent, add up every call in a run. Lesson 5.1’s trace showed three model calls per question; each call re-reads the growing transcript. So an agent’s input tokens grow faster than you might think - and that is where cutting calls (like the repeat guard’s extra call) saves the most.
Our numbers make the shape clear. A support answer (the question plus two handbook sections) was about 11 output tokens. The "300 words" request read 17 tokens and wrote 358. Output tokens are slower AND usually more expensive - so short answers save on both.
PRICE_IN_PER_M = 1.00 # example only: price per 1,000,000 input tokens
PRICE_OUT_PER_M = 4.00 # example only: price per 1,000,000 output tokens
def cost(replies):
"""Add up every model call of one run."""
tokens_in = sum(r.usage_metadata["input_tokens"] for r in replies)
tokens_out = sum(r.usage_metadata["output_tokens"] for r in replies)
return tokens_in, tokens_out, tokens_in / 1e6 * PRICE_IN_PER_M + tokens_out / 1e6 * PRICE_OUT_PER_MBenchmark two models - and how ours lied
Is gemma3 (4.3 billion parameters) better than llama3 (8 billion) for our support bot? A benchmark should measure speed AND quality on the same test set. We reused Lesson 5.4: the 14 handbook questions, the same retrieval, the same llama3 judge for both models. Only the model that writes the answer changed.
The first result: same quality (11/14 each), but gemma3 took 5.29 seconds per answer against 0.99 for llama3 - five times slower, for a smaller model. That made no sense, so we checked.
The cause was our benchmark. After each gemma3 answer, the llama3 judge ran. Only one model stays in memory here, so Ollama unloaded gemma3 to run the judge, then loaded gemma3 again for the next answer. We logged gemma3’s load time with the judge in between: 3.7 to 4.7 seconds before every answer. We were measuring model swapping, not gemma3.
The fix: answer all questions with one model first, judge afterwards. Then gemma3 took 0.80 s per answer and llama3 1.07 s - gemma3 is the faster one, with the same score. The first benchmark was wrong by about five times, and would have chosen the wrong model.
import statistics, time
from langchain_ollama import ChatOllama
import rag # the RAG bot from Lesson 5.4
from dataset import DATASET # its 14 cases
from judge import judge # the llama3 judge, the same for both
for model in ["llama3", "gemma3"]:
rag.llm = ChatOllama(model=model, temperature=0) # swap only the model that WRITES the answers
rag.answer("warm-up question", k=2) # load it first, so load time is not counted
walls, loads, answers = [], [], []
for case in DATASET: # 1. ONLY answers - one model in memory
start = time.perf_counter()
docs = rag.store.similarity_search(case["q"], k=2)
reply = rag.llm.invoke(rag.PROMPT.format(context="\n\n".join(d.page_content for d in docs), question=case["q"]))
walls.append(time.perf_counter() - start)
loads.append(reply.response_metadata["load_duration"] / 1e9)
answers.append(reply.content.strip())
passed = sum(judge(c, a) for c, a in zip(DATASET, answers)) # 2. judge afterwards
print(f"{model:7} answers alone: median {statistics.median(walls):4.2f}s, total {sum(walls):4.1f}s, "
f"median load {statistics.median(loads):4.2f}s | judge {passed}/{len(DATASET)}")judge after every answer (bench.py):
llama3 judge 11/14 | median 0.99s per answer | total 14.2s | median 11 output tokens
gemma3 judge 11/14 | median 5.29s per answer | total 69.8s | median 12 output tokens
gemma3 load time per answer, judge in between: [4.42, 3.67, 4.7, 4.44, 4.69]
one model at a time, judge afterwards (bench2.py):
llama3 answers alone: median 1.07s, total 15.0s, median load 0.00s | judge 11/14
gemma3 answers alone: median 0.80s, total 12.0s, median load 0.00s | judge 11/14Watch out: The same trap hits real apps: an agent that uses one model to answer and another to judge or route will swap them on every step if memory only fits one. Check load_duration in your traces - our "slow model" was really "slow swapping".
Making an agent faster: what to try, in order
Use the numbers, not guesses. In our measurements the order of impact was clear.
1. Write less~18 tokens/s: 358 tokens took 20 s, 42 took 2 s. Ask for short answers; set num_predict as a limit.2. Fewer callsEvery call adds 1-2 s. Lesson 5.1 found an extra call per question.3. Stable prompt startA repeated start read in 0.06-0.20 s instead of 4.2 s.4. No cold starts or swapsLoading took 1-4.7 s. Keep models loaded; avoid switching models per step.5. A smaller modelgemma3 answered 25% faster with the same score - measure quality too.6. Stream the answerSame total time, but the user sees the first words sooner (Lesson 3.9).Cost and latency at a glance
Token countsSame names for local and hosted models.
reply.usage_metadata["input_tokens"], ["output_tokens"]
Ollama timingsIn nanoseconds.
reply.response_metadata["eval_duration"] / 1e9
Limit outputHard stop - may cut a sentence.
ChatOllama(model="llama3", num_predict=60)
Cold startUnload, then measure the first call.
ollama stop llama3
Keep awake while measuringmacOS.
caffeinate -i python measure2.py
Cost of a callPer million tokens.
in/1e6*price_in + out/1e6*price_out
Fair benchmarkSame dataset and judge; one model at a time; medians.
Try it yourself
The code does not change. Swap the content string and the program does something else entirely.
“Run measure2.py on your machine. What is your tokens-per-second for writing?”
“Add "Answer in at most 20 words" to the RAG prompt. How much faster is the 5.4 evaluation - and does the judge score change?”
“Put the current time at the start of the RAG prompt. Measure read time before and after.”
“Add usage_metadata counting to the Lesson 5.1 tracer. How many input and output tokens does one question cost?”
“Pull another small model and add it to bench2.py. Is it faster, and is it as good?”
What usually goes wrong
One of our calls "took" 607 seconds. Run each test several times, use the median, and compare with a stopwatch.
Interleaving the judge made gemma3 look 5x slower: it was reloaded before every answer.
✗ for case in cases:
answer = gemma.invoke(...)
judge(case, answer) # llama3 -> swap✓ answers = [gemma.invoke(...) for case in cases]
scores = [judge(c, a) for c, a in zip(cases, answers)]A new ticket id at the start made every 756-token prompt take over 4 s to read. Put changing parts at the end.
Writing (~18 tokens/s) dominated almost every call. Shorter answers save more than shorter prompts.
A faster model that answers worse is not cheaper. Always run the evaluation (Lesson 5.4) in the same benchmark.
Key points
- A call’s time is load + read prompt + write answer; writing (~18 tokens/s here) is usually the biggest part.
- Shorter answers and fewer calls are the biggest levers; num_predict is a hard limit.
- Prompt caching made a repeated 740-token prompt start read in 0.06 s instead of ~4 s - keep stable parts first.
- Cold starts and model swaps cost 1 to 4.7 s each.
- Cost = input tokens x price + output tokens x price, summed over every call of a run.
- Measure several times, use medians, check against a stopwatch, and benchmark one model at a time with the same quality test.
Quick check before you move on
Quiz
- 1.
At 18 tokens per second, about how long does a 360-token answer take to write?
- 2.
What does num_predict=60 do, and what is its downside?
- 3.
Where should a changing ticket number go in a long prompt, and why?
- 4.
After fixing the benchmark, which model would you choose for the support bot?
Interview questions
How do you reduce the latency of an LLM agent?
Measure first with traces. Then cut output length, cut the number of model calls, structure prompts so stable prefixes are cached, avoid cold starts and model swaps, consider a smaller model if evaluation quality holds, and stream output to cut perceived latency.
How do you benchmark two models fairly?
Same dataset, same evaluator, same prompts; warm both models; avoid interleaving other models that cause swaps; repeat runs and use medians; report quality and latency together, and sanity-check reported timings against wall-clock time.
What is prompt caching and how do you design for it?
Reusing computation for a prompt prefix seen before. Put stable content - system prompt, tool definitions, documents - first and variable content last, and avoid volatile values like timestamps at the start.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...