Evaluation Basics
Write a test set of 14 questions for a RAG bot, score every answer with code checks and an LLM judge, and compare two versions. Along the way: a test that passed by luck, a keyword check fooled 5 times out of 5, and a judge that was wrong too.
What you will be able to do
- Explain why "it looked fine in a few tries" is not testing
- Write a small evaluation dataset, including questions the bot cannot answer
- Write evaluators in code: retrieval check, keyword check, "I don’t know" check
- Use an LLM as a judge - and test the judge itself
- Compare two versions of a system on the same dataset
- Spot lucky passes and flaky results, and read failures by hand
The idea, in plain English
You change your agent’s prompt. Is it better now? Most people ask it three questions, read the answers, and decide. That is not testing - three questions miss most problems, and you cannot compare today’s three answers with last week’s.
Evaluation means checking a fixed set of questions every time, scoring each answer the same way, and comparing scores. The set of questions with known good answers is the dataset. The function that scores one answer is the evaluator. Change one thing, run the same dataset, compare - like a school exam for your bot.
In this lesson we evaluate the support-handbook RAG bot (the handbook and stand-in embeddings from Lesson 2.15, with llama3 writing the answer). We start with simple checks in code, find their limits, add an LLM judge, test the judge, and compare two prompts. LangSmith can store datasets and run evaluations for you; we do it locally first so you see every step. Versions: chromadb 1.5.9, langchain-chroma 1.1.0, llama3 on Ollama.
Worked example: The support-handbook RAG bot: 14 questions (10 answerable, 4 not), retrieval checked in code, answers checked by a llama3 judge, prompt v1 vs v2.
1 - First run: a perfect score
With simple code checks and k=2, every case passed. It looks like we are done.
Every case goes through the bot; every answer through the same evaluators. Real results from our runs.
Words you will see in this lesson
A few new words. You will meet them in every evaluation tool, including LangSmith.
EvaluationChecking a system on a fixed set of cases and scoring the results.DatasetThe fixed set of cases: inputs, and what a good output must be.ReferenceThe correct answer for one case, written by you.EvaluatorA function that scores one output: pass/fail, or a number.LLM-as-judgeUsing a model as the evaluator: "Is this answer correct? PASS or FAIL."RegressionSomething that worked before and is broken after a change.FlakyA test that passes or fails by chance, without any change.An everyday example: a driving test
A driving school does not decide you can drive because you drove well around one block. The examiner uses a fixed list: park, reverse, stop at a sign, give way, merge. Every student gets the same list, scored the same way. That is why the scores mean something.
Your bot needs the same: a fixed list of questions (the dataset), scored the same way every time (the evaluators). And like a driving test, it must include hard parts - questions the bot cannot answer and should refuse.
Step 1 - the system we test
Our bot is a small RAG pipeline (Lesson 2.12): search the handbook from Lesson 2.15 for the 2 closest chunks, then ask llama3 to answer using only them - or say "I don’t know".
The answer function returns the answer AND which sections it used. Returning the intermediate step lets us test search and writing separately. If an answer is wrong, we can tell whether search found the wrong text, or llama3 misused the right text.
from langchain_chroma import Chroma
from langchain_ollama import ChatOllama
from langchain_text_splitters import RecursiveCharacterTextSplitter
from common import HANDBOOK, TopicEmbeddings # the handbook and stand-in model from Lesson 2.15
chunks = RecursiveCharacterTextSplitter(chunk_size=200, chunk_overlap=0).split_text(HANDBOOK)
store = Chroma(collection_name="handbook_eval", embedding_function=TopicEmbeddings(),
collection_metadata={"hnsw:space": "cosine"})
store.add_texts(chunks, ids=[f"handbook#{i}" for i in range(len(chunks))],
metadatas=[{"section": c.split(".")[0].lower()} for c in chunks])
llm = ChatOllama(model="llama3", temperature=0)
PROMPT = """Answer the question using ONLY the context.
If the context does not contain the answer, say exactly: I don't know.
Context:
{context}
Question: {question}
Answer in one or two sentences."""
def answer(question: str, k: int = 2) -> dict:
docs = store.similarity_search(question, k=k)
context = "\n\n".join(d.page_content for d in docs)
reply = llm.invoke(PROMPT.format(context=context, question=question)).content.strip()
return {"answer": reply, "sections": [d.metadata["section"] for d in docs]}Step 2 - write the dataset
Each case is a question plus what a good answer must be: the handbook section that holds the answer, a short reference answer, and key words. Write the reference yourself - from the handbook, not from the bot’s output, or you only test that the bot agrees with itself.
Include the hard cases on purpose. Four of our fourteen questions cannot be answered from the handbook: the price of express shipping, shipping to Mars, a phone number, PayPal. For those, the only correct answer is "I don’t know". A bot that invents an answer to them is worse than a bot that refuses.
Start small - 10 to 30 cases you understand well. Add a case every time you find a real bug, so it can never come back unnoticed.
# Each case: a question, the section that holds the answer, and words the answer must contain.
# answerable=False: the handbook does not say - the right answer is "I don't know".
DATASET = [
{"q": "How long do I have to ask for a refund?", "section": "refunds", "reference": "You can ask for a refund within 30 days.", "must": ["30"]},
{"q": "Can I get a refund for a digital item I downloaded?", "section": "refunds", "reference": "No. Refunds for digital items are not possible after download.", "must": ["not"]},
{"q": "How many days does standard shipping take?", "section": "shipping", "reference": "Standard shipping takes 3 to 5 days.", "must": ["3", "5"]},
{"q": "When is shipping free?", "section": "shipping", "reference": "Shipping is free for orders over 50 euros.", "must": ["50"]},
{"q": "My delivery is late. What should I do?", "section": "delivery problems", "reference": "Wait 2 more days, then contact us.", "must": ["2"]},
{"q": "My parcel arrived broken. What should I do?", "section": "delivery problems", "reference": "Send a photo within 48 hours.", "must": ["photo"]},
{"q": "How long must a password be?", "section": "passwords", "reference": "At least 12 characters.", "must": ["12"]},
{"q": "In which format can I download an invoice?", "section": "invoices", "reference": "As a PDF.", "must": ["pdf"]},
{"q": "What does a company invoice need?", "section": "invoices", "reference": "Your VAT number.", "must": ["vat"]},
{"q": "When are my orders deleted after I close my account?", "section": "closing an account", "reference": "90 days after the account is closed.", "must": ["90"]},
{"q": "How much does express shipping cost?", "section": "shipping", "must": [], "answerable": False},
{"q": "Do you ship to Mars?", "section": "shipping", "must": [], "answerable": False},
{"q": "What is your phone number?", "section": None, "must": [], "answerable": False},
{"q": "Can I pay with PayPal?", "section": None, "must": [], "answerable": False},
]Step 3 - evaluators in code
Start with checks that need no model - they are fast, free, and never change their mind. We wrote three. Retrieval: is the expected section among the retrieved ones? Keywords: does the answer contain every "must" word? Refusal: for unanswerable questions, does the answer say it does not know?
The first run with k=2 was perfect: retrieval 12/12, answers 14/14. That should make you suspicious, not happy. So we looked closer.
def says_dont_know(text):
t = text.lower()
return "don't know" in t or "do not know" in t or "not specified" in t or "does not" in t and "mention" in t
def score(case, out):
"""Three simple checks, all in code - no model needed."""
retrieval = None if case["section"] is None else case["section"] in out["sections"]
if case.get("answerable", True):
correct = all(w in out["answer"].lower() for w in case["must"]) and not says_dont_know(out["answer"])
else:
correct = says_dont_know(out["answer"]) # the right answer is "I don't know"
return retrieval, correctok ok | How long do I have to ask for a refund? | You can ask for a refund within 30 days.
ok ok | Can I get a refund for a digital item I downloaded | Refunds for digital items are not possible after download.
ok ok | How many days does standard shipping take? | Standard shipping takes 3 to 5 days.
ok ok | When is shipping free? | Shipping is free for orders over 50 euros.
ok ok | My delivery is late. What should I do? | Wait 2 more days, then contact us.
ok ok | My parcel arrived broken. What should I do? | Send a photo of the broken parcel within 48 hours.
ok ok | How long must a password be? | A password must have at least 12 characters.
ok ok | In which format can I download an invoice? | You can download an invoice as a PDF.
ok ok | What does a company invoice need? | A company invoice needs your VAT number.
ok ok | When are my orders deleted after I close my accoun | According to the context, your orders are deleted after 90 days after closing yo
ok ok | How much does express shipping cost? | According to the context, express shipping takes 1 day and costs extra, but the
ok ok | Do you ship to Mars? | I don't know.
- ok | What is your phone number? | I don't know.
- ok | Can I pay with PayPal? | I don't know.
k=2: retrieval 12/12 answers 14/14 (12 s)A pass that was luck
"My parcel arrived broken" does not contain any of the stand-in model’s topic words (refund, shipping, password, invoice, delivery, account). So its vector is all zeros, and every section scores exactly 0.0. When everything ties, which section comes first is chance.
We ran the same evaluation with k=1 and k=3. With k=1, the parcel question failed - wrong section, and llama3 rightly said "I don’t know". "Do you ship to Mars?" has the same problem: "ship" is not "shipping". And later, with the same code and the same k=2, retrieval scored 10/12 instead of 12/12 - the tie came out differently. A test that changes without any code change is flaky.
The fix for the evaluator: require a real match - the right section AND a score above 0. The fix for the bot (with a real embedding model, Lesson 2.9) would be a model that understands "parcel" and "broken". The evaluation found a weakness the perfect first score hid.
vector: [0.0, 0.0, 0.0, 0.0, 0.0, 0.0]
[(0.0, 'refunds'), (0.0, 'shipping'), (0.0, 'delivery problems'), (0.0, 'passwords'), (0.0, 'invoices'), (0.0, 'closing an account')]
FAIL FAIL | My parcel arrived broken. What should I do? | I don't know.
FAIL ok | Do you ship to Mars? | I don't know.
k=1: retrieval 10/12 answers 13/14 (10 s)
FAIL ok | Do you ship to Mars? | I don't know.
k=3: retrieval 11/12 answers 14/14 (12 s)def retrieval_ok(case, question):
"""Right section in the top 2 - AND a real match (score above 0), not a lucky tie."""
if case["section"] is None:
return None
hits = rag.store.similarity_search_with_relevance_scores(question, k=2)
return any(d.metadata["section"] == case["section"] and s > 0 for d, s in hits)Keyword checks are easy to fool
Keyword checks are cheap, but they only look for words, not meaning. We tried prompt v2 - the same prompt without the "I don’t know" rule. llama3 refused in new words: "I’m not aware of any information about shipping to Mars." That is a fine refusal, but our keyword list did not know "not aware", so it was marked FAIL.
Worse, wrong answers can contain the right words. We wrote five wrong answers that still contain the key word - "You must wait 30 days before you can ask for a refund", "Shipping is free for orders under 50 euros". The keyword check passed all five. Keywords catch missing facts; they cannot catch wrong facts.
Step 4 - an LLM as the judge
An LLM judge reads the question, the reference and the bot’s answer, and replies PASS or FAIL. It understands meaning, so it can accept "I’m not aware of..." and reject "wait 30 days before". But a judge is a model, and models make mistakes. So before you trust a judge, TEST THE JUDGE: give it answers you already know are right or wrong.
Our llama3 judge got 12 of 12 simple cases right. Then the five tricky wrong answers. With a weak reference - only "The answer must contain: 30" - it was fooled twice ("wait 30 days before", "under 50 euros"). With a full reference sentence - "You can ask for a refund within 30 days." - it caught all five. The judge is only as good as the reference you give it.
from langchain_ollama import ChatOllama
judge_llm = ChatOllama(model="llama3", temperature=0)
JUDGE = """You are grading an answer from a support bot.
Question: {question}
Reference: {reference}
Bot answer: {answer}
Is the bot answer correct according to the reference? If the reference says the bot should not know,
the bot must clearly say it does not know or that the information is not available.
Reply with one word: PASS or FAIL."""
def reference_for(case):
if not case.get("answerable", True):
return "The handbook does not contain this information. The bot should say it does not know."
return case.get("reference") or "The answer must contain: " + ", ".join(case["must"])
def judge(case, answer: str) -> bool:
reply = judge_llm.invoke(JUDGE.format(question=case["q"], reference=reference_for(case), answer=answer))
return reply.content.strip().upper().startswith("PASS")right answer -> judge says PASS | You can ask for a refund within 30 days.
WRONG answer -> judge says FAIL | You can ask for a refund within 60 days.
right answer -> judge says PASS | Shipping is free for orders over 50 euros.
WRONG answer -> judge says FAIL | Shipping is always free.
right answer -> judge says PASS | A password must have at least 12 characters.
WRONG answer -> judge says FAIL | A password must have at least 8 characters.
right answer -> judge says PASS | Your orders are deleted 90 days after you close your account.
WRONG answer -> judge says FAIL | Your orders are deleted immediately.
right answer -> judge says PASS | I don't know.
WRONG answer -> judge says FAIL | Yes, we ship to Mars with express shipping.
right answer -> judge says PASS | I'm not aware of any phone number being mentioned.
WRONG answer -> judge says FAIL | You can call us at 040-1234-5678.
judge agreed with the truth on 12/12With a short reference ("The answer must contain: 30"):
keyword check: PASS judge: PASS | You must wait 30 days before you can ask for a refund.
keyword check: PASS judge: PASS | Shipping is free for orders under 50 euros.
keyword check: PASS judge: FAIL | A password can have at most 12 characters.
keyword check: PASS judge: FAIL | Your orders are kept for 90 years after you close your account.
keyword check: PASS judge: FAIL | Wait 2 weeks, then contact us.
With a full reference ("You can ask for a refund within 30 days."):
keyword check: PASS judge: FAIL | You must wait 30 days before you can ask for a refund.
keyword check: PASS judge: FAIL | Shipping is free for orders under 50 euros.
keyword check: PASS judge: FAIL | A password can have at most 12 characters.
keyword check: PASS judge: FAIL | Your orders are kept for 90 years after you close your account.
keyword check: PASS judge: FAIL | Wait 2 weeks, then contact us.Step 5 - compare two versions
Now the real use of evaluation: deciding between two versions. Prompt v1 has the rule "If the context does not contain the answer, say exactly: I don’t know." Prompt v2 does not. Same dataset, same retrieval check, same judge.
v1 scored 12/14 with the judge; v2 scored 10/14. Retrieval was 10/12 for both - with the stricter check, the parcel and Mars questions now fail honestly instead of passing by luck.
v1 (with I-don't-know rule): retrieval 10/12, judge 12/14 (17 s)
retrieval=False judge=PASS | My parcel arrived broken. What should I do? | Send a photo within 48 hours.
retrieval=True judge=FAIL | What does a company invoice need? | A company invoice needs your VAT number.
retrieval=True judge=FAIL | How much does express shipping cost? | According to the context, express shipping takes 1 day and c
retrieval=False judge=PASS | Do you ship to Mars? | I don't know.
v2 (no rule): retrieval 10/12, judge 10/14 (21 s)
retrieval=False judge=PASS | My parcel arrived broken. What should I do? | Send a photo of the damaged parcel within 48 hours to report
retrieval=True judge=FAIL | What does a company invoice need? | A company invoice needs your VAT number.
retrieval=True judge=FAIL | How much does express shipping cost? | According to the context, express shipping takes 1 day, but
retrieval=False judge=PASS | Do you ship to Mars? | We don't have any information about shipping to Mars, as our
retrieval=- judge=FAIL | What is your phone number? | We don't provide a phone number in this context, as the inst
retrieval=- judge=FAIL | Can I pay with PayPal? | No, the context does not mention PayPal as a payment option.Always read the failures yourself
Do not stop at the scores. Read every failure. In v1, "A company invoice needs your VAT number." is correct - the judge failed it anyway. "Express shipping ... costs extra, but the exact cost is not specified" is a fair answer to a question the handbook cannot answer - the judge failed that too. So v1’s real score is 14/14; the judge made two mistakes.
In v2, look at PayPal: "No, the context does not mention PayPal". The second half is true, but "No" is a claim the handbook does not support - maybe the shop does accept PayPal. That is exactly the kind of answer the "I don’t know" rule prevents. So v1 is better - and the evaluation shows it, as long as you read the failures and not only the numbers.
Also notice the parcel question. Retrieval failed (no real match), yet the answer passed - with k=2, the right section came along by chance. That is not something to trust; it is something to fix.
Lucky passParcel and Mars found the right section by tie - flaky. Check scores, not only results.Keyword checksFooled by 5 of 5 wrong answers that contained the right words.LLM judgeNeeds full references: fooled 2/5 with short ones, 0/5 with full ones.Judge mistakesFailed 2 correct answers in v1. Read every failure by hand.DecisionKeep prompt v1: it refuses cleanly; v2 made an unsupported claim.Which evaluator when?
Use code checks wherever they can work, and an LLM judge only where meaning matters. Code is free, fast and stable. A judge costs a model call per case, can be wrong, and must itself be tested. A good evaluation usually has both.
Was the right document found?Code: section in results, with a real score.Is a required fact present?Code: keyword or regex - but it cannot catch wrong facts.Valid JSON, length, format?Code (Lesson 5.5).Is the answer correct in meaning?LLM judge with a full reference - tested first.Did it refuse when it should?LLM judge, or code with a list of refusal phrases.Evaluation at a glance
Dataset caseInput + what good looks like.
{"q": ..., "section": ..., "reference": ..., "answerable": False}Code evaluatorFast, free, stable.
def score(case, out) -> bool
Retrieval checkRight source AND a real match.
section matches and score > 0
LLM judgeMeaning, not words - test it first.
judge(case, answer) -> PASS / FAIL
Compare versionsSame dataset, same evaluators.
for label, prompt in PROMPTS.items(): ...
Read failuresScores lie; failures explain.
Try it yourself
The code does not change. Swap the content string and the program does something else entirely.
“Add three questions of your own, including one the handbook cannot answer. Do they pass?”
“Change "30 days" to "14 days" in the handbook but not in the dataset. Which evaluator notices?”
“Write five more tricky wrong answers. Does the judge with full references still catch them?”
“Run run_eval.py three times without changes. Which cases change? Those are your flaky tests.”
“Compare k=1, 2 and 3 with run_eval.py. Which k gives the best judge score?”
What usually goes wrong
Three questions miss most problems, and you cannot compare today’s answers with last week’s. Use a fixed dataset.
Our first run was 14/14; one pass was luck from tied scores. Look at how each case passed.
They passed five wrong answers that contained the right words.
Test the judge on answers you know are right and wrong. Give it full references - short ones fooled it twice.
✗ reference = "The answer must contain: 30"✓ reference = "You can ask for a refund within 30 days."Two of v1’s failures were judge mistakes; one of v2’s "passes" made an unsupported claim. Read every failure.
A bot that invents answers looks great on answerable questions. Include cases where the right answer is "I don’t know".
Key points
- Evaluation = a fixed dataset + evaluators, run after every change, and compared.
- Write references yourself; include questions the bot must refuse.
- Return intermediate steps (like retrieved sections) so you can test each part.
- Code checks are fast and stable but match words, not meaning.
- An LLM judge understands meaning - test it first, and give it full references.
- Watch for lucky passes and flaky cases; check scores, not only results.
- Always read the failures: the judge can be wrong too.
Quick check before you move on
Quiz
- 1.
Prompt v1 scored 12/14 with the judge. What was its real score, and why?
- 2.
Why is "No, the context does not mention PayPal" a bad answer?
- 3.
Retrieval scored 12/12 in one run and 10/12 in another with the same code. Why?
- 4.
Why return the retrieved sections from the answer function?
Interview questions
How would you evaluate a RAG system?
Build a dataset of questions with references, including unanswerable ones. Evaluate retrieval (right source, real match), faithfulness (answer supported by context), correctness against the reference, and refusal behaviour - with code checks where possible and a validated LLM judge for meaning. Run it on every change and compare versions.
What are the pitfalls of LLM-as-judge?
Judges can be wrong, inconsistent, and fooled by answers that look similar to the reference; quality depends on the reference and rubric. Calibrate the judge on labelled examples, use full references, keep temperature low, and spot-check failures by hand.
How do you prevent regressions in an LLM app?
Keep a growing regression dataset - add every real bug as a case - and run the evaluation in CI on prompt, model or code changes, comparing scores and reviewing new failures before shipping.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...