Putting It Together: RAG
Combine embeddings, a vector store and a retriever into one chain that answers questions about a PDF - and see where it fails.
What you will be able to do
- Explain what RAG is and the problem it solves
- Separate the two phases: indexing and answering
- Load a PDF, split it into chunks, and choose chunk_size and chunk_overlap
- Build the RAG chain with LCEL: retriever, prompt, LLM, parser
- Write a prompt that says "I don’t know" when the context lacks the answer
- Find where a RAG system fails - usually retrieval
The idea, in plain English
A model knows what it was trained on, not your documents. Asked how much annual leave "our company" gives, llama3 said it has no access to your company’s policies. RAG - Retrieval-Augmented Generation - fixes that without changing the model: retrieve the relevant passages, add them to the prompt, and let the model generate an answer from them.
It has two phases. Indexing happens once: load the documents, split them into chunks, embed the chunks, store them (Lessons 2.9 and 2.10). Answering happens per question: the retriever finds the closest chunks (2.11), they are formatted into a prompt with the question, and the LLM answers (2.3 and 2.4).
The whole answering phase is one LCEL chain: {"context": retriever | format_documents, "question": passthrough} | prompt | llm | StrOutputParser(). Ask it a string, get a string back.
Generation in this lesson is real llama3 output. Retrieval needs an embedding model, which this machine does not have (Lesson 2.9) - so the runs use a keyword retriever, BM25, in Chroma’s place. Because a retriever is just question in, Documents out, the rest of the chain is unchanged. The one place keyword search behaves differently turned out to be the most instructive result of the lesson.
Worked example: Ask questions about a PDF.
1 - Indexing: load and split
PyPDFLoader returns one Document per page, with source and page in the metadata. The splitter cuts the pages into chunks of at most 200 characters; each chunk keeps its page number.
A two-page employee handbook, indexed once, then asked "How many sick days do I get?". Real runs; retrieval by BM25 in place of Chroma.
Why RAG
Without RAG, the model has its training data and your question. Asked "How many days of annual leave do employees at our company receive?", llama3 said it does not have access to your company’s policies and offered general information instead. Honest - and useless. Another model, or another question, might produce a confident invented number.
RAG suits information that is private (policies, customer records), changes often (prices, documentation), is too large for one prompt (a hundred PDFs), or is specific to your domain (course material). The documents stay in your store; the model sees only the few chunks relevant to each question.
Two phases
Indexing: load, split, embed, store. It runs once, and again when documents change - the expensive part, kept out of the question path. Answering: retrieve, augment, generate. It runs per question and touches only the top k chunks.
Keep them as separate scripts or functions. An app that re-indexes every document on every question is a common first mistake - and with persist_directory (Lesson 2.10) there is no need.
Chunking
A whole document is too large to embed as one vector and too large to put in every prompt. RecursiveCharacterTextSplitter cuts text into chunks of at most chunk_size characters, preferring paragraph breaks, then line breaks, then spaces - so it rarely cuts mid-word. Our four one-sentence policies stayed four chunks at chunk_size=500.
chunk_overlap repeats the end of one chunk at the start of the next. On a three-sentence leave policy with chunk_size=100: without overlap, one chunk ended "Unused annual leave of up to 5 days can" and the next began "be carried over to the next year." - the rule split in half. With chunk_overlap=30, the third chunk carried "within 2 working days. Unused annual leave of up to 5 days can be carried over to the next year." whole.
overlap 0, chunk 2Managers approve or reject requests within 2 working days. Unused annual leave of up to 5 days canoverlap 0, chunk 3be carried over to the next year.overlap 30, chunk 2days before the leave starts. Managers approve or reject requests within 2 working days. Unusedoverlap 30, chunk 3within 2 working days. Unused annual leave of up to 5 days can be carried over to the next year.Loading a PDF
PyPDFLoader (pip install langchain-community pypdf) returns one Document per page, with metadata including source, page (counting from 0) and page_label. split_documents() - not create_documents() - splits those Documents and copies the metadata onto every chunk, so each chunk still knows its page.
Our handbook: 2 pages, 5 chunks of 127 to 190 characters. Check what extraction produced before anything else - scanned PDFs, tables and two-column layouts often come out as garbled or empty text, and no later step can recover it. Note: importing from langchain-community printed a notice that the package is being sunset in favour of standalone integration packages; the loader still worked.
The RAG chain
The chain starts with a dict of runnables, run in parallel on the question: "context" is retriever | format_documents - Documents in, one string out - and "question" passes the question through unchanged (lambda x: x or RunnablePassthrough(), both gave identical answers). The dict feeds the prompt, then the model, then StrOutputParser.
format_documents decides what the model sees. Joining page_content with blank lines works; prefixing each chunk with its page - "[page 1] Sick leave..." - lets the answer cite its source.
"If it is not in the context, say you don’t know"
For "What is the company’s maternity leave policy?" the retriever still returned three chunks - annual leave, travel, laptops - because top-k always returns something. With the instruction, llama3 answered "I don’t know." Asked about bringing a dog to the office: "I don’t know."
We also removed the instruction, to see if llama3 would invent a policy. It did not: it said the context did not mention maternity leave, or pets. That is this model on these questions at temperature 0 - not a guarantee. Keep the instruction; it costs nothing and it is the behaviour you want specified, not hoped for.
Where RAG fails: usually retrieval
"How much vacation do I get?" The handbook says 20 days of annual leave. Our keyword retriever returned the laptop and remote-work chunks - "vacation" appears nowhere in the PDF - and llama3, correctly given what it saw, said "I don’t know." The answer was in the document; it never reached the prompt.
This is exactly the gap embeddings close: "vacation" and "annual leave" mean the same thing, and a semantic retriever would match them where keyword search cannot. It is also the general lesson. Most bad RAG answers are retrieval failures, and the model cannot answer from a chunk it was never shown - so when an answer is wrong, print what was retrieved first.
ExtractionThe PDF’s text comes out empty or garbled.ChunkingA rule is split across chunks, or chunks are too big to be specific.EmbeddingThe model does not place question and answer close together.RetrievalThe right chunk is not in the top k - "vacation" vs "annual leave".PromptNo instruction for missing answers; context formatted badly.GenerationThe model misreads the context or adds to it.RAG is not training
Indexing a PDF changes the vector store, not the model. The model receives the chunks for one call and keeps nothing. Update a policy and you re-index the document; the next answer uses the new text.
Fine-tuning is different: it changes the model’s weights, needs a training run, and suits style or task behaviour more than facts that change. For questions over documents, RAG is the natural fit.
What changesRAG: the document index. Fine-tuning: the model’s weights.New informationRAG: re-index. Fine-tuning: train again.Good forRAG: questions over documents. Fine-tuning: style, format, task behaviour.CitationsRAG: the retrieved chunks are the sources. Fine-tuning: none.RAG on a learning platform
Index lesson content with course_id, lesson_id - and, for many academies on one platform, tenant_id - in the metadata. "Ask this course" is then a retriever with filter={"tenant_id": ..., "course_id": ...}; "ask this lesson" narrows the filter. Quiz generation and module summaries are the same chain with a different prompt.
Step-by-step code
from langchain_chroma import Chroma
from langchain_ollama import OllamaEmbeddings
from langchain_text_splitters import RecursiveCharacterTextSplitter
documents = [
"Employees receive 20 days of annual leave per year.",
"Employees can work remotely up to three days per week.",
"All company laptops must use disk encryption.",
"Travel expenses must be submitted within 30 days.",
]
splitter = RecursiveCharacterTextSplitter(chunk_size=500, chunk_overlap=50)
chunks = splitter.create_documents(documents) # 4 Documents - each sentence fits in one chunk
vector_store = Chroma(
collection_name="company_docs",
embedding_function=OllamaEmbeddings(model="nomic-embed-text"),
)
vector_store.add_documents(chunks)
retriever = vector_store.as_retriever(search_kwargs={"k": 3})from langchain_core.output_parsers import StrOutputParser
from langchain_core.prompts import ChatPromptTemplate
from langchain_ollama import ChatOllama
prompt = ChatPromptTemplate.from_template("""
Answer the question using only the context below.
Context:
{context}
Question:
{question}
If the answer is not present in the context,
say that you don't know.
""")
def format_documents(documents):
return "\n\n".join(document.page_content for document in documents)
llm = ChatOllama(model="llama3.1", temperature=0)
chain = (
{"context": retriever | format_documents, "question": lambda x: x}
| prompt
| llm
| StrOutputParser()
)
print(chain.invoke("How many annual leave days do employees receive?"))"How many annual leave days do employees receive?"
retrieved Employees receive 20 days of annual leave... / Travel expenses... / All company laptops...
answer According to the context, employees receive 20 days of annual leave per year.
"How often can I work from home?"
answer According to the context, you can work from home up to three days per week.
"What is the company's maternity leave policy?"
retrieved annual leave, travel expenses, laptops - top-k always returns something
answer I don't know. The context only mentions annual leave, travel expenses, and disk
encryption, but does not mention maternity leave.
Without RAG - llama3 alone:
"How many days of annual leave do employees at our company receive?"
I'm happy to help! However, I'm a large language model, I don't have access to specific
information about your company's policies. ...# pip install langchain-community pypdf
from langchain_community.document_loaders import PyPDFLoader
pages = PyPDFLoader("employee_handbook.pdf").load() # one Document per page
print(len(pages), pages[0].metadata["page"]) # 2 0
chunks = RecursiveCharacterTextSplitter(chunk_size=200, chunk_overlap=30).split_documents(pages)
print(len(chunks)) # 5 - each keeps its page metadata
vector_store = Chroma(collection_name="handbook", embedding_function=OllamaEmbeddings(model="nomic-embed-text"))
vector_store.add_documents(chunks)
retriever = vector_store.as_retriever(search_kwargs={"k": 2})
def format_with_pages(documents):
return "\n\n".join(f"[page {d.metadata['page'] + 1}] {d.page_content}" for d in documents)
chain = {"context": retriever | format_with_pages, "question": lambda x: x} | prompt | llm | StrOutputParser()How many days before my leave do I need to submit a request?
retrieved page 1: Requesting leave... / page 1: Annual leave...
answer you need to submit a request at least 7 days before your leave starts.
How many sick days do I get?
retrieved page 1: Sick leave... / page 1: Annual leave...
answer you get 10 days of paid sick leave per year.
Do I need a doctor's note?
answer you need a doctor's note if your absence is longer than 3 days.
How much vacation do I get?
retrieved page 2: Laptops... / page 2: Remote work... <- the leave chunk was missed
answer I don't know. The context only provides information about company laptops,
disk encryption, and travel expenses, but does not mention vacation time.
Can I bring my dog to the office?
answer I don't know.# pip install rank_bm25
from langchain_community.retrievers import BM25Retriever
retriever = BM25Retriever.from_documents(chunks, k=2) # keyword search - question in, Documents out
# The chain above works unchanged: a retriever is a retriever.
# Keyword search matches words, not meaning - "vacation" never finds "annual leave".Tip: When an answer is wrong, print retriever.invoke(question) before touching the prompt. If the right chunk is not there, the model never had a chance.
Watch out: The "say you don’t know" instruction reduces invented answers; it does not prevent them. Test questions your documents cannot answer, and keep testing when you change the model.
RAG at a glance
PyPDFLoaderOne Document per page, with page metadata.
PyPDFLoader("file.pdf").load()RecursiveCharacterTextSplitterSplit text into chunks at natural boundaries.
RecursiveCharacterTextSplitter(chunk_size=500, chunk_overlap=50)
create_documentsSplit strings into Documents.
splitter.create_documents(texts)
split_documentsSplit Documents, keeping metadata.
splitter.split_documents(pages)
add_documentsEmbed and store chunks.
vector_store.add_documents(chunks)
The chainRetrieve, augment, generate.
{"context": retriever | fmt, "question": ...} | prompt | llm | parserTry it yourself
The code does not change. Swap the content string and the program does something else entirely.
“Index "Python is used for backend development.", "React is used for building user interfaces." and "PostgreSQL is a relational database.", then ask "What technology is used to build user interfaces?"”
“Ask the policy chain about maternity leave, with and without the "say you don’t know" line.”
“Split a paragraph with chunk_size=100 and overlap 0, then 30. Find a rule that gets cut in half.”
“Load a PDF of your own and format the context with page numbers. Ask the model to cite them.”
What usually goes wrong
Indexing is a separate phase. Persist the store and only re-index when documents change.
create_documents on extracted text drops the metadata. Split the loaded Documents.
✗ splitter.create_documents([p.page_content for p in pages])✓ splitter.split_documents(pages)Our "vacation" question failed in retrieval; no prompt change could fix it. Look at the retrieved chunks first.
Top-k always returns something. Tell the model what to do when the answer is not in it.
The model learns nothing. Change the documents and re-index.
Key points
- RAG: retrieve relevant chunks, add them to the prompt, generate the answer.
- Indexing (load, split, embed, store) runs once; answering runs per question.
- chunk_overlap keeps rules that cross a boundary readable.
- The chain: {"context": retriever | format, "question": passthrough} | prompt | llm | parser.
- Tell the model to say it does not know when the context lacks the answer.
- Most failures are retrieval failures - print what was retrieved.
- RAG does not train the model; re-index to update knowledge.
Quick check before you move on
Quiz
- 1.
"How much vacation do I get?" got "I don’t know" although the handbook lists 20 days of annual leave. Which step failed?
- 2.
For the maternity-leave question, the retriever returned three chunks. Why, if none was relevant?
- 3.
Why split_documents(pages) rather than create_documents(texts)?
- 4.
What does "question": lambda x: x do in the chain?
Interview questions
What is RAG?
A pattern that retrieves relevant information from an external source at query time and gives it to an LLM as context for its answer.
What are the main steps?
Load and chunk documents, embed and store them; at query time retrieve the top chunks, build a prompt with them and the question, and generate.
RAG or fine-tuning for questions over company documents?
RAG: knowledge stays in an index you can update and cite. Fine-tuning changes weights and suits behaviour, not frequently changing facts.
How do you debug a wrong RAG answer?
Check retrieval first - were the right chunks returned? Then extraction and chunking, then the prompt, and only then the model.
Can RAG still hallucinate?
Yes. Instructions to stay within the context reduce it; evaluation on questions the documents cannot answer is how you find out.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...
AI
System Design
Backend
- GraphQL8 modules · 69 lessons planned
- Core Python13 modules · 75 lessons planned
- FastAPI5 sections · 20 lessons
- Node.js14 modules · 206 lessons planned
- Node.js Performance7 chapters · 36 topics
- Event Loop Lifecycle6 phases · 3 scenarios
- Docker & Containerization11 modules · 144 lessons planned
- AWS for Developers14 modules · 219 lessons planned
- CI/CD & DevOps Automation10 modules · 134 lessons planned