← Back to Agentic AI map
Lesson 2.12 · Building Agents with LangChain

Putting It Together: RAG

Combine embeddings, a vector store and a retriever into one chain that answers questions about a PDF - and see where it fails.

rag

What you will be able to do

  • Explain what RAG is and the problem it solves
  • Separate the two phases: indexing and answering
  • Load a PDF, split it into chunks, and choose chunk_size and chunk_overlap
  • Build the RAG chain with LCEL: retriever, prompt, LLM, parser
  • Write a prompt that says "I don’t know" when the context lacks the answer
  • Find where a RAG system fails - usually retrieval

The idea, in plain English

A model knows what it was trained on, not your documents. Asked how much annual leave "our company" gives, llama3 said it has no access to your company’s policies. RAG - Retrieval-Augmented Generation - fixes that without changing the model: retrieve the relevant passages, add them to the prompt, and let the model generate an answer from them.

It has two phases. Indexing happens once: load the documents, split them into chunks, embed the chunks, store them (Lessons 2.9 and 2.10). Answering happens per question: the retriever finds the closest chunks (2.11), they are formatted into a prompt with the question, and the LLM answers (2.3 and 2.4).

The whole answering phase is one LCEL chain: {"context": retriever | format_documents, "question": passthrough} | prompt | llm | StrOutputParser(). Ask it a string, get a string back.

Generation in this lesson is real llama3 output. Retrieval needs an embedding model, which this machine does not have (Lesson 2.9) - so the runs use a keyword retriever, BM25, in Chroma’s place. Because a retriever is just question in, Documents out, the rest of the chain is unchanged. The one place keyword search behaves differently turned out to be the most instructive result of the lesson.

Worked example: Ask questions about a PDF.

request flowAsk questions about a PDFstep 1 / 4

1 - Indexing: load and split

PyPDFLoader returns one Document per page, with source and page in the metadata. The splitter cuts the pages into chunks of at most 200 characters; each chunk keeps its page number.

pages
2
chunks
5
chunk_size / overlap
200 / 30
metadata kept
source, page

A two-page employee handbook, indexed once, then asked "How many sick days do I get?". Real runs; retrieval by BM25 in place of Chroma.

Why RAG

Without RAG, the model has its training data and your question. Asked "How many days of annual leave do employees at our company receive?", llama3 said it does not have access to your company’s policies and offered general information instead. Honest - and useless. Another model, or another question, might produce a confident invented number.

RAG suits information that is private (policies, customer records), changes often (prices, documentation), is too large for one prompt (a hundred PDFs), or is specific to your domain (course material). The documents stay in your store; the model sees only the few chunks relevant to each question.

Two phases

Indexing: load, split, embed, store. It runs once, and again when documents change - the expensive part, kept out of the question path. Answering: retrieve, augment, generate. It runs per question and touches only the top k chunks.

Keep them as separate scripts or functions. An app that re-indexes every document on every question is a common first mistake - and with persist_directory (Lesson 2.10) there is no need.

Chunking

A whole document is too large to embed as one vector and too large to put in every prompt. RecursiveCharacterTextSplitter cuts text into chunks of at most chunk_size characters, preferring paragraph breaks, then line breaks, then spaces - so it rarely cuts mid-word. Our four one-sentence policies stayed four chunks at chunk_size=500.

chunk_overlap repeats the end of one chunk at the start of the next. On a three-sentence leave policy with chunk_size=100: without overlap, one chunk ended "Unused annual leave of up to 5 days can" and the next began "be carried over to the next year." - the rule split in half. With chunk_overlap=30, the third chunk carried "within 2 working days. Unused annual leave of up to 5 days can be carried over to the next year." whole.

Splitting the leave policy, chunk_size=100 - checked
overlap 0, chunk 2Managers approve or reject requests within 2 working days. Unused annual leave of up to 5 days can
overlap 0, chunk 3be carried over to the next year.
overlap 30, chunk 2days before the leave starts. Managers approve or reject requests within 2 working days. Unused
overlap 30, chunk 3within 2 working days. Unused annual leave of up to 5 days can be carried over to the next year.

Loading a PDF

PyPDFLoader (pip install langchain-community pypdf) returns one Document per page, with metadata including source, page (counting from 0) and page_label. split_documents() - not create_documents() - splits those Documents and copies the metadata onto every chunk, so each chunk still knows its page.

Our handbook: 2 pages, 5 chunks of 127 to 190 characters. Check what extraction produced before anything else - scanned PDFs, tables and two-column layouts often come out as garbled or empty text, and no later step can recover it. Note: importing from langchain-community printed a notice that the package is being sunset in favour of standalone integration packages; the loader still worked.

The RAG chain

The chain starts with a dict of runnables, run in parallel on the question: "context" is retriever | format_documents - Documents in, one string out - and "question" passes the question through unchanged (lambda x: x or RunnablePassthrough(), both gave identical answers). The dict feeds the prompt, then the model, then StrOutputParser.

format_documents decides what the model sees. Joining page_content with blank lines works; prefixing each chunk with its page - "[page 1] Sick leave..." - lets the answer cite its source.

"If it is not in the context, say you don’t know"

For "What is the company’s maternity leave policy?" the retriever still returned three chunks - annual leave, travel, laptops - because top-k always returns something. With the instruction, llama3 answered "I don’t know." Asked about bringing a dog to the office: "I don’t know."

We also removed the instruction, to see if llama3 would invent a policy. It did not: it said the context did not mention maternity leave, or pets. That is this model on these questions at temperature 0 - not a guarantee. Keep the instruction; it costs nothing and it is the behaviour you want specified, not hoped for.

Where RAG fails: usually retrieval

"How much vacation do I get?" The handbook says 20 days of annual leave. Our keyword retriever returned the laptop and remote-work chunks - "vacation" appears nowhere in the PDF - and llama3, correctly given what it saw, said "I don’t know." The answer was in the document; it never reached the prompt.

This is exactly the gap embeddings close: "vacation" and "annual leave" mean the same thing, and a semantic retriever would match them where keyword search cannot. It is also the general lesson. Most bad RAG answers are retrieval failures, and the model cannot answer from a chunk it was never shown - so when an answer is wrong, print what was retrieved first.

Where it can go wrong
ExtractionThe PDF’s text comes out empty or garbled.
ChunkingA rule is split across chunks, or chunks are too big to be specific.
EmbeddingThe model does not place question and answer close together.
RetrievalThe right chunk is not in the top k - "vacation" vs "annual leave".
PromptNo instruction for missing answers; context formatted badly.
GenerationThe model misreads the context or adds to it.

RAG is not training

Indexing a PDF changes the vector store, not the model. The model receives the chunks for one call and keeps nothing. Update a policy and you re-index the document; the next answer uses the new text.

Fine-tuning is different: it changes the model’s weights, needs a training run, and suits style or task behaviour more than facts that change. For questions over documents, RAG is the natural fit.

RAG or fine-tuning
What changesRAG: the document index. Fine-tuning: the model’s weights.
New informationRAG: re-index. Fine-tuning: train again.
Good forRAG: questions over documents. Fine-tuning: style, format, task behaviour.
CitationsRAG: the retrieved chunks are the sources. Fine-tuning: none.

RAG on a learning platform

Index lesson content with course_id, lesson_id - and, for many academies on one platform, tenant_id - in the metadata. "Ask this course" is then a retriever with filter={"tenant_id": ..., "course_id": ...}; "ask this lesson" narrows the filter. Quiz generation and module summaries are the same chain with a different prompt.

Step-by-step code

Phase 1 - index the documents
from langchain_chroma import Chroma from langchain_ollama import OllamaEmbeddings from langchain_text_splitters import RecursiveCharacterTextSplitter documents = [ "Employees receive 20 days of annual leave per year.", "Employees can work remotely up to three days per week.", "All company laptops must use disk encryption.", "Travel expenses must be submitted within 30 days.", ] splitter = RecursiveCharacterTextSplitter(chunk_size=500, chunk_overlap=50) chunks = splitter.create_documents(documents) # 4 Documents - each sentence fits in one chunk vector_store = Chroma( collection_name="company_docs", embedding_function=OllamaEmbeddings(model="nomic-embed-text"), ) vector_store.add_documents(chunks) retriever = vector_store.as_retriever(search_kwargs={"k": 3})
Phase 2 - the RAG chain
from langchain_core.output_parsers import StrOutputParser from langchain_core.prompts import ChatPromptTemplate from langchain_ollama import ChatOllama prompt = ChatPromptTemplate.from_template(""" Answer the question using only the context below. Context: {context} Question: {question} If the answer is not present in the context, say that you don't know. """) def format_documents(documents): return "\n\n".join(document.page_content for document in documents) llm = ChatOllama(model="llama3.1", temperature=0) chain = ( {"context": retriever | format_documents, "question": lambda x: x} | prompt | llm | StrOutputParser() ) print(chain.invoke("How many annual leave days do employees receive?"))
Output - llama3, BM25 retriever in place of Chroma
"How many annual leave days do employees receive?" retrieved Employees receive 20 days of annual leave... / Travel expenses... / All company laptops... answer According to the context, employees receive 20 days of annual leave per year. "How often can I work from home?" answer According to the context, you can work from home up to three days per week. "What is the company's maternity leave policy?" retrieved annual leave, travel expenses, laptops - top-k always returns something answer I don't know. The context only mentions annual leave, travel expenses, and disk encryption, but does not mention maternity leave. Without RAG - llama3 alone: "How many days of annual leave do employees at our company receive?" I'm happy to help! However, I'm a large language model, I don't have access to specific information about your company's policies. ...
Ask questions about a PDF
# pip install langchain-community pypdf from langchain_community.document_loaders import PyPDFLoader pages = PyPDFLoader("employee_handbook.pdf").load() # one Document per page print(len(pages), pages[0].metadata["page"]) # 2 0 chunks = RecursiveCharacterTextSplitter(chunk_size=200, chunk_overlap=30).split_documents(pages) print(len(chunks)) # 5 - each keeps its page metadata vector_store = Chroma(collection_name="handbook", embedding_function=OllamaEmbeddings(model="nomic-embed-text")) vector_store.add_documents(chunks) retriever = vector_store.as_retriever(search_kwargs={"k": 2}) def format_with_pages(documents): return "\n\n".join(f"[page {d.metadata['page'] + 1}] {d.page_content}" for d in documents) chain = {"context": retriever | format_with_pages, "question": lambda x: x} | prompt | llm | StrOutputParser()
Output - the handbook, llama3, BM25 retriever
How many days before my leave do I need to submit a request? retrieved page 1: Requesting leave... / page 1: Annual leave... answer you need to submit a request at least 7 days before your leave starts. How many sick days do I get? retrieved page 1: Sick leave... / page 1: Annual leave... answer you get 10 days of paid sick leave per year. Do I need a doctor's note? answer you need a doctor's note if your absence is longer than 3 days. How much vacation do I get? retrieved page 2: Laptops... / page 2: Remote work... <- the leave chunk was missed answer I don't know. The context only provides information about company laptops, disk encryption, and travel expenses, but does not mention vacation time. Can I bring my dog to the office? answer I don't know.
Testing the chain without an embedding model
# pip install rank_bm25 from langchain_community.retrievers import BM25Retriever retriever = BM25Retriever.from_documents(chunks, k=2) # keyword search - question in, Documents out # The chain above works unchanged: a retriever is a retriever. # Keyword search matches words, not meaning - "vacation" never finds "annual leave".

Tip: When an answer is wrong, print retriever.invoke(question) before touching the prompt. If the right chunk is not there, the model never had a chance.

Watch out: The "say you don’t know" instruction reduces invented answers; it does not prevent them. Test questions your documents cannot answer, and keep testing when you change the model.

RAG at a glance

PyPDFLoader

One Document per page, with page metadata.

PyPDFLoader("file.pdf").load()
RecursiveCharacterTextSplitter

Split text into chunks at natural boundaries.

RecursiveCharacterTextSplitter(chunk_size=500, chunk_overlap=50)
create_documents

Split strings into Documents.

splitter.create_documents(texts)
split_documents

Split Documents, keeping metadata.

splitter.split_documents(pages)
add_documents

Embed and store chunks.

vector_store.add_documents(chunks)
The chain

Retrieve, augment, generate.

{"context": retriever | fmt, "question": ...} | prompt | llm | parser

Try it yourself

The code does not change. Swap the content string and the program does something else entirely.

Quick exercise

“Index "Python is used for backend development.", "React is used for building user interfaces." and "PostgreSQL is a relational database.", then ask "What technology is used to build user interfaces?"”

Missing answer

“Ask the policy chain about maternity leave, with and without the "say you don’t know" line.”

Overlap

“Split a paragraph with chunk_size=100 and overlap 0, then 30. Find a rule that gets cut in half.”

Cite pages

“Load a PDF of your own and format the context with page numbers. Ask the model to cite them.”

What usually goes wrong

Re-indexing on every question

Indexing is a separate phase. Persist the store and only re-index when documents change.

Losing the page numbers

create_documents on extracted text drops the metadata. Split the loaded Documents.

✗ splitter.create_documents([p.page_content for p in pages])
✓ splitter.split_documents(pages)
Debugging the prompt when retrieval failed

Our "vacation" question failed in retrieval; no prompt change could fix it. Look at the retrieved chunks first.

No instruction for missing answers

Top-k always returns something. Tell the model what to do when the answer is not in it.

Expecting RAG to update the model

The model learns nothing. Change the documents and re-index.

Key points

  • RAG: retrieve relevant chunks, add them to the prompt, generate the answer.
  • Indexing (load, split, embed, store) runs once; answering runs per question.
  • chunk_overlap keeps rules that cross a boundary readable.
  • The chain: {"context": retriever | format, "question": passthrough} | prompt | llm | parser.
  • Tell the model to say it does not know when the context lacks the answer.
  • Most failures are retrieval failures - print what was retrieved.
  • RAG does not train the model; re-index to update knowledge.

Quick check before you move on

What is RAG?
Retrieval-Augmented Generation: retrieve relevant information, add it to the prompt, and let the LLM generate the answer from it.
What are the two phases?
Indexing (load, split, embed, store) and answering (retrieve, augment, generate).
Why chunk documents?
So retrieval can find the specific passage, and the prompt gets focused context instead of whole documents.
What is chunk overlap for?
Repeating a little text between neighbouring chunks so information crossing a boundary is not cut in half.
Does RAG train the model?
No. The chunks are given as context for each call; the model is unchanged.
Can RAG still give wrong answers?
Yes - bad extraction, chunking or retrieval, or a model that misreads the context.

Quiz

  1. 1.

    "How much vacation do I get?" got "I don’t know" although the handbook lists 20 days of annual leave. Which step failed?

  2. 2.

    For the maternity-leave question, the retriever returned three chunks. Why, if none was relevant?

  3. 3.

    Why split_documents(pages) rather than create_documents(texts)?

  4. 4.

    What does "question": lambda x: x do in the chain?

Interview questions

What is RAG?

A pattern that retrieves relevant information from an external source at query time and gives it to an LLM as context for its answer.

What are the main steps?

Load and chunk documents, embed and store them; at query time retrieve the top chunks, build a prompt with them and the question, and generate.

RAG or fine-tuning for questions over company documents?

RAG: knowledge stays in an index you can update and cite. Fine-tuning changes weights and suits behaviour, not frequently changing facts.

How do you debug a wrong RAG answer?

Check retrieval first - were the right chunks returned? Then extraction and chunking, then the prompt, and only then the model.

Can RAG still hallucinate?

Yes. Instructions to stay within the context reduce it; evaluation on questions the documents cannot answer is how you find out.

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...