Retrievers
Turn the vector store into a retriever, ask for the top k, and see why MMR can return better context than plain similarity.
What you will be able to do
- Explain the difference between a vector store and a retriever
- Create a retriever with as_retriever() and call it with invoke()
- Choose k, and know the default
- Compare similarity search and MMR on the same question
- Tune MMR with fetch_k and lambda_mult
- Place the retriever inside RAG
The idea, in plain English
A retriever takes a question and returns relevant Documents. That is its whole interface: retriever.invoke("Which technology stores data in memory?") gives back a list of Documents with page_content and metadata. Where they come from - Chroma, pgvector, a keyword index, a web search - is hidden behind it.
For a vector store, vector_store.as_retriever() builds one. search_kwargs={"k": 2} says how many to return; search_type picks the strategy: "similarity" for the closest, "mmr" for relevant results that are also different from each other, "similarity_score_threshold" for everything above a score.
A retriever is also a runnable - the same interface as prompts and models in Lesson 2.4. That is what lets Lesson 2.12 put it in a chain: question -> retriever -> prompt -> LLM.
Still no embedding model on this machine (Lesson 2.9), so the runs below use hand-made vectors with readable dimensions. Chroma, the retriever and LangChain’s MMR are the real code; only the vectors are chosen by us, and the lesson says so wherever a result depends on them.
Worked example: Find the 2 notes most relevant to this question.
1 - Rank the candidates
Cosine similarity to the question. The three backend notes are closest - they are near-copies of each other.
"What can Python be used for?" over five Python notes - three about backend, one about data science, one about machine learning. Hand-made vectors; scores computed by LangChain’s MMR (lambda_mult=0.5).
Vector store vs retriever
The vector store is where documents and vectors live and how they are searched. The retriever is the interface the rest of your application uses: question in, Documents out. Swap Chroma for pgvector and the code that calls retriever.invoke() does not change.
Not every retriever is a vector store underneath - keyword search, a SQL query or an API can sit behind the same interface. In this course it is Chroma.
Vector storeStores records and searches vectors.RetrieverQuestion in, Documents out - a runnable.as_retriever() and top-k
vector_store.as_retriever() returns a VectorStoreRetriever with search_type "similarity" and no search_kwargs - which means k defaults to 4: our five-note store returned four Documents. Pass search_kwargs={"k": 2} for two.
Asking for more than exists is not an error: k=10 on five notes returned five. A filter goes in search_kwargs too - {"k": 2, "filter": {"topic": "database"}} - and so does a score_threshold for search_type="similarity_score_threshold". An unknown search_type fails immediately: "search_type of fuzzy not allowed. Valid values are: (‘similarity’, ‘similarity_score_threshold’, ‘mmr’)".
The two notes most relevant
The lesson’s example: five notes from Lesson 2.10, the question "Which technology stores data in memory?", k=2. Our hand-made vectors give the notes dimensions you can read - programming, database, devops, frontend, memory - and put Redis on both database and memory.
The retriever returned note-5 (Redis) and note-2 (PostgreSQL), with metadata. Redis is the answer; PostgreSQL came second because it is also a database. That second result is typical: top-k always returns k documents, whether or not k of them are relevant.
Similarity search
search_type="similarity" ranks by closeness to the question and takes the first k. Simple, predictable, and usually the right starting point.
Its weakness shows when the collection has near-duplicates. Three of our five Python notes say almost the same thing about backend development, and they are also the closest to "What can Python be used for?" - so k=3 returned all three. Three slots of context, one fact.
MMR - Maximal Marginal Relevance
search_type="mmr" picks documents one at a time. The first is the most relevant. Each next one maximizes lambda_mult x relevance - (1 - lambda_mult) x its highest similarity to anything already picked. A document that repeats a chosen one is penalized, however relevant it is.
On the same question, MMR returned server-side development, data science, machine learning - all three uses of Python. The simulation shows the arithmetic: the backend near-copies scored -0.071 and -0.082 in the second round, data science 0.220.
fetch_k and lambda_mult
MMR first fetches fetch_k candidates by similarity (default 20), then chooses k of them. The pool matters: with fetch_k=3 the only candidates were the three backend notes, and MMR returned exactly what similarity did. Diversity needs something different in the pool.
lambda_mult (default 0.5) sets the balance. At 1.0 MMR is plain similarity - three backend notes. At 0.75: two backend notes and data science. At 0.5 and below: one of each.
lambda_mult=1.0server-side, backend applications, backend development - same as similarity.lambda_mult=0.75server-side, backend applications, data science.lambda_mult=0.5 (default)server-side, data science, machine learning.lambda_mult=0.0server-side, data science, machine learning.fetch_k=3Only backend notes in the pool - same as similarity.Choosing k
Too small and the document with the answer may be left out. Too large and the prompt fills with loosely related text: more tokens, more latency, and more for the model to be distracted by. There is no universal value - try a few on real questions and look at what comes back.
A score threshold is the other lever: search_type="similarity_score_threshold" with score_threshold=0.5 returned only Redis and PostgreSQL out of five, because the others scored below it. It answers "how relevant?" instead of "how many?" - and can return nothing at all.
Retriever vs RAG
A retriever only finds. RAG - Lesson 2.12 - puts what it found into the prompt and asks the LLM to answer from it. The retriever decides what the model gets to see; if the right document is not retrieved, no prompt can make the answer right.
Step-by-step code
retriever = vector_store.as_retriever(search_kwargs={"k": 2}) # the store from Lesson 2.10
documents = retriever.invoke("Which technology stores data in memory?")
for document in documents:
print(document.id, document.page_content, document.metadata)
# note-5 Redis is an in-memory data store. {'topic': 'database'}
# note-2 PostgreSQL is a relational database. {'topic': 'database'}from langchain_chroma import Chroma
from langchain_core.embeddings import Embeddings
class ToyEmbeddings(Embeddings):
"""Hand-made vectors so results are readable. A real model learns hundreds of dimensions."""
def __init__(self, vectors):
self.vectors = vectors
def embed_documents(self, texts):
return [self.vectors[t] for t in texts]
def embed_query(self, text):
return self.vectors[text]
# [backend, data science, machine learning, python]
notes = {
"Python is used for backend development.": [1.0, 0.0, 0.0, 0.5],
"Python is commonly used to build backend applications.": [0.95, 0.05, 0.0, 0.5],
"Python is popular for server-side development.": [0.9, 0.0, 0.05, 0.5],
"Python is widely used in data science.": [0.0, 1.0, 0.3, 0.5],
"Python is commonly used in machine learning.": [0.0, 0.3, 1.0, 0.5],
"What can Python be used for?": [0.6, 0.35, 0.3, 0.6],
}
vector_store = Chroma(
collection_name="python_notes",
embedding_function=ToyEmbeddings(notes),
collection_metadata={"hnsw:space": "cosine"},
)
vector_store.add_texts(list(notes)[:5], ids=[f"py-{i}" for i in range(1, 6)])question = "What can Python be used for?"
similarity = vector_store.as_retriever(search_type="similarity", search_kwargs={"k": 3})
mmr = vector_store.as_retriever(search_type="mmr", search_kwargs={"k": 3, "fetch_k": 20, "lambda_mult": 0.5})
for name, retriever in [("similarity", similarity), ("mmr", mmr)]:
print(name)
for document in retriever.invoke(question):
print(" ", document.page_content)similarity
Python is popular for server-side development.
Python is commonly used to build backend applications.
Python is used for backend development.
mmr
Python is popular for server-side development.
Python is widely used in data science.
Python is commonly used in machine learning.vector_store.as_retriever() search_type='similarity', search_kwargs={}
.invoke(question) 4 Documents - k defaults to 4
search_kwargs={"k": 10} on 5 notes 5 Documents - no error
search_kwargs={"k": 2, "filter": {"topic": "database"}}
Redis, PostgreSQL
search_type="similarity_score_threshold",
search_kwargs={"score_threshold": 0.5, "k": 5} Redis, PostgreSQL - the other three scored below 0.5
search_type="fuzzy"
ValidationError: search_type of fuzzy not allowed.
Valid values are: ('similarity', 'similarity_score_threshold', 'mmr')
retriever.batch([q1, q2]) one list of Documents per questionTip: Print what the retriever returns before you build RAG on it. Most bad RAG answers are retrieval problems: the right document was never in the context.
Watch out: MMR only diversifies within its candidate pool. If fetch_k is too small, or the collection really has only one kind of document, MMR returns the same as similarity.
Retrievers at a glance
as_retrieverWrap a vector store as a retriever.
vector_store.as_retriever()
invokeQuestion in, list of Documents out.
retriever.invoke(question)
kHow many Documents - default 4.
search_kwargs={"k": 2}similarityThe k closest.
search_type="similarity"
mmrRelevant and different.
search_type="mmr"
fetch_kMMR candidate pool - default 20.
{"k": 3, "fetch_k": 20}lambda_mult1 = relevance only, 0 = diversity only; default 0.5.
{"lambda_mult": 0.5}score thresholdEverything above a relevance score.
search_type="similarity_score_threshold"
Try it yourself
The code does not change. Swap the content string and the program does something else entirely.
“Call as_retriever().invoke() on the five notes and count the results.”
“Run similarity and MMR with k=3 on the Python notes and list the topics each covers.”
“Set fetch_k=3 for MMR. Why does the result change?”
“Try lambda_mult 1.0, 0.75 and 0.5 and watch when data science and machine learning appear.”
What usually goes wrong
as_retriever() with no search_kwargs returns four Documents. Set k deliberately.
✗ vector_store.as_retriever()✓ vector_store.as_retriever(search_kwargs={"k": 2})k, fetch_k, lambda_mult and filter go inside search_kwargs.
✗ as_retriever(search_type="mmr", k=3)✓ as_retriever(search_type="mmr", search_kwargs={"k": 3})With fetch_k equal to k, MMR has nothing to choose between and returns the similarity result.
Top-k always returns k documents if it can. PostgreSQL came second for "stores data in memory" because something had to.
Key points
- A retriever is question in, Documents out - and a runnable.
- vector_store.as_retriever() builds one; k defaults to 4.
- Settings go in search_kwargs: k, fetch_k, lambda_mult, filter.
- Similarity returns the closest k - including near-duplicates.
- MMR balances relevance against similarity to what it already picked.
- fetch_k is MMR’s pool; lambda_mult is the balance.
- The retriever decides what the LLM sees in RAG.
Quick check before you move on
Quiz
- 1.
vector_store.as_retriever().invoke(q) returns four Documents. Why four?
- 2.
MMR with k=3 and fetch_k=3 returned the same as similarity. Why?
- 3.
What does lambda_mult=1.0 do?
- 4.
For "stores data in memory", PostgreSQL came second. Is that a bug?
Interview questions
What is a retriever?
An abstraction that takes a query and returns relevant documents from a knowledge source. In RAG it fetches the chunks that go into the prompt.
What is top-k retrieval?
Returning the k most relevant documents for a query - for example search_kwargs={"k": 3}.
Similarity search vs MMR?
Similarity returns the closest documents. MMR selects iteratively, trading relevance against similarity to already-selected documents, which reduces redundancy.
What is fetch_k?
The number of candidates MMR fetches by similarity before choosing the final k. Too small and MMR cannot diversify.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...
AI
System Design
Backend
- GraphQL8 modules · 69 lessons planned
- Core Python13 modules · 75 lessons planned
- FastAPI5 sections · 20 lessons
- Node.js14 modules · 206 lessons planned
- Node.js Performance7 chapters · 36 topics
- Event Loop Lifecycle6 phases · 3 scenarios
- Docker & Containerization11 modules · 144 lessons planned
- AWS for Developers14 modules · 219 lessons planned
- CI/CD & DevOps Automation10 modules · 134 lessons planned