Vector Stores Deep Dive
Go past the basics of Lesson 2.10: cut documents into good chunks, give them ids you can rebuild, update a file without leaving old pieces behind, and filter with every Chroma operator. Every result comes from a real run.
What you will be able to do
- Explain why a long document must be cut into chunks, and choose a chunk size
- Use RecursiveCharacterTextSplitter with chunk_size and chunk_overlap
- Give chunks ids you can rebuild, so indexing twice does not make duplicates
- Update a changed file without leaving old chunks behind
- Filter with Chroma’s operators: $eq, $ne, $gt, $gte, $lt, $lte, $in, $nin, $and, $or - and on the text itself
- Use a filter and a score threshold in a retriever
- Avoid the traps: vector size, distance type, and settings that are silently ignored
The idea, in plain English
In Lesson 2.10 you put five short notes into Chroma, searched them, and tried a filter, an update and a delete. That is enough for a demo. Real documents are different: they are long, they change, and old versions must not come back in answers.
This lesson is about keeping a vector store correct over time. We take a small support handbook, cut it into chunks, index it, change it the next day, and index it again. On the way we meet the problems that appear in real projects - duplicates, old chunks that stay forever, an old FAQ that answers instead of the new handbook - and fix each one.
A note about the model. No embedding model is installed on the machine these lessons were checked on (Lesson 2.9). So, as in Lessons 2.10 and 2.11, we use a stand-in: a tiny embedding class that only counts topic words such as "refund" and "shipping". It does not understand meaning - a real model would. But everything this lesson is about - chunks, ids, updates, deletes, filters, errors, saving to disk - is Chroma’s own behaviour, and all of it was really run. Versions: chromadb 1.5.9, langchain-chroma 1.1.0, langchain-text-splitters 1.1.3.
Worked example: A support handbook in Chroma: chunked, indexed, changed (30 days becomes 14), re-indexed and searched with filters.
1 - Index version 1
The splitter cuts the 2025 handbook into 6 chunks, about one section each. Ids are "handbook.txt#0" to "handbook.txt#5".
Real runs. The handbook had 6 sections in 2025. In 2026 only 3 are left, and "30 days" became "14 days".
Words you will see in this lesson
Most of these words appeared in Lesson 2.10. A few are new.
ChunkA small piece of a long document. Each chunk gets its own vector.chunk_sizeThe longest a chunk may be, in characters.chunk_overlapHow many characters the end of one chunk repeats at the start of the next.SourceThe file a chunk came from. We store it in the metadata.Stale chunkAn old chunk that should have been removed when its file changed.OperatorA word like $gt or $in that says how a filter compares values.Distance typeHow Chroma measures "close": l2 (the default) or cosine.An everyday example: a library card index
Imagine a library that finds books with small index cards. Nobody writes one card for a whole book - one card for 300 pages helps no one. Instead, each card describes one chapter. That is chunking.
Each card also has labels: which book, which edition, which year. Those labels are metadata. When you want only the newest edition, you look only at cards with that label. That is a filter.
When a new edition arrives, the librarian removes ALL cards of the old edition, then writes new ones. If she only replaces the cards that have the same chapter number, the cards for chapters that were removed stay in the box - and someone will find them. That is the stale chunk problem, and it is the most common bug in real RAG systems.
Why we cut documents into chunks
A vector store keeps one vector for each record. If the record is a whole handbook, its one vector is a mix of every topic in it: refunds, shipping, passwords, invoices. A question about refunds matches that mix only weakly - and if it does match, you give the LLM the whole handbook instead of the one paragraph it needs.
We tested this on our 812-character handbook with the question "How do I get a refund?". With chunks of up to 800 characters, the best match was a 684-character chunk containing all six topics, with a relevance score of 0.46. With chunks of up to 200 characters, the best match was the 154-character refund section alone, with a relevance of 1.00.
import math
import re
from langchain_core.embeddings import Embeddings
# A stand-in for a real embedding model. It only COUNTS topic words - it does not
# understand meaning. A real model (Lesson 2.9) gives 768 numbers; this gives 6 we can read.
TOPICS = ["refund", "shipping", "password", "invoice", "delivery", "account"]
class TopicEmbeddings(Embeddings):
def _vector(self, text):
words = re.findall(r"[a-z]+", text.lower())
counts = [sum(w.startswith(t) for w in words) for t in TOPICS] # "refunds" counts as refund
length = math.sqrt(sum(c * c for c in counts)) or 1.0
return [c / length for c in counts] # same length for every text
def embed_documents(self, texts):
return [self._vector(t) for t in texts]
def embed_query(self, text):
return self._vector(text)
HANDBOOK = """Refunds. You can ask for a refund within 30 days. A refund goes back to the card you paid with. Refunds for digital items are not possible after download.
Shipping. Standard shipping takes 3 to 5 days. Express shipping takes 1 day and costs extra. Shipping is free for orders over 50 euros.
Delivery problems. If a delivery is late, wait 2 more days, then contact us. If the delivery is damaged, send a photo within 48 hours.
Passwords. To reset your password, click Forgot password on the login page. A password must have at least 12 characters.
Invoices. Every order has an invoice in your account. You can download the invoice as a PDF. A company invoice needs your VAT number.
Closing an account. You can close your account in Settings. Closing an account deletes your orders and invoices after 90 days."""from langchain_chroma import Chroma
from langchain_text_splitters import RecursiveCharacterTextSplitter
from common import HANDBOOK, TopicEmbeddings
question = "How do I get a refund?"
for size in [800, 200]:
chunks = RecursiveCharacterTextSplitter(chunk_size=size, chunk_overlap=0).split_text(HANDBOOK)
store = Chroma(collection_name=f"handbook_{size}", embedding_function=TopicEmbeddings(),
collection_metadata={"hnsw:space": "cosine"})
store.add_texts(chunks, ids=[f"c{i}" for i in range(len(chunks))])
doc, score = store.similarity_search_with_relevance_scores(question, k=1)[0]
print(f"chunk_size={size}: {len(chunks)} chunks | best match {len(doc.page_content)} chars, relevance {score:.2f}")
print(" starts:", doc.page_content.replace("\n", " ")[:90], "...")
print(" topics in it:", [t for t in ["Refund", "Shipping", "Delivery", "Password", "Invoice", "account"] if t in doc.page_content])chunk_size=800: 2 chunks | best match 684 chars, relevance 0.46
starts: Refunds. You can ask for a refund within 30 days. A refund goes back to the card you paid ...
topics in it: ['Refund', 'Shipping', 'Delivery', 'Password', 'Invoice', 'account']
chunk_size=200: 6 chunks | best match 154 chars, relevance 1.00
starts: Refunds. You can ask for a refund within 30 days. A refund goes back to the card you paid ...
topics in it: ['Refund']Choosing chunk_size and chunk_overlap
RecursiveCharacterTextSplitter tries to cut at the best place first: between paragraphs ("\n\n"), then between lines, then between words. It only cuts in the middle of a sentence when a piece is still too long.
Too big, and one chunk holds many topics: with chunk_size=300 our first chunk held both the Refunds and the Shipping sections. Too small, and sentences are cut in half: with chunk_size=150, one chunk ended "...not possible after" and the next one was only "download." - 9 characters that mean nothing alone.
chunk_overlap repeats the end of one chunk at the start of the next, so a sentence cut at a border appears whole in at least one chunk. With overlap=40, that 9-character piece became "digital items are not possible after download." - readable again. Overlap costs a little space; it is most useful when your chunks are small.
A good starting point: make a chunk about one paragraph or one section, then test with real questions. If your documents have clear sections, as our handbook does, cut on those.
from langchain_text_splitters import RecursiveCharacterTextSplitter
from common import HANDBOOK
print("handbook:", len(HANDBOOK), "characters")
for size, overlap in [(800, 0), (300, 0), (150, 0), (150, 40)]:
splitter = RecursiveCharacterTextSplitter(chunk_size=size, chunk_overlap=overlap)
chunks = splitter.split_text(HANDBOOK)
lengths = [len(c) for c in chunks]
print(f"chunk_size={size:3} overlap={overlap:2} -> {len(chunks):2} chunks, longest {max(lengths)}, shortest {min(lengths)}")
splitter = RecursiveCharacterTextSplitter(chunk_size=150, chunk_overlap=40)
print("\nfirst three chunks (150 / 40):")
for c in splitter.split_text(HANDBOOK)[:3]:
print(" |", c.replace("\n", " "))handbook: 812 characters
chunk_size=800 overlap= 0 -> 2 chunks, longest 684, shortest 126
chunk_size=300 overlap= 0 -> 3 chunks, longest 291, shortest 256
chunk_size=150 overlap= 0 -> 7 chunks, longest 144, shortest 9
chunk_size=150 overlap=40 -> 7 chunks, longest 144, shortest 46
first three chunks (150 / 40):
| Refunds. You can ask for a refund within 30 days. A refund goes back to the card you paid with. Refunds for digital items are not possible after
| digital items are not possible after download.
| Shipping. Standard shipping takes 3 to 5 days. Express shipping takes 1 day and costs extra. Shipping is free for orders over 50 euros.Ids you can rebuild
Lesson 2.10 showed that adding a record with an id that already exists replaces it. That only helps if you can produce the same id again tomorrow. A random id (uuid4) is new every time, so indexing the same file twice gives every chunk twice.
Build the id from things that do not change: the file name and the chunk’s position, for example "handbook.txt#3". Index the file again, and the ids are the same, so the records are replaced instead of copied. We tested both: random ids, indexed twice -> 12 records. Rebuildable ids, indexed twice -> 6 records.
random ids, indexed twice -> 12 records
rebuildable ids, indexed twice -> 6 recordsTip: Rebuildable ids stop duplicates. They do NOT remove chunks that disappeared from the file - for that you need the next section.
When a file changes: the stale chunk problem
Our handbook changed: refunds now take 14 days, and three sections were removed. The new version has 3 chunks instead of 6. We added the new chunks with the same ids. Chroma replaced #0, #1 and #2 - and left #3, #4 and #5 from 2025 exactly where they were. The store held 6 handbook records: versions 1, 1, 1, 2, 2, 2.
Then we searched for "password reset". The top result was the 2025 Passwords section - a section that is no longer in the handbook. No error, no warning. In a RAG app, the LLM would answer from it.
The fix is simple: before adding a file, delete every chunk of that file, by metadata, not by id: store.delete(where={"source": "handbook.txt"}). Then add the new chunks. Other files are not touched. After the fix the store held 3 handbook records, all version 2.
v1 indexed: 7 records
after adding v2 with the same ids: 6 handbook records; versions: [1, 1, 1, 2, 2, 2]
search 'password reset' -> {'source': 'handbook.txt', 'section': 'passwords', 'year': 2025, 'version': 1} | Passwords. To reset your password, click Forg
delete by source, then add v2: 3 records; versions: [2, 2, 2]
total now: 4Example 3 - index a file the right way
Here is everything so far in one small program: one function, index_file(), that works the first time and every time after. It cuts the text into chunks, deletes the file’s old chunks, and adds the new ones with rebuildable ids and metadata. We index the handbook, an old FAQ, then the changed handbook - twice - and search.
from langchain_chroma import Chroma
from langchain_text_splitters import RecursiveCharacterTextSplitter
from common import HANDBOOK, TopicEmbeddings # stand-in model + the sample handbook
splitter = RecursiveCharacterTextSplitter(chunk_size=200, chunk_overlap=0) # about one section each
store = Chroma(
collection_name="support_docs",
embedding_function=TopicEmbeddings(),
persist_directory="kb", # saved to ./kb on disk
collection_metadata={"hnsw:space": "cosine"}, # chosen once, when the collection is created
)
def index_file(source: str, text: str, year: int):
"""Put one file into the store - the first time, or again after it changed."""
chunks = splitter.split_text(text)
store.delete(where={"source": source}) # 1. remove EVERY old chunk of this file
store.add_texts( # 2. add the new chunks
chunks,
metadatas=[{"source": source, "section": c.split(".")[0].lower(), "year": year}
for c in chunks],
ids=[f"{source}#{i}" for i in range(len(chunks))], # ids we can rebuild: file + position
)
print(f"indexed {source}: {len(chunks)} chunks | store now has {store._collection.count()} records")
# Day 1: the 2025 handbook, plus an old FAQ from 2022
index_file("handbook.txt", HANDBOOK, 2025)
index_file("faq.txt", "Old FAQ. Refunds take 10 days to arrive.", 2022)
# Day 2: the handbook changes - 30 days becomes 14, and only 3 sections are left
new_handbook = "\n\n".join(HANDBOOK.split("\n\n")[:3]).replace("30 days", "14 days")
index_file("handbook.txt", new_handbook, 2026)
index_file("handbook.txt", new_handbook, 2026) # running it twice changes nothing
question = "How long do I have to ask for a refund?"
print("\nno filter:")
for doc in store.similarity_search(question, k=2):
print(" ", doc.metadata["source"], doc.metadata["year"], "|", doc.page_content[:60])
print("only the current handbook:")
for doc in store.similarity_search(question, k=2, filter={"year": {"$gte": 2026}}):
print(" ", doc.metadata["source"], doc.metadata["year"], "|", doc.page_content[:60])
print("\n'password reset' ->", store.similarity_search("password reset", k=1)[0].page_content[:50])indexed handbook.txt: 6 chunks | store now has 6 records
indexed faq.txt: 1 chunks | store now has 7 records
indexed handbook.txt: 3 chunks | store now has 4 records
indexed handbook.txt: 3 chunks | store now has 4 records
no filter:
faq.txt 2022 | Old FAQ. Refunds take 10 days to arrive.
handbook.txt 2026 | Refunds. You can ask for a refund within 14 days. A refund g
only the current handbook:
handbook.txt 2026 | Refunds. You can ask for a refund within 14 days. A refund g
handbook.txt 2026 | Shipping. Standard shipping takes 3 to 5 days. Express shipp
'password reset' -> Delivery problems. If a delivery is late, wait 2 mReading the output: two more problems
Look at "no filter". The first answer is the 2022 FAQ: "Refunds take 10 days". It is about refunds, so it matches well - but it is old and wrong. Similar is not the same as correct. The year filter removed it, and the current handbook answered. Store a date or version in the metadata of every chunk, so you can filter like this.
Now look at the last line. We asked about "password reset", but the Passwords section was removed in 2026. The store still returned something: the Delivery problems section, which has nothing to do with passwords. similarity_search(k=1) always returns k results if the store has them, relevant or not. The next sections show how to filter, and how to drop results whose score is too low.
Filters: every operator, tested
A filter checks the metadata before Chroma returns a record. You can use the same filter in store.get(where=...), store.delete(where=...) and in similarity_search(..., filter=...).
A simple filter is {"key": value}, which means "equal". For anything else, use an operator. Two rules caught us while testing. First, a filter with two keys must be wrapped in $and - {"source": "handbook.txt", "year": 2026} raised ValueError: "Expected where to have exactly one operator". Second, $gt, $gte, $lt and $lte only work with numbers. With the string "2026" it raised an error whose message shows "got 2026" without quotes - easy to misread. Store years and versions as numbers.
There is also where_document, which filters on the chunk’s text, not its metadata. {"$contains": "days"} keeps only chunks whose text contains the word "days".
The results below come from the store built by Example 3: 3 handbook chunks from 2026 (refunds, shipping, delivery problems) and the old FAQ from 2022. The $not_contains line is empty because every chunk mentions "days".
def show(label, **kw):
result = store.get(**kw) # get() returns ids, texts, metadatas
print(f"{label:44} -> {[m['section'] for m in result['metadatas']]}")
show('{"source": "faq.txt"}', where={"source": "faq.txt"})
show('{"year": {"$gte": 2026}}', where={"year": {"$gte": 2026}})
show('{"section": {"$in": ["refunds", "old faq"]}}', where={"section": {"$in": ["refunds", "old faq"]}})
show('{"section": {"$ne": "refunds"}}', where={"section": {"$ne": "refunds"}})
show('$and: source handbook AND year < 2026', where={"$and": [{"source": "handbook.txt"}, {"year": {"$lt": 2026}}]})
show('$or: faq OR section shipping', where={"$or": [{"source": "faq.txt"}, {"section": "shipping"}]})
show('where_document {"$contains": "days"}', where_document={"$contains": "days"})
show('where_document {"$not_contains": "days"}', where_document={"$not_contains": "days"})
for label, where in [("two keys without $and", {"source": "handbook.txt", "year": 2026}),
("$gte with a string", {"year": {"$gte": "2026"}})]:
try:
store.get(where=where)
except ValueError as e:
print(f"{label:21} -> ValueError: {e}"){"source": "faq.txt"} -> ['old faq']
{"year": {"$gte": 2026}} -> ['refunds', 'shipping', 'delivery problems']
{"section": {"$in": ["refunds", "old faq"]}} -> ['old faq', 'refunds']
{"section": {"$ne": "refunds"}} -> ['old faq', 'shipping', 'delivery problems']
$and: source handbook AND year < 2026 -> []
$or: faq OR section shipping -> ['old faq', 'shipping']
where_document {"$contains": "days"} -> ['old faq', 'refunds', 'shipping', 'delivery problems']
where_document {"$not_contains": "days"} -> []
two keys without $and -> ValueError: Expected where to have exactly one operator, got {'source': 'handbook.txt', 'year': 2026} in get.
$gte with a string -> ValueError: Expected operand value to be an int or a float for operator $gte, got 2026 in get.{"source": "faq.txt"}Equal - the same as {"source": {"$eq": "faq.txt"}}.$neNot equal.$gt $gte $lt $lteGreater / less than (or equal). Numbers only.$in $ninThe value is (not) in a list.$and $orCombine filters. A list of filters inside.where_document $containsThe chunk text contains a word. Also $not_contains.Tip: The $and filter returned an empty list - and that is good news. It means no 2025 handbook chunks are left after the fix. A filter like this makes a quick health check for your store.
Filters and score thresholds in a retriever
Lesson 2.11 turned a store into a retriever with as_retriever(). The filter goes into search_kwargs, next to k. Now the RAG chain from Lesson 2.12 only ever sees chunks that pass the filter.
To drop results that are not relevant enough, use search_type="similarity_score_threshold" with a score_threshold between 0 and 1. With a threshold of 0.5, the refund question returned only the two refund chunks; the question "What is the weather today?" returned an empty list, and LangChain printed the warning "No relevant docs were retrieved using the relevance score threshold 0.5". An empty list is better than a wrong chunk: your app can then answer "I could not find this in our documents."
Remember that relevance scores depend on the embedding model. With our stand-in, a chunk either shares a topic word (score 1.0) or not (0.0). A real model gives values in between, and you choose the threshold by testing real questions.
q = "How long does a refund take?"
retriever = store.as_retriever(search_kwargs={"k": 2, "filter": {"source": "handbook.txt"}})
print("filtered retriever ->", [d.page_content[:40] for d in retriever.invoke(q)])
strict = store.as_retriever(search_type="similarity_score_threshold",
search_kwargs={"score_threshold": 0.5, "k": 4})
print("threshold 0.5 ->", [d.page_content[:30] for d in strict.invoke(q)])
print("weather question ->", strict.invoke("What is the weather today?"))
print("all scores ->", [(round(s, 2), d.page_content[:25])
for d, s in store.similarity_search_with_relevance_scores(q, k=4)])filtered retriever -> ['Refunds. You can ask for a refund within', 'Shipping. Standard shipping takes 3 to 5']
threshold 0.5 -> ['Old FAQ. Refunds take 10 days ', 'Refunds. You can ask for a ref']
weather question -> []
UserWarning: No relevant docs were retrieved using the relevance score threshold 0.5
all scores -> [(1.0, 'Old FAQ. Refunds take 10 '), (1.0, 'Refunds. You can ask for '), (0.0, 'Shipping. Standard shippi'), (0.0, 'Delivery problems. If a d')]Watch out: Look at the filtered retriever: with k=2 its second result is "Shipping", which scored 0.0. A filter chooses WHICH chunks may be returned; it does not check relevance. Use a filter for "which documents", and a threshold for "how relevant".
Three settings that bite later
Vector size. A collection remembers the size of the first vectors you add - 6 for our stand-in, 768 for nomic-embed-text. We opened the same collection with a 768-number model: both search and add failed with "Collection expecting embedding with dimension of 6, got 768". If you change the embedding model, you must re-index everything into a new collection.
Distance type. Chroma’s default is l2 (straight-line distance). Lessons 2.10 and 2.11 used cosine. You choose it with collection_metadata={"hnsw:space": "cosine"} when the collection is created. We then opened the existing default collection again and asked for cosine: no error - and the collection was still l2. The setting is silently ignored for a collection that already exists.
On disk. With persist_directory="kb", the data stays after Python stops. We opened ./kb again in a new process: the collection was there, with the same 4 records.
query with a 768-number model -> InvalidArgumentError: Collection expecting embedding with dimension of 6, got 768
add with a 768-number model -> InvalidArgumentError: Collection expecting embedding with dimension of 6, got 768
space=default: config space = l2 | raw scores [('Refunds take', 0.0), ('Shipping is ', 2.0)]
space=cosine: config space = cosine | raw scores [('Refunds take', 0.0), ('Shipping is ', 1.0)]
reopen the default collection asking for cosine -> space is l2
collections in ./db: ['handbook']
records after reopening: 4Tip: Put the embedding model in the collection name, for example "support_docs_nomic_v1". When you change models, you create a new collection, re-index into it, and switch - the old one keeps working until you are ready.
A checklist for a real project
Before you put a vector store behind a real app, check these points. Each one comes from a problem in this lesson.
ChunksAbout one paragraph or section each. Test with real questions.Metadata on every chunksource, section, and a date or version - as numbers if you compare them.IdsRebuildable: "file#position". Never random.Changed filedelete(where={"source": file}), then add. Run it twice to check.Old documentsRemove them, or filter them out by date or version.SearchFilter for "which documents", score threshold for "how relevant".Embedding modelOne collection per model. Changing the model means re-indexing.Distance typeSet when the collection is created. Cannot be changed later.Vector stores at a glance
Cut into chunksUp to chunk_size characters, cut at paragraphs first.
RecursiveCharacterTextSplitter(chunk_size=200, chunk_overlap=0).split_text(text)
Rebuildable idsSame file and position, same id.
ids=[f"{source}#{i}" for i in range(len(chunks))]Re-index a fileRemove every old chunk first.
store.delete(where={"source": source}); store.add_texts(...)Read recordsids, texts and metadata, with a filter.
store.get(where={"year": {"$gte": 2026}})Two conditionsMust be wrapped in $and (or $or).
{"$and": [{"source": "a.txt"}, {"year": 2026}]}Filter on textThe chunk contains a word.
store.get(where_document={"$contains": "days"})Filtered searchOnly matching chunks can be returned.
store.similarity_search(q, k=2, filter={"year": 2026})Score thresholdDrop results that are not relevant enough.
as_retriever(search_type="similarity_score_threshold", search_kwargs={"score_threshold": 0.5})Distance typeOnly when the collection is created.
collection_metadata={"hnsw:space": "cosine"}Try it yourself
The code does not change. Swap the content string and the program does something else entirely.
“Split a page of your own notes with chunk_size 200, 400 and 800. Ask three questions and see which size gives the best first result.”
“In Example 3, remove the store.delete(...) line, index both versions, and search "password reset". Then put the line back.”
“Write a function that prints, for each source, how many chunks it has and which years they come from.”
“Find all handbook chunks from 2026 whose text contains "days". You need $and in where, and where_document.”
“Make five questions your documents can answer and five they cannot. Find the threshold that keeps the first five and drops the rest.”
What usually goes wrong
Chunks that disappeared from the file stay in the store and keep answering questions. In our run, the removed Passwords section was still the top result.
✗ store.add_texts(new_chunks, ids=new_ids)✓ store.delete(where={"source": source})
store.add_texts(new_chunks, metadatas=metas, ids=new_ids)Every run creates new records. Indexing our 6 chunks twice gave 12.
✗ ids=[str(uuid.uuid4()) for _ in chunks]✓ ids=[f"{source}#{i}" for i in range(len(chunks))]The vector becomes a mix of every topic. Our 684-character chunk scored 0.46 for a refund question; the 154-character refund section scored 1.00.
Chroma wants exactly one operator at the top level.
✗ where={"source": "handbook.txt", "year": 2026}✓ where={"$and": [{"source": "handbook.txt"}, {"year": 2026}]}$gt, $gte, $lt and $lte only accept numbers. Store numbers, and compare with numbers.
✗ {"year": "2026"} ... where={"year": {"$gte": "2026"}}✓ {"year": 2026} ... where={"year": {"$gte": 2026}}Asking for cosine on an existing l2 collection is ignored without an error, and a new embedding model with a different vector size fails. Create a new collection and re-index.
Key points
- Cut long documents into chunks of about one topic - small chunks match better and give the LLM less noise.
- chunk_overlap keeps sentences readable when a chunk border cuts them.
- Build ids from the file and position, so indexing twice replaces instead of copying.
- When a file changes, delete all its chunks by source, then add - same ids alone leave stale chunks.
- Similar is not the same as correct: store a date or version and filter out old documents.
- Two filter conditions need $and or $or; $gt and friends need numbers.
- A filter decides which chunks may come back; a score threshold decides which are relevant enough.
- Vector size and distance type are fixed when a collection is created - a new model means a new collection.
Quick check before you move on
Quiz
- 1.
You add 6 chunks with uuid4 ids, then run the same script again. How many records are there?
- 2.
What does Chroma do with where={"source": "a.txt", "year": 2026}?
- 3.
similarity_search(q, k=1) for a topic that is not in the store returns...?
- 4.
You open an existing l2 collection with collection_metadata={"hnsw:space": "cosine"}. Which distance does it use?
- 5.
You switch from a 6-number model to a 768-number model on the same collection. What happens?
Interview questions
How do you keep a vector index in sync with documents that change?
Store a source key (and version or date) in every chunk’s metadata and use deterministic ids. On change, delete all chunks for that source by metadata, then insert the new chunks. That removes chunks that disappeared - replacing by id alone leaves them behind - and makes re-indexing idempotent.
How do you choose a chunk size?
Aim for one coherent unit - a paragraph or section - so each vector represents one topic, and add overlap if natural boundaries are not available. Then measure retrieval on real questions; in our test, 200-character section chunks beat 800-character ones clearly.
Metadata filtering vs a relevance threshold - when do you use each?
Filters enforce hard rules: tenant, permissions, document type, current version. Thresholds handle quality: drop weak matches so the model is not fed unrelated text, and let the app say it found nothing. Production systems usually need both.
What must happen when you change the embedding model?
Everything must be re-embedded. Vectors from different models are not comparable and often differ in size - Chroma rejects the mismatch. Build a new collection, re-index, compare quality, then switch.
How would you stop outdated content from reaching users?
Delete it when the source changes, and filter by version or date at query time as a second guard. In our test an old FAQ ranked first for a refund question until a year filter removed it.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...