← Back to Agentic AI map
Lesson 2.15 · Building Agents with LangChain

Vector Stores Deep Dive

Go past the basics of Lesson 2.10: cut documents into good chunks, give them ids you can rebuild, update a file without leaving old pieces behind, and filter with every Chroma operator. Every result comes from a real run.

vector stores

What you will be able to do

  • Explain why a long document must be cut into chunks, and choose a chunk size
  • Use RecursiveCharacterTextSplitter with chunk_size and chunk_overlap
  • Give chunks ids you can rebuild, so indexing twice does not make duplicates
  • Update a changed file without leaving old chunks behind
  • Filter with Chroma’s operators: $eq, $ne, $gt, $gte, $lt, $lte, $in, $nin, $and, $or - and on the text itself
  • Use a filter and a score threshold in a retriever
  • Avoid the traps: vector size, distance type, and settings that are silently ignored

The idea, in plain English

In Lesson 2.10 you put five short notes into Chroma, searched them, and tried a filter, an update and a delete. That is enough for a demo. Real documents are different: they are long, they change, and old versions must not come back in answers.

This lesson is about keeping a vector store correct over time. We take a small support handbook, cut it into chunks, index it, change it the next day, and index it again. On the way we meet the problems that appear in real projects - duplicates, old chunks that stay forever, an old FAQ that answers instead of the new handbook - and fix each one.

A note about the model. No embedding model is installed on the machine these lessons were checked on (Lesson 2.9). So, as in Lessons 2.10 and 2.11, we use a stand-in: a tiny embedding class that only counts topic words such as "refund" and "shipping". It does not understand meaning - a real model would. But everything this lesson is about - chunks, ids, updates, deletes, filters, errors, saving to disk - is Chroma’s own behaviour, and all of it was really run. Versions: chromadb 1.5.9, langchain-chroma 1.1.0, langchain-text-splitters 1.1.3.

Worked example: A support handbook in Chroma: chunked, indexed, changed (30 days becomes 14), re-indexed and searched with filters.

workflowA file changes: two ways to re-index itstep 1 / 4

1 - Index version 1

The splitter cuts the 2025 handbook into 6 chunks, about one section each. Ids are "handbook.txt#0" to "handbook.txt#5".

chunks
6
ids
handbook.txt#0 ... #5
records
6
year
2025

Real runs. The handbook had 6 sections in 2025. In 2026 only 3 are left, and "30 days" became "14 days".

Words you will see in this lesson

Most of these words appeared in Lesson 2.10. A few are new.

Small dictionary
ChunkA small piece of a long document. Each chunk gets its own vector.
chunk_sizeThe longest a chunk may be, in characters.
chunk_overlapHow many characters the end of one chunk repeats at the start of the next.
SourceThe file a chunk came from. We store it in the metadata.
Stale chunkAn old chunk that should have been removed when its file changed.
OperatorA word like $gt or $in that says how a filter compares values.
Distance typeHow Chroma measures "close": l2 (the default) or cosine.

An everyday example: a library card index

Imagine a library that finds books with small index cards. Nobody writes one card for a whole book - one card for 300 pages helps no one. Instead, each card describes one chapter. That is chunking.

Each card also has labels: which book, which edition, which year. Those labels are metadata. When you want only the newest edition, you look only at cards with that label. That is a filter.

When a new edition arrives, the librarian removes ALL cards of the old edition, then writes new ones. If she only replaces the cards that have the same chapter number, the cards for chapters that were removed stay in the box - and someone will find them. That is the stale chunk problem, and it is the most common bug in real RAG systems.

Why we cut documents into chunks

A vector store keeps one vector for each record. If the record is a whole handbook, its one vector is a mix of every topic in it: refunds, shipping, passwords, invoices. A question about refunds matches that mix only weakly - and if it does match, you give the LLM the whole handbook instead of the one paragraph it needs.

We tested this on our 812-character handbook with the question "How do I get a refund?". With chunks of up to 800 characters, the best match was a 684-character chunk containing all six topics, with a relevance score of 0.46. With chunks of up to 200 characters, the best match was the 154-character refund section alone, with a relevance of 1.00.

The stand-in embedding model and the handbook - common.py
import math import re from langchain_core.embeddings import Embeddings # A stand-in for a real embedding model. It only COUNTS topic words - it does not # understand meaning. A real model (Lesson 2.9) gives 768 numbers; this gives 6 we can read. TOPICS = ["refund", "shipping", "password", "invoice", "delivery", "account"] class TopicEmbeddings(Embeddings): def _vector(self, text): words = re.findall(r"[a-z]+", text.lower()) counts = [sum(w.startswith(t) for w in words) for t in TOPICS] # "refunds" counts as refund length = math.sqrt(sum(c * c for c in counts)) or 1.0 return [c / length for c in counts] # same length for every text def embed_documents(self, texts): return [self._vector(t) for t in texts] def embed_query(self, text): return self._vector(text) HANDBOOK = """Refunds. You can ask for a refund within 30 days. A refund goes back to the card you paid with. Refunds for digital items are not possible after download. Shipping. Standard shipping takes 3 to 5 days. Express shipping takes 1 day and costs extra. Shipping is free for orders over 50 euros. Delivery problems. If a delivery is late, wait 2 more days, then contact us. If the delivery is damaged, send a photo within 48 hours. Passwords. To reset your password, click Forgot password on the login page. A password must have at least 12 characters. Invoices. Every order has an invoice in your account. You can download the invoice as a PDF. A company invoice needs your VAT number. Closing an account. You can close your account in Settings. Closing an account deletes your orders and invoices after 90 days."""
Example 1 - big chunks vs small chunks
from langchain_chroma import Chroma from langchain_text_splitters import RecursiveCharacterTextSplitter from common import HANDBOOK, TopicEmbeddings question = "How do I get a refund?" for size in [800, 200]: chunks = RecursiveCharacterTextSplitter(chunk_size=size, chunk_overlap=0).split_text(HANDBOOK) store = Chroma(collection_name=f"handbook_{size}", embedding_function=TopicEmbeddings(), collection_metadata={"hnsw:space": "cosine"}) store.add_texts(chunks, ids=[f"c{i}" for i in range(len(chunks))]) doc, score = store.similarity_search_with_relevance_scores(question, k=1)[0] print(f"chunk_size={size}: {len(chunks)} chunks | best match {len(doc.page_content)} chars, relevance {score:.2f}") print(" starts:", doc.page_content.replace("\n", " ")[:90], "...") print(" topics in it:", [t for t in ["Refund", "Shipping", "Delivery", "Password", "Invoice", "account"] if t in doc.page_content])
Output
chunk_size=800: 2 chunks | best match 684 chars, relevance 0.46 starts: Refunds. You can ask for a refund within 30 days. A refund goes back to the card you paid ... topics in it: ['Refund', 'Shipping', 'Delivery', 'Password', 'Invoice', 'account'] chunk_size=200: 6 chunks | best match 154 chars, relevance 1.00 starts: Refunds. You can ask for a refund within 30 days. A refund goes back to the card you paid ... topics in it: ['Refund']

Choosing chunk_size and chunk_overlap

RecursiveCharacterTextSplitter tries to cut at the best place first: between paragraphs ("\n\n"), then between lines, then between words. It only cuts in the middle of a sentence when a piece is still too long.

Too big, and one chunk holds many topics: with chunk_size=300 our first chunk held both the Refunds and the Shipping sections. Too small, and sentences are cut in half: with chunk_size=150, one chunk ended "...not possible after" and the next one was only "download." - 9 characters that mean nothing alone.

chunk_overlap repeats the end of one chunk at the start of the next, so a sentence cut at a border appears whole in at least one chunk. With overlap=40, that 9-character piece became "digital items are not possible after download." - readable again. Overlap costs a little space; it is most useful when your chunks are small.

A good starting point: make a chunk about one paragraph or one section, then test with real questions. If your documents have clear sections, as our handbook does, cut on those.

Example 2 - the same handbook, four settings
from langchain_text_splitters import RecursiveCharacterTextSplitter from common import HANDBOOK print("handbook:", len(HANDBOOK), "characters") for size, overlap in [(800, 0), (300, 0), (150, 0), (150, 40)]: splitter = RecursiveCharacterTextSplitter(chunk_size=size, chunk_overlap=overlap) chunks = splitter.split_text(HANDBOOK) lengths = [len(c) for c in chunks] print(f"chunk_size={size:3} overlap={overlap:2} -> {len(chunks):2} chunks, longest {max(lengths)}, shortest {min(lengths)}") splitter = RecursiveCharacterTextSplitter(chunk_size=150, chunk_overlap=40) print("\nfirst three chunks (150 / 40):") for c in splitter.split_text(HANDBOOK)[:3]: print(" |", c.replace("\n", " "))
Output
handbook: 812 characters chunk_size=800 overlap= 0 -> 2 chunks, longest 684, shortest 126 chunk_size=300 overlap= 0 -> 3 chunks, longest 291, shortest 256 chunk_size=150 overlap= 0 -> 7 chunks, longest 144, shortest 9 chunk_size=150 overlap=40 -> 7 chunks, longest 144, shortest 46 first three chunks (150 / 40): | Refunds. You can ask for a refund within 30 days. A refund goes back to the card you paid with. Refunds for digital items are not possible after | digital items are not possible after download. | Shipping. Standard shipping takes 3 to 5 days. Express shipping takes 1 day and costs extra. Shipping is free for orders over 50 euros.

Ids you can rebuild

Lesson 2.10 showed that adding a record with an id that already exists replaces it. That only helps if you can produce the same id again tomorrow. A random id (uuid4) is new every time, so indexing the same file twice gives every chunk twice.

Build the id from things that do not change: the file name and the chunk’s position, for example "handbook.txt#3". Index the file again, and the ids are the same, so the records are replaced instead of copied. We tested both: random ids, indexed twice -> 12 records. Rebuildable ids, indexed twice -> 6 records.

Output - the same 6 chunks indexed twice
random ids, indexed twice -> 12 records rebuildable ids, indexed twice -> 6 records

Tip: Rebuildable ids stop duplicates. They do NOT remove chunks that disappeared from the file - for that you need the next section.

When a file changes: the stale chunk problem

Our handbook changed: refunds now take 14 days, and three sections were removed. The new version has 3 chunks instead of 6. We added the new chunks with the same ids. Chroma replaced #0, #1 and #2 - and left #3, #4 and #5 from 2025 exactly where they were. The store held 6 handbook records: versions 1, 1, 1, 2, 2, 2.

Then we searched for "password reset". The top result was the 2025 Passwords section - a section that is no longer in the handbook. No error, no warning. In a RAG app, the LLM would answer from it.

The fix is simple: before adding a file, delete every chunk of that file, by metadata, not by id: store.delete(where={"source": "handbook.txt"}). Then add the new chunks. Other files are not touched. After the fix the store held 3 handbook records, all version 2.

Output - same ids only, then delete by source
v1 indexed: 7 records after adding v2 with the same ids: 6 handbook records; versions: [1, 1, 1, 2, 2, 2] search 'password reset' -> {'source': 'handbook.txt', 'section': 'passwords', 'year': 2025, 'version': 1} | Passwords. To reset your password, click Forg delete by source, then add v2: 3 records; versions: [2, 2, 2] total now: 4

Example 3 - index a file the right way

Here is everything so far in one small program: one function, index_file(), that works the first time and every time after. It cuts the text into chunks, deletes the file’s old chunks, and adds the new ones with rebuildable ids and metadata. We index the handbook, an old FAQ, then the changed handbook - twice - and search.

Example 3 - index_files.py
from langchain_chroma import Chroma from langchain_text_splitters import RecursiveCharacterTextSplitter from common import HANDBOOK, TopicEmbeddings # stand-in model + the sample handbook splitter = RecursiveCharacterTextSplitter(chunk_size=200, chunk_overlap=0) # about one section each store = Chroma( collection_name="support_docs", embedding_function=TopicEmbeddings(), persist_directory="kb", # saved to ./kb on disk collection_metadata={"hnsw:space": "cosine"}, # chosen once, when the collection is created ) def index_file(source: str, text: str, year: int): """Put one file into the store - the first time, or again after it changed.""" chunks = splitter.split_text(text) store.delete(where={"source": source}) # 1. remove EVERY old chunk of this file store.add_texts( # 2. add the new chunks chunks, metadatas=[{"source": source, "section": c.split(".")[0].lower(), "year": year} for c in chunks], ids=[f"{source}#{i}" for i in range(len(chunks))], # ids we can rebuild: file + position ) print(f"indexed {source}: {len(chunks)} chunks | store now has {store._collection.count()} records") # Day 1: the 2025 handbook, plus an old FAQ from 2022 index_file("handbook.txt", HANDBOOK, 2025) index_file("faq.txt", "Old FAQ. Refunds take 10 days to arrive.", 2022) # Day 2: the handbook changes - 30 days becomes 14, and only 3 sections are left new_handbook = "\n\n".join(HANDBOOK.split("\n\n")[:3]).replace("30 days", "14 days") index_file("handbook.txt", new_handbook, 2026) index_file("handbook.txt", new_handbook, 2026) # running it twice changes nothing question = "How long do I have to ask for a refund?" print("\nno filter:") for doc in store.similarity_search(question, k=2): print(" ", doc.metadata["source"], doc.metadata["year"], "|", doc.page_content[:60]) print("only the current handbook:") for doc in store.similarity_search(question, k=2, filter={"year": {"$gte": 2026}}): print(" ", doc.metadata["source"], doc.metadata["year"], "|", doc.page_content[:60]) print("\n'password reset' ->", store.similarity_search("password reset", k=1)[0].page_content[:50])
Output
indexed handbook.txt: 6 chunks | store now has 6 records indexed faq.txt: 1 chunks | store now has 7 records indexed handbook.txt: 3 chunks | store now has 4 records indexed handbook.txt: 3 chunks | store now has 4 records no filter: faq.txt 2022 | Old FAQ. Refunds take 10 days to arrive. handbook.txt 2026 | Refunds. You can ask for a refund within 14 days. A refund g only the current handbook: handbook.txt 2026 | Refunds. You can ask for a refund within 14 days. A refund g handbook.txt 2026 | Shipping. Standard shipping takes 3 to 5 days. Express shipp 'password reset' -> Delivery problems. If a delivery is late, wait 2 m

Reading the output: two more problems

Look at "no filter". The first answer is the 2022 FAQ: "Refunds take 10 days". It is about refunds, so it matches well - but it is old and wrong. Similar is not the same as correct. The year filter removed it, and the current handbook answered. Store a date or version in the metadata of every chunk, so you can filter like this.

Now look at the last line. We asked about "password reset", but the Passwords section was removed in 2026. The store still returned something: the Delivery problems section, which has nothing to do with passwords. similarity_search(k=1) always returns k results if the store has them, relevant or not. The next sections show how to filter, and how to drop results whose score is too low.

Filters: every operator, tested

A filter checks the metadata before Chroma returns a record. You can use the same filter in store.get(where=...), store.delete(where=...) and in similarity_search(..., filter=...).

A simple filter is {"key": value}, which means "equal". For anything else, use an operator. Two rules caught us while testing. First, a filter with two keys must be wrapped in $and - {"source": "handbook.txt", "year": 2026} raised ValueError: "Expected where to have exactly one operator". Second, $gt, $gte, $lt and $lte only work with numbers. With the string "2026" it raised an error whose message shows "got 2026" without quotes - easy to misread. Store years and versions as numbers.

There is also where_document, which filters on the chunk’s text, not its metadata. {"$contains": "days"} keeps only chunks whose text contains the word "days".

The results below come from the store built by Example 3: 3 handbook chunks from 2026 (refunds, shipping, delivery problems) and the old FAQ from 2022. The $not_contains line is empty because every chunk mentions "days".

Example 4 - filters on the store from Example 3
def show(label, **kw): result = store.get(**kw) # get() returns ids, texts, metadatas print(f"{label:44} -> {[m['section'] for m in result['metadatas']]}") show('{"source": "faq.txt"}', where={"source": "faq.txt"}) show('{"year": {"$gte": 2026}}', where={"year": {"$gte": 2026}}) show('{"section": {"$in": ["refunds", "old faq"]}}', where={"section": {"$in": ["refunds", "old faq"]}}) show('{"section": {"$ne": "refunds"}}', where={"section": {"$ne": "refunds"}}) show('$and: source handbook AND year < 2026', where={"$and": [{"source": "handbook.txt"}, {"year": {"$lt": 2026}}]}) show('$or: faq OR section shipping', where={"$or": [{"source": "faq.txt"}, {"section": "shipping"}]}) show('where_document {"$contains": "days"}', where_document={"$contains": "days"}) show('where_document {"$not_contains": "days"}', where_document={"$not_contains": "days"}) for label, where in [("two keys without $and", {"source": "handbook.txt", "year": 2026}), ("$gte with a string", {"year": {"$gte": "2026"}})]: try: store.get(where=where) except ValueError as e: print(f"{label:21} -> ValueError: {e}")
Output
{"source": "faq.txt"} -> ['old faq'] {"year": {"$gte": 2026}} -> ['refunds', 'shipping', 'delivery problems'] {"section": {"$in": ["refunds", "old faq"]}} -> ['old faq', 'refunds'] {"section": {"$ne": "refunds"}} -> ['old faq', 'shipping', 'delivery problems'] $and: source handbook AND year < 2026 -> [] $or: faq OR section shipping -> ['old faq', 'shipping'] where_document {"$contains": "days"} -> ['old faq', 'refunds', 'shipping', 'delivery problems'] where_document {"$not_contains": "days"} -> [] two keys without $and -> ValueError: Expected where to have exactly one operator, got {'source': 'handbook.txt', 'year': 2026} in get. $gte with a string -> ValueError: Expected operand value to be an int or a float for operator $gte, got 2026 in get.
Chroma filter operators
{"source": "faq.txt"}Equal - the same as {"source": {"$eq": "faq.txt"}}.
$neNot equal.
$gt $gte $lt $lteGreater / less than (or equal). Numbers only.
$in $ninThe value is (not) in a list.
$and $orCombine filters. A list of filters inside.
where_document $containsThe chunk text contains a word. Also $not_contains.

Tip: The $and filter returned an empty list - and that is good news. It means no 2025 handbook chunks are left after the fix. A filter like this makes a quick health check for your store.

Filters and score thresholds in a retriever

Lesson 2.11 turned a store into a retriever with as_retriever(). The filter goes into search_kwargs, next to k. Now the RAG chain from Lesson 2.12 only ever sees chunks that pass the filter.

To drop results that are not relevant enough, use search_type="similarity_score_threshold" with a score_threshold between 0 and 1. With a threshold of 0.5, the refund question returned only the two refund chunks; the question "What is the weather today?" returned an empty list, and LangChain printed the warning "No relevant docs were retrieved using the relevance score threshold 0.5". An empty list is better than a wrong chunk: your app can then answer "I could not find this in our documents."

Remember that relevance scores depend on the embedding model. With our stand-in, a chunk either shares a topic word (score 1.0) or not (0.0). A real model gives values in between, and you choose the threshold by testing real questions.

Example 5 - a filtered retriever, and a threshold
q = "How long does a refund take?" retriever = store.as_retriever(search_kwargs={"k": 2, "filter": {"source": "handbook.txt"}}) print("filtered retriever ->", [d.page_content[:40] for d in retriever.invoke(q)]) strict = store.as_retriever(search_type="similarity_score_threshold", search_kwargs={"score_threshold": 0.5, "k": 4}) print("threshold 0.5 ->", [d.page_content[:30] for d in strict.invoke(q)]) print("weather question ->", strict.invoke("What is the weather today?")) print("all scores ->", [(round(s, 2), d.page_content[:25]) for d, s in store.similarity_search_with_relevance_scores(q, k=4)])
Output
filtered retriever -> ['Refunds. You can ask for a refund within', 'Shipping. Standard shipping takes 3 to 5'] threshold 0.5 -> ['Old FAQ. Refunds take 10 days ', 'Refunds. You can ask for a ref'] weather question -> [] UserWarning: No relevant docs were retrieved using the relevance score threshold 0.5 all scores -> [(1.0, 'Old FAQ. Refunds take 10 '), (1.0, 'Refunds. You can ask for '), (0.0, 'Shipping. Standard shippi'), (0.0, 'Delivery problems. If a d')]

Watch out: Look at the filtered retriever: with k=2 its second result is "Shipping", which scored 0.0. A filter chooses WHICH chunks may be returned; it does not check relevance. Use a filter for "which documents", and a threshold for "how relevant".

Three settings that bite later

Vector size. A collection remembers the size of the first vectors you add - 6 for our stand-in, 768 for nomic-embed-text. We opened the same collection with a 768-number model: both search and add failed with "Collection expecting embedding with dimension of 6, got 768". If you change the embedding model, you must re-index everything into a new collection.

Distance type. Chroma’s default is l2 (straight-line distance). Lessons 2.10 and 2.11 used cosine. You choose it with collection_metadata={"hnsw:space": "cosine"} when the collection is created. We then opened the existing default collection again and asked for cosine: no error - and the collection was still l2. The setting is silently ignored for a collection that already exists.

On disk. With persist_directory="kb", the data stays after Python stops. We opened ./kb again in a new process: the collection was there, with the same 4 records.

Output - the three settings, checked
query with a 768-number model -> InvalidArgumentError: Collection expecting embedding with dimension of 6, got 768 add with a 768-number model -> InvalidArgumentError: Collection expecting embedding with dimension of 6, got 768 space=default: config space = l2 | raw scores [('Refunds take', 0.0), ('Shipping is ', 2.0)] space=cosine: config space = cosine | raw scores [('Refunds take', 0.0), ('Shipping is ', 1.0)] reopen the default collection asking for cosine -> space is l2 collections in ./db: ['handbook'] records after reopening: 4

Tip: Put the embedding model in the collection name, for example "support_docs_nomic_v1". When you change models, you create a new collection, re-index into it, and switch - the old one keeps working until you are ready.

A checklist for a real project

Before you put a vector store behind a real app, check these points. Each one comes from a problem in this lesson.

Vector store checklist
ChunksAbout one paragraph or section each. Test with real questions.
Metadata on every chunksource, section, and a date or version - as numbers if you compare them.
IdsRebuildable: "file#position". Never random.
Changed filedelete(where={"source": file}), then add. Run it twice to check.
Old documentsRemove them, or filter them out by date or version.
SearchFilter for "which documents", score threshold for "how relevant".
Embedding modelOne collection per model. Changing the model means re-indexing.
Distance typeSet when the collection is created. Cannot be changed later.

Vector stores at a glance

Cut into chunks

Up to chunk_size characters, cut at paragraphs first.

RecursiveCharacterTextSplitter(chunk_size=200, chunk_overlap=0).split_text(text)
Rebuildable ids

Same file and position, same id.

ids=[f"{source}#{i}" for i in range(len(chunks))]
Re-index a file

Remove every old chunk first.

store.delete(where={"source": source}); store.add_texts(...)
Read records

ids, texts and metadata, with a filter.

store.get(where={"year": {"$gte": 2026}})
Two conditions

Must be wrapped in $and (or $or).

{"$and": [{"source": "a.txt"}, {"year": 2026}]}
Filter on text

The chunk contains a word.

store.get(where_document={"$contains": "days"})
Filtered search

Only matching chunks can be returned.

store.similarity_search(q, k=2, filter={"year": 2026})
Score threshold

Drop results that are not relevant enough.

as_retriever(search_type="similarity_score_threshold", search_kwargs={"score_threshold": 0.5})
Distance type

Only when the collection is created.

collection_metadata={"hnsw:space": "cosine"}

Try it yourself

The code does not change. Swap the content string and the program does something else entirely.

Your own document

“Split a page of your own notes with chunk_size 200, 400 and 800. Ask three questions and see which size gives the best first result.”

Break it, then fix it

“In Example 3, remove the store.delete(...) line, index both versions, and search "password reset". Then put the line back.”

A health check

“Write a function that prints, for each source, how many chunks it has and which years they come from.”

Two conditions

“Find all handbook chunks from 2026 whose text contains "days". You need $and in where, and where_document.”

Set a threshold

“Make five questions your documents can answer and five they cannot. Find the threshold that keeps the first five and drops the rest.”

What usually goes wrong

Re-indexing with the same ids only

Chunks that disappeared from the file stay in the store and keep answering questions. In our run, the removed Passwords section was still the top result.

✗ store.add_texts(new_chunks, ids=new_ids)
✓ store.delete(where={"source": source})
store.add_texts(new_chunks, metadatas=metas, ids=new_ids)
Random ids

Every run creates new records. Indexing our 6 chunks twice gave 12.

✗ ids=[str(uuid.uuid4()) for _ in chunks]
✓ ids=[f"{source}#{i}" for i in range(len(chunks))]
One chunk for a whole document

The vector becomes a mix of every topic. Our 684-character chunk scored 0.46 for a refund question; the 154-character refund section scored 1.00.

Two keys in one filter

Chroma wants exactly one operator at the top level.

✗ where={"source": "handbook.txt", "year": 2026}
✓ where={"$and": [{"source": "handbook.txt"}, {"year": 2026}]}
Years and versions stored as text

$gt, $gte, $lt and $lte only accept numbers. Store numbers, and compare with numbers.

✗ {"year": "2026"} ... where={"year": {"$gte": "2026"}}
✓ {"year": 2026} ... where={"year": {"$gte": 2026}}
Changing a setting on an existing collection

Asking for cosine on an existing l2 collection is ignored without an error, and a new embedding model with a different vector size fails. Create a new collection and re-index.

Key points

  • Cut long documents into chunks of about one topic - small chunks match better and give the LLM less noise.
  • chunk_overlap keeps sentences readable when a chunk border cuts them.
  • Build ids from the file and position, so indexing twice replaces instead of copying.
  • When a file changes, delete all its chunks by source, then add - same ids alone leave stale chunks.
  • Similar is not the same as correct: store a date or version and filter out old documents.
  • Two filter conditions need $and or $or; $gt and friends need numbers.
  • A filter decides which chunks may come back; a score threshold decides which are relevant enough.
  • Vector size and distance type are fixed when a collection is created - a new model means a new collection.

Quick check before you move on

Why not store each document as one record?
Its vector mixes all its topics, so questions match it weakly, and the LLM receives far more text than it needs.
What does chunk_overlap do?
It repeats the end of each chunk at the start of the next, so a sentence cut at the border is whole in at least one chunk.
Why build ids like "handbook.txt#3"?
The same file and position give the same id every time, so indexing again replaces the records instead of adding copies.
A file shrank from 6 chunks to 3. How do you re-index it correctly?
Delete every chunk of that file by metadata - delete(where={"source": file}) - then add the 3 new chunks.
What is the difference between a filter and a score threshold?
A filter limits which chunks can be returned, by metadata. A threshold drops returned chunks that are not similar enough.

Quiz

  1. 1.

    You add 6 chunks with uuid4 ids, then run the same script again. How many records are there?

  2. 2.

    What does Chroma do with where={"source": "a.txt", "year": 2026}?

  3. 3.

    similarity_search(q, k=1) for a topic that is not in the store returns...?

  4. 4.

    You open an existing l2 collection with collection_metadata={"hnsw:space": "cosine"}. Which distance does it use?

  5. 5.

    You switch from a 6-number model to a 768-number model on the same collection. What happens?

Interview questions

How do you keep a vector index in sync with documents that change?

Store a source key (and version or date) in every chunk’s metadata and use deterministic ids. On change, delete all chunks for that source by metadata, then insert the new chunks. That removes chunks that disappeared - replacing by id alone leaves them behind - and makes re-indexing idempotent.

How do you choose a chunk size?

Aim for one coherent unit - a paragraph or section - so each vector represents one topic, and add overlap if natural boundaries are not available. Then measure retrieval on real questions; in our test, 200-character section chunks beat 800-character ones clearly.

Metadata filtering vs a relevance threshold - when do you use each?

Filters enforce hard rules: tenant, permissions, document type, current version. Thresholds handle quality: drop weak matches so the model is not fed unrelated text, and let the app say it found nothing. Production systems usually need both.

What must happen when you change the embedding model?

Everything must be re-embedded. Vectors from different models are not comparable and often differ in size - Chroma rejects the mismatch. Build a new collection, re-index, compare quality, then switch.

How would you stop outdated content from reaching users?

Delete it when the source changes, and filter by version or date at query time as a second guard. In our test an old FAQ ranked first for a refund question until a year filter removed it.

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...