Vector Stores
Index five short notes into Chroma, inspect exactly what got stored, and search them by meaning.
What you will be able to do
- Explain what a vector store keeps, and why it keeps the text as well as the vector
- Index documents into Chroma with ids and metadata
- Inspect what got stored - including the vectors
- Search with similarity_search, and read what the score means
- Filter by metadata, update and delete by id, and persist to disk
- Place Chroma among the alternatives, such as PostgreSQL with pgvector
The idea, in plain English
Lesson 2.9 ended with vectors in a Python list and a loop that compared the query with every one. A vector store does that job properly. For each record it keeps an id, the original text, metadata, and the vector - plus an index that finds the nearest vectors to a query without comparing against every record.
Chroma is the vector store this course uses. Through LangChain you create a Chroma collection with an embedding function, call add_texts() - which embeds each text with that function and stores the lot - and later call similarity_search(question, k=2), which embeds the question the same way and returns the two closest Documents.
The vector finds the record; the text is what you use. After a search you want "Redis is an in-memory data store." to give to an LLM, not 768 numbers. That is why the store keeps both, and why metadata - topic, source, tenant - travels with them.
No embedding model is installed on the machine these lessons were checked on (see Lesson 2.9). Chroma’s behaviour - storing, inspecting, ids, metadata, filters, distances, persistence - was all run, with LangChain’s fake embeddings in place of the model. Rankings for real questions need a real model, so those are illustrative and labelled.
Worked example: Index 5 short notes into Chroma.
1 - Index the notes
add_texts() embeds each note with the embedding function and stores id, text, metadata and vector together.
add_texts() on the way in, similarity_search() on the way out - both go through the same embedding model.
What a vector store keeps
Four things per record: an id, the original text (Chroma calls it the document), metadata, and the embedding vector. And one thing for the collection: an index that finds near neighbours quickly. Chroma’s index is HNSW, a graph that jumps towards close vectors instead of measuring every one.
A normal database answers "which rows match this condition?" - WHERE title ILIKE ‘%python%’. A vector store answers "which records are closest in meaning to this?" - and "How can I run applications in isolated containers?" can find the Docker note without sharing the word Docker.
Indexing with Chroma
pip install langchain-chroma langchain-ollama. Create the store with Chroma(collection_name="tech_notes", embedding_function=embeddings) - a collection is a named set of records - then call add_texts(texts=..., metadatas=..., ids=...). Indexing means exactly this: embed each text, store the vector with its text and metadata, make it searchable.
You do not call embed_documents() yourself. The store calls the embedding function you gave it, for documents now and for queries later - which is what guarantees both are in the same space. If that function is a chat model, indexing fails at the first add_texts: with llama3, "This server does not support embeddings."
Ids decide whether you update or duplicate
add_texts() returns the ids it stored. Give your own - note-1 to note-5 - and adding the same five notes again replaced them: the count stayed 5. Leave ids out and Chroma generates random UUIDs, so running the same indexing script twice stored every note twice: the count went to 10.
With ids you can also change single records: vector_store.delete(ids=["note-3"]) left four, and update_document("note-5", ...) replaced the Redis note’s text (and re-embedded it). Use ids that come from your data - a file path, a lesson id - so re-indexing is safe.
add_texts(texts, ids=ids) twiceCount stays 5 - the second call replaces.add_texts(texts) twiceCount 10 - random ids, every note stored twice.delete(ids=["note-3"])Count 4: note-1, note-2, note-4, note-5.update_document("note-5", doc)New text, new vector, same id.Inspecting what got stored
Do not assume indexing worked - look. vector_store._collection.count() gave 5. vector_store.get() returned the ids, the five documents and their metadatas - and "embeddings": None. Chroma leaves the vectors out of get() unless you ask: get(ids=["note-5"], include=["documents", "metadatas", "embeddings"]) returned the Redis note with its 768 numbers.
That is also where you catch the classic indexing bugs: duplicates from missing ids, empty or truncated documents, missing metadata, the wrong collection.
Searching, and what the score means
similarity_search(query, k=2) returns the two nearest records as Documents: page_content is the text, metadata the dict, id the id. similarity_search_with_score returns (Document, score) pairs - and in Chroma that score is a distance. The default space is l2: searching for the exact text of the Redis note returned it with 0.0, the next note far behind. Lower is closer.
Create the collection with collection_metadata={"hnsw:space": "cosine"} and the score becomes cosine distance - 1 minus cosine similarity: exact match 0.0, unrelated around 1. similarity_search_with_relevance_scores converts distances to "higher is better", but only makes sense for normalized vectors; with our fake ones it produced -1069 and a warning. Always check which kind of number you have.
similarity_search_with_score, l2Euclidean distance. 0 = identical, lower = closer.similarity_search_with_score, cosine1 - cosine similarity. 0 = identical, about 1 = unrelated.similarity_search_with_relevance_scoresConverted to 0-1, higher = closer - assumes normalized vectors.Metadata and filters
Metadata is a dict per record. With topic on each note, similarity_search(query, k=5, filter={"topic": "database"}) returned only the PostgreSQL and Redis notes - the filter is applied before ranking, so nothing outside it can appear.
In a real application metadata carries source, document type, dates - and, for a multi-tenant platform, tenant_id and course_id. Filtering on tenant_id is how one store can serve many academies without one tenant’s documents showing up in another’s answers. Treat that filter as a security rule: apply it in your code on every query, never as an option.
In memory or on disk
Without a persist_directory, Chroma keeps the collection in memory - and shares it within the process: a second Chroma(collection_name="tech_notes") saw the first one’s records. Re-running a notebook cell that indexes without ids piles up duplicates this way.
Chroma(..., persist_directory="./chroma_db") writes a chroma.sqlite3 file. We indexed two notes in one Python process and read them back in a second: count 2, both texts. Index once, search many times - you do not re-embed your documents on every start.
Chroma is one option
Chroma is easy to start with: a pip install, no server. It is not the only choice. PostgreSQL with the pgvector extension stores vectors in a normal table, next to users, courses and lessons, with SQL filters and joins - attractive if PostgreSQL is already your database. Dedicated services exist too.
The concepts carry over unchanged: records with text, metadata and a vector; the same embedding model for documents and queries; nearest-neighbour search with a k and a filter.
Step-by-step code
# pip install langchain-chroma langchain-ollama
# ollama pull nomic-embed-text
from langchain_chroma import Chroma
from langchain_ollama import OllamaEmbeddings
texts = [
"Python is a programming language.",
"PostgreSQL is a relational database.",
"Docker packages applications into containers.",
"React is a JavaScript UI library.",
"Redis is an in-memory data store.",
]
metadatas = [
{"topic": "programming"},
{"topic": "database"},
{"topic": "devops"},
{"topic": "frontend"},
{"topic": "database"},
]
ids = ["note-1", "note-2", "note-3", "note-4", "note-5"]
embeddings = OllamaEmbeddings(model="nomic-embed-text")
vector_store = Chroma(
collection_name="tech_notes",
embedding_function=embeddings,
)
print(vector_store.add_texts(texts=texts, metadatas=metadatas, ids=ids))
# ['note-1', 'note-2', 'note-3', 'note-4', 'note-5']print(vector_store._collection.count()) # 5
stored = vector_store.get()
print(stored["ids"]) # ['note-1', 'note-2', 'note-3', 'note-4', 'note-5']
print(stored["documents"]) # ['Python is a programming language.', 'PostgreSQL is a relational database.', ...]
print(stored["metadatas"]) # [{'topic': 'programming'}, {'topic': 'database'}, ...]
print(stored["embeddings"]) # None - vectors are left out unless you ask
note = vector_store.get(ids=["note-5"], include=["documents", "metadatas", "embeddings"])
print(note["documents"]) # ['Redis is an in-memory data store.']
print(len(note["embeddings"][0])) # the vector length, e.g. 768results = vector_store.similarity_search("Which technology stores data in memory?", k=2)
for document in results:
print(document.id, document.page_content, document.metadata)
for document, score in vector_store.similarity_search_with_score(
"Which technology stores data in memory?", k=2
):
print(round(score, 3), document.page_content) # a distance: lower = closernote-5 Redis is an in-memory data store. {'topic': 'database'}
note-2 PostgreSQL is a relational database. {'topic': 'database'}
With a real embedding model, expect Redis first. The second result and every
distance depend on the model.Distance space l2 (Chroma's default)
Search for the exact text "Redis is an in-memory data store.", k=2
0.0 Redis is an in-memory data store. <- identical text, distance 0
1513.5 PostgreSQL is a relational database.
Search "Which technology stores data in memory?" with fake (meaningless) vectors
React first, Redis second, Python last - mechanics fine, ranking random
filter={"topic": "database"}, k=5
Redis is an in-memory data store.
PostgreSQL is a relational database. - only the two database notes
collection_metadata={"hnsw:space": "cosine"}
exact match 0.0, unrelated note 1.029 - cosine distance = 1 - similarity
OllamaEmbeddings(model="llama3") + add_texts
ResponseError: This server does not support embeddings.from langchain_core.documents import Document
database_notes = vector_store.similarity_search(
"Which technology stores data in memory?",
k=5,
filter={"topic": "database"},
)
vector_store.update_document(
"note-5",
Document(page_content="Redis is an in-memory key-value store.", metadata={"topic": "database"}),
)
vector_store.delete(ids=["note-3"])
print(vector_store.get()["ids"]) # ['note-1', 'note-2', 'note-4', 'note-5']vector_store = Chroma(
collection_name="tech_notes",
embedding_function=embeddings,
persist_directory="./chroma_db", # writes ./chroma_db/chroma.sqlite3
collection_metadata={"hnsw:space": "cosine"}, # optional: cosine distance instead of l2
)
# Process 1: add_texts(...) -> WROTE 2
# Process 2: Chroma(... same directory) -> READ 2, both texts backTip: Use ids from your data - "lesson-12", "postgresql.md#3" - so re-running the indexer updates records instead of duplicating them.
Watch out: The space is set when the collection is created. To switch an existing collection from l2 to cosine, create a new collection and re-index.
Chroma at a glance
Chroma(...)Open or create a collection.
Chroma(collection_name="tech_notes", embedding_function=embeddings)
add_textsEmbed and store; returns the ids.
add_texts(texts=..., metadatas=..., ids=...)
getInspect records; vectors only if included.
get(ids=[...], include=["embeddings"])
similarity_searchTop k Documents for a question.
similarity_search(query, k=2)
with_score(Document, distance) pairs - lower is closer.
similarity_search_with_score(query, k=2)
filterRestrict by metadata before ranking.
filter={"topic": "database"}update / deleteChange single records by id.
update_document(id, doc) / delete(ids=[...])
persist_directoryKeep the collection on disk.
persist_directory="./chroma_db"
Try it yourself
The code does not change. Swap the content string and the program does something else entirely.
“Index the five notes, then print count() and get()["ids"].”
“Run add_texts twice without ids and count; then twice with ids and count.”
“Search "store data" with filter={"topic": "database"} and then without it.”
“Index into ./chroma_db, stop the program, open the collection again and search.”
What usually goes wrong
Indexing fails on the first add_texts. Use an embedding model.
✗ OllamaEmbeddings(model="llama3.1")✓ OllamaEmbeddings(model="nomic-embed-text")Every run adds another copy of every note - our count went from 5 to 10.
✗ vector_store.add_texts(texts)✓ vector_store.add_texts(texts, ids=ids)Chroma’s with_score returns a distance. 0.0 is the best possible result, not the worst.
Old vectors and new queries would live in different spaces. Re-index everything with the new model.
If one store serves several tenants, a forgotten filter leaks documents. Apply it in code on every query.
✗ similarity_search(query, k=4)✓ similarity_search(query, k=4, filter={"tenant_id": tenant_id})Key points
- A vector store keeps id, text, metadata and vector - and an index for nearest-neighbour search.
- add_texts() embeds and stores; similarity_search() embeds the question and returns Documents.
- Documents and queries must use the same embedding model.
- Ids make re-indexing an update instead of a duplicate.
- get() shows what got stored; vectors only with include=["embeddings"].
- Chroma’s score is a distance: lower is closer.
- Metadata filters restrict the search - including by tenant.
Quick check before you move on
Quiz
- 1.
You run add_texts(texts) twice without ids. How many records are there?
- 2.
similarity_search_with_score returns 0.0 for one result. Good or bad?
- 3.
vector_store.get()["embeddings"] is None. Did indexing fail?
- 4.
How do you search only the database notes?
Interview questions
What is a vector store?
A system that stores vector representations with their original content and metadata, and retrieves the nearest ones to a query vector using an index.
Why do you need a vector store after generating embeddings?
To persist the vectors once and search them efficiently for every future query, with metadata filters, updates and deletes.
What happens during a similarity search?
The query is embedded with the same model, the index finds the nearest stored vectors, and their documents and metadata are returned, ranked by distance.
What is metadata used for?
Filtering, source tracking, authorization and tenant isolation - anything you need to know about a record besides its meaning.
Can PostgreSQL do vector search?
Yes, with the pgvector extension - vectors in a normal table, alongside your other data. Chroma is a simple way to learn the same concepts.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...
AI
System Design
Backend
- GraphQL8 modules · 69 lessons planned
- Core Python13 modules · 75 lessons planned
- FastAPI5 sections · 20 lessons
- Node.js14 modules · 206 lessons planned
- Node.js Performance7 chapters · 36 topics
- Event Loop Lifecycle6 phases · 3 scenarios
- Docker & Containerization11 modules · 144 lessons planned
- AWS for Developers14 modules · 219 lessons planned
- CI/CD & DevOps Automation10 modules · 134 lessons planned