← Back to Agentic AI map
Lesson 2.10 · Building Agents with LangChain

Vector Stores

Index five short notes into Chroma, inspect exactly what got stored, and search them by meaning.

vector stores

What you will be able to do

  • Explain what a vector store keeps, and why it keeps the text as well as the vector
  • Index documents into Chroma with ids and metadata
  • Inspect what got stored - including the vectors
  • Search with similarity_search, and read what the score means
  • Filter by metadata, update and delete by id, and persist to disk
  • Place Chroma among the alternatives, such as PostgreSQL with pgvector

The idea, in plain English

Lesson 2.9 ended with vectors in a Python list and a loop that compared the query with every one. A vector store does that job properly. For each record it keeps an id, the original text, metadata, and the vector - plus an index that finds the nearest vectors to a query without comparing against every record.

Chroma is the vector store this course uses. Through LangChain you create a Chroma collection with an embedding function, call add_texts() - which embeds each text with that function and stores the lot - and later call similarity_search(question, k=2), which embeds the question the same way and returns the two closest Documents.

The vector finds the record; the text is what you use. After a search you want "Redis is an in-memory data store." to give to an LLM, not 768 numbers. That is why the store keeps both, and why metadata - topic, source, tenant - travels with them.

No embedding model is installed on the machine these lessons were checked on (see Lesson 2.9). Chroma’s behaviour - storing, inspecting, ids, metadata, filters, distances, persistence - was all run, with LangChain’s fake embeddings in place of the model. Rankings for real questions need a real model, so those are illustrative and labelled.

Worked example: Index 5 short notes into Chroma.

request flowIndex five notes, then search themstep 1 / 4

1 - Index the notes

add_texts() embeds each note with the embedding function and stores id, text, metadata and vector together.

records
5
ids
note-1 ... note-5
returns
the ids
you call embed_documents
no - Chroma does

add_texts() on the way in, similarity_search() on the way out - both go through the same embedding model.

What a vector store keeps

Four things per record: an id, the original text (Chroma calls it the document), metadata, and the embedding vector. And one thing for the collection: an index that finds near neighbours quickly. Chroma’s index is HNSW, a graph that jumps towards close vectors instead of measuring every one.

A normal database answers "which rows match this condition?" - WHERE title ILIKE ‘%python%’. A vector store answers "which records are closest in meaning to this?" - and "How can I run applications in isolated containers?" can find the Docker note without sharing the word Docker.

Indexing with Chroma

pip install langchain-chroma langchain-ollama. Create the store with Chroma(collection_name="tech_notes", embedding_function=embeddings) - a collection is a named set of records - then call add_texts(texts=..., metadatas=..., ids=...). Indexing means exactly this: embed each text, store the vector with its text and metadata, make it searchable.

You do not call embed_documents() yourself. The store calls the embedding function you gave it, for documents now and for queries later - which is what guarantees both are in the same space. If that function is a chat model, indexing fails at the first add_texts: with llama3, "This server does not support embeddings."

Ids decide whether you update or duplicate

add_texts() returns the ids it stored. Give your own - note-1 to note-5 - and adding the same five notes again replaced them: the count stayed 5. Leave ids out and Chroma generates random UUIDs, so running the same indexing script twice stored every note twice: the count went to 10.

With ids you can also change single records: vector_store.delete(ids=["note-3"]) left four, and update_document("note-5", ...) replaced the Redis note’s text (and re-embedded it). Use ids that come from your data - a file path, a lesson id - so re-indexing is safe.

Ids in practice - checked
add_texts(texts, ids=ids) twiceCount stays 5 - the second call replaces.
add_texts(texts) twiceCount 10 - random ids, every note stored twice.
delete(ids=["note-3"])Count 4: note-1, note-2, note-4, note-5.
update_document("note-5", doc)New text, new vector, same id.

Inspecting what got stored

Do not assume indexing worked - look. vector_store._collection.count() gave 5. vector_store.get() returned the ids, the five documents and their metadatas - and "embeddings": None. Chroma leaves the vectors out of get() unless you ask: get(ids=["note-5"], include=["documents", "metadatas", "embeddings"]) returned the Redis note with its 768 numbers.

That is also where you catch the classic indexing bugs: duplicates from missing ids, empty or truncated documents, missing metadata, the wrong collection.

Searching, and what the score means

similarity_search(query, k=2) returns the two nearest records as Documents: page_content is the text, metadata the dict, id the id. similarity_search_with_score returns (Document, score) pairs - and in Chroma that score is a distance. The default space is l2: searching for the exact text of the Redis note returned it with 0.0, the next note far behind. Lower is closer.

Create the collection with collection_metadata={"hnsw:space": "cosine"} and the score becomes cosine distance - 1 minus cosine similarity: exact match 0.0, unrelated around 1. similarity_search_with_relevance_scores converts distances to "higher is better", but only makes sense for normalized vectors; with our fake ones it produced -1069 and a warning. Always check which kind of number you have.

Three kinds of score
similarity_search_with_score, l2Euclidean distance. 0 = identical, lower = closer.
similarity_search_with_score, cosine1 - cosine similarity. 0 = identical, about 1 = unrelated.
similarity_search_with_relevance_scoresConverted to 0-1, higher = closer - assumes normalized vectors.

Metadata and filters

Metadata is a dict per record. With topic on each note, similarity_search(query, k=5, filter={"topic": "database"}) returned only the PostgreSQL and Redis notes - the filter is applied before ranking, so nothing outside it can appear.

In a real application metadata carries source, document type, dates - and, for a multi-tenant platform, tenant_id and course_id. Filtering on tenant_id is how one store can serve many academies without one tenant’s documents showing up in another’s answers. Treat that filter as a security rule: apply it in your code on every query, never as an option.

In memory or on disk

Without a persist_directory, Chroma keeps the collection in memory - and shares it within the process: a second Chroma(collection_name="tech_notes") saw the first one’s records. Re-running a notebook cell that indexes without ids piles up duplicates this way.

Chroma(..., persist_directory="./chroma_db") writes a chroma.sqlite3 file. We indexed two notes in one Python process and read them back in a second: count 2, both texts. Index once, search many times - you do not re-embed your documents on every start.

Chroma is one option

Chroma is easy to start with: a pip install, no server. It is not the only choice. PostgreSQL with the pgvector extension stores vectors in a normal table, next to users, courses and lessons, with SQL filters and joins - attractive if PostgreSQL is already your database. Dedicated services exist too.

The concepts carry over unchanged: records with text, metadata and a vector; the same embedding model for documents and queries; nearest-neighbour search with a k and a filter.

Step-by-step code

Index five notes into Chroma
# pip install langchain-chroma langchain-ollama # ollama pull nomic-embed-text from langchain_chroma import Chroma from langchain_ollama import OllamaEmbeddings texts = [ "Python is a programming language.", "PostgreSQL is a relational database.", "Docker packages applications into containers.", "React is a JavaScript UI library.", "Redis is an in-memory data store.", ] metadatas = [ {"topic": "programming"}, {"topic": "database"}, {"topic": "devops"}, {"topic": "frontend"}, {"topic": "database"}, ] ids = ["note-1", "note-2", "note-3", "note-4", "note-5"] embeddings = OllamaEmbeddings(model="nomic-embed-text") vector_store = Chroma( collection_name="tech_notes", embedding_function=embeddings, ) print(vector_store.add_texts(texts=texts, metadatas=metadatas, ids=ids)) # ['note-1', 'note-2', 'note-3', 'note-4', 'note-5']
Inspect what got stored
print(vector_store._collection.count()) # 5 stored = vector_store.get() print(stored["ids"]) # ['note-1', 'note-2', 'note-3', 'note-4', 'note-5'] print(stored["documents"]) # ['Python is a programming language.', 'PostgreSQL is a relational database.', ...] print(stored["metadatas"]) # [{'topic': 'programming'}, {'topic': 'database'}, ...] print(stored["embeddings"]) # None - vectors are left out unless you ask note = vector_store.get(ids=["note-5"], include=["documents", "metadatas", "embeddings"]) print(note["documents"]) # ['Redis is an in-memory data store.'] print(len(note["embeddings"][0])) # the vector length, e.g. 768
Search by meaning
results = vector_store.similarity_search("Which technology stores data in memory?", k=2) for document in results: print(document.id, document.page_content, document.metadata) for document, score in vector_store.similarity_search_with_score( "Which technology stores data in memory?", k=2 ): print(round(score, 3), document.page_content) # a distance: lower = closer
Output - illustrative, not from a run
note-5 Redis is an in-memory data store. {'topic': 'database'} note-2 PostgreSQL is a relational database. {'topic': 'database'} With a real embedding model, expect Redis first. The second result and every distance depend on the model.
What Chroma did - from real runs with fake embeddings
Distance space l2 (Chroma's default) Search for the exact text "Redis is an in-memory data store.", k=2 0.0 Redis is an in-memory data store. <- identical text, distance 0 1513.5 PostgreSQL is a relational database. Search "Which technology stores data in memory?" with fake (meaningless) vectors React first, Redis second, Python last - mechanics fine, ranking random filter={"topic": "database"}, k=5 Redis is an in-memory data store. PostgreSQL is a relational database. - only the two database notes collection_metadata={"hnsw:space": "cosine"} exact match 0.0, unrelated note 1.029 - cosine distance = 1 - similarity OllamaEmbeddings(model="llama3") + add_texts ResponseError: This server does not support embeddings.
Filter, update, delete
from langchain_core.documents import Document database_notes = vector_store.similarity_search( "Which technology stores data in memory?", k=5, filter={"topic": "database"}, ) vector_store.update_document( "note-5", Document(page_content="Redis is an in-memory key-value store.", metadata={"topic": "database"}), ) vector_store.delete(ids=["note-3"]) print(vector_store.get()["ids"]) # ['note-1', 'note-2', 'note-4', 'note-5']
Persist to disk
vector_store = Chroma( collection_name="tech_notes", embedding_function=embeddings, persist_directory="./chroma_db", # writes ./chroma_db/chroma.sqlite3 collection_metadata={"hnsw:space": "cosine"}, # optional: cosine distance instead of l2 ) # Process 1: add_texts(...) -> WROTE 2 # Process 2: Chroma(... same directory) -> READ 2, both texts back

Tip: Use ids from your data - "lesson-12", "postgresql.md#3" - so re-running the indexer updates records instead of duplicating them.

Watch out: The space is set when the collection is created. To switch an existing collection from l2 to cosine, create a new collection and re-index.

Chroma at a glance

Chroma(...)

Open or create a collection.

Chroma(collection_name="tech_notes", embedding_function=embeddings)
add_texts

Embed and store; returns the ids.

add_texts(texts=..., metadatas=..., ids=...)
get

Inspect records; vectors only if included.

get(ids=[...], include=["embeddings"])
similarity_search

Top k Documents for a question.

similarity_search(query, k=2)
with_score

(Document, distance) pairs - lower is closer.

similarity_search_with_score(query, k=2)
filter

Restrict by metadata before ranking.

filter={"topic": "database"}
update / delete

Change single records by id.

update_document(id, doc) / delete(ids=[...])
persist_directory

Keep the collection on disk.

persist_directory="./chroma_db"

Try it yourself

The code does not change. Swap the content string and the program does something else entirely.

Index and count

“Index the five notes, then print count() and get()["ids"].”

Duplicate on purpose

“Run add_texts twice without ids and count; then twice with ids and count.”

Filter

“Search "store data" with filter={"topic": "database"} and then without it.”

Persist

“Index into ./chroma_db, stop the program, open the collection again and search.”

What usually goes wrong

A chat model as the embedding function

Indexing fails on the first add_texts. Use an embedding model.

✗ OllamaEmbeddings(model="llama3.1")
✓ OllamaEmbeddings(model="nomic-embed-text")
Indexing without ids

Every run adds another copy of every note - our count went from 5 to 10.

✗ vector_store.add_texts(texts)
✓ vector_store.add_texts(texts, ids=ids)
Reading a distance as a similarity

Chroma’s with_score returns a distance. 0.0 is the best possible result, not the worst.

Changing the embedding model after indexing

Old vectors and new queries would live in different spaces. Re-index everything with the new model.

Optional tenant filters

If one store serves several tenants, a forgotten filter leaks documents. Apply it in code on every query.

✗ similarity_search(query, k=4)
✓ similarity_search(query, k=4, filter={"tenant_id": tenant_id})

Key points

  • A vector store keeps id, text, metadata and vector - and an index for nearest-neighbour search.
  • add_texts() embeds and stores; similarity_search() embeds the question and returns Documents.
  • Documents and queries must use the same embedding model.
  • Ids make re-indexing an update instead of a duplicate.
  • get() shows what got stored; vectors only with include=["embeddings"].
  • Chroma’s score is a distance: lower is closer.
  • Metadata filters restrict the search - including by tenant.

Quick check before you move on

What is a vector store?
A store for vectors together with their documents and metadata, with an index for similarity search.
What is Chroma doing in this lesson?
It is the vector store: it embeds the notes through the embedding function, stores them, and searches them.
What does indexing mean?
Embedding each document and storing the vector with its text and metadata so it can be searched.
Why store the original text with the vector?
The vector is for finding; the text is what you return and give to the LLM.
What does k=2 mean in similarity_search(query, k=2)?
Return the two closest records.
Embedding model or vector store - what is the difference?
The embedding model turns text into vectors. The vector store keeps vectors with their documents and searches them.

Quiz

  1. 1.

    You run add_texts(texts) twice without ids. How many records are there?

  2. 2.

    similarity_search_with_score returns 0.0 for one result. Good or bad?

  3. 3.

    vector_store.get()["embeddings"] is None. Did indexing fail?

  4. 4.

    How do you search only the database notes?

Interview questions

What is a vector store?

A system that stores vector representations with their original content and metadata, and retrieves the nearest ones to a query vector using an index.

Why do you need a vector store after generating embeddings?

To persist the vectors once and search them efficiently for every future query, with metadata filters, updates and deletes.

What happens during a similarity search?

The query is embedded with the same model, the index finds the nearest stored vectors, and their documents and metadata are returned, ranked by distance.

What is metadata used for?

Filtering, source tracking, authorization and tenant isolation - anything you need to know about a record besides its meaning.

Can PostgreSQL do vector search?

Yes, with the pgvector extension - vectors in a normal table, alongside your other data. Chroma is a simple way to learn the same concepts.

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...