Ollama Embeddings
Turn text into vectors with OllamaEmbeddings and compare them with cosine similarity: which of three sentences is closest in meaning?
What you will be able to do
- Explain what an embedding is, and how it differs from an LLM’s output
- Picture embeddings as positions in a meaning space
- Compute cosine similarity and read what the score means
- Use OllamaEmbeddings: embed_query and embed_documents
- Rank sentences by meaning - a tiny semantic search
- Know the limits: an embedding model is required, and similarity is not truth
The idea, in plain English
An LLM turns text into text. An embedding model turns text into a list of numbers - a vector - built so that texts with similar meaning get similar vectors. "I like Python programming" and "I enjoy writing computer programs in Python" share few words but should land close together; "The weather is rainy today" should land far away.
Once meaning is a position, comparing meaning is geometry. Cosine similarity measures the angle between two vectors: near 1 when they point the same way, near 0 when they are unrelated. Rank documents by that score and you have semantic search - the foundation of vector stores (2.10), retrievers (2.11) and RAG (2.12).
In LangChain, OllamaEmbeddings(model=...) gives you embed_query() for a question and embed_documents() for a list of texts. It needs an embedding model pulled into Ollama, such as nomic-embed-text. Chat models do not do this job: asked for an embedding, our llama3 and gemma3 both returned "This server does not support embeddings."
No embedding model is installed on the machine these lessons were checked on, so the model’s similarity scores below are illustrative and labelled as such. The cosine similarity code, the hand-made vectors, the keyword comparison and the LangChain calls were all run.
Worked example: Which of these three sentences is closest in meaning?
1 - Four texts
One query, three candidates. A person sees at once that two are about programming and one is about the weather.
A query and three sentences, in a hand-made 3-D meaning space: [programming, food, weather]. The scores are computed from these vectors.
Text in, numbers out
An embedding model reads text and returns a fixed-length list of numbers. Every text gets a vector of the same length, whatever its size, so any two can be compared. The individual numbers mean nothing to a person - dimension 17 is not "about food" - only the relationships between vectors do.
So an embedding model is not asked to answer anything. It produces a representation for your code to compare, store and search.
LLMText -> text. "Explain Docker" -> "Docker is a platform for..."Embedding modelText -> vector. "Docker is a container platform" -> [0.23, -0.71, ...]A meaning space
Picture a map where similar things sit near each other: Python, JavaScript and C++ in one region, pizza, pasta and biryani in another. An embedding gives each text a position on such a map - with hundreds of dimensions instead of two.
The simulation above uses three dimensions you can read - programming, food, weather - so you can see why the scores come out as they do. A real model learns its dimensions from data, and they are not labelled.
Cosine similarity
Cosine similarity is the dot product of two vectors divided by the product of their lengths: A · B / (|A| |B|). It depends only on the angle between them. We checked the edges: two arrows pointing almost the same way scored 0.994, perpendicular arrows 0.000, opposite arrows -1.000.
Length does not matter: [1, 2, 3] and [2, 4, 6] scored exactly 1.0. That is why cosine is the usual choice for text - a longer document does not get a bigger vector score just for being longer.
1.0Same direction - the closest possible.around 0Unrelated directions.-1.0Opposite directions.In practiceCompare scores from the same model; what counts as "high" differs between models.OllamaEmbeddings
from langchain_ollama import OllamaEmbeddings, then embeddings = OllamaEmbeddings(model="nomic-embed-text"). embed_query(text) returns one vector for a search query; embed_documents(texts) returns one vector per text. Keep the model fixed: vectors from different models live in different spaces and cannot be compared.
The model must be an embedding model, pulled first with ollama pull nomic-embed-text. With a model that is not installed, Ollama answers "model “nomic-embed-text” not found, try pulling it first". With a chat model like llama3 (Ollama 0.34), "This server does not support embeddings."
Keyword search vs semantic search
Keyword search compares words. We compared the words of "How can I change my password?" and "Steps for updating your login credentials." - after removing words like how, my and for, they share none. A keyword search finds nothing. "How do I reset my password?" against "Changing account credentials": also none.
"I want to learn how to improve database performance" and "PostgreSQL Database Optimization" share one word, database - which they would also share with "Database Backups for Beginners". Semantic search compares meaning instead, which is what an embedding model is built to capture. Keyword search is still useful - for exact names, codes and IDs it is often better - and real systems frequently use both.
The meaning comes from the model
We ran the whole lesson’s code with LangChain’s DeterministicFakeEmbedding, which turns each text into a random-looking 768-number vector. Everything worked - 768 dimensions, three document vectors, cosine scores, a sorted list. And the scores were meaningless: 0.103, 0.019 and 0.002 - the weather sentence happened to come last, but none of the three was any closer to "I like Python" than random vectors are to each other, and the paraphrase scored lower than the weaker match.
That is the useful lesson. The code is a few lines of arithmetic; the quality of semantic search is entirely the quality of the embedding model. Same text, same model gives the same vector (the fake scored exactly 1.0 against itself); a good model also puts different wordings of the same idea close together.
Similarity is not truth
"The Earth orbits the Sun." and "The Sun orbits the Earth." use the same words about the same topic, and an embedding model is likely to place them close together. One is false. Embeddings measure closeness of meaning and topic, not correctness.
In RAG this matters: retrieval finds text about the question, not text that is right. What you put in the document collection is what the model will be shown.
Watch out: High similarity means "about the same thing", not "true", and not even "agrees".
Where this goes next
Looping over every document works for three sentences. For a hundred thousand, you store the vectors once and search them with an index - a vector store, Lesson 2.10. A retriever (2.11) wraps that search, and RAG (2.12) gives the retrieved text to the LLM.
Conversation memoryThe recent messages of one conversation (Lesson 2.8).EmbeddingText turned into a vector.Vector storeStores vectors and finds the nearest ones (2.10).RetrieverReturns relevant documents for a question (2.11).RAGRetrieved text plus the question, sent to the LLM (2.12).Step-by-step code
# ollama pull nomic-embed-text
import numpy as np
from langchain_ollama import OllamaEmbeddings
embeddings = OllamaEmbeddings(model="nomic-embed-text")
query = "I like Python programming."
sentences = [
"Python is useful for building software.",
"I enjoy writing computer programs in Python.",
"The weather is rainy today.",
]
query_vector = embeddings.embed_query(query) # one vector
sentence_vectors = embeddings.embed_documents(sentences) # three vectors
def cosine_similarity(a, b):
a = np.array(a)
b = np.array(b)
return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
results = []
for sentence, vector in zip(sentences, sentence_vectors):
results.append((cosine_similarity(query_vector, vector), sentence))
results.sort(reverse=True)
for score, sentence in results:
print(f"{score:.2f} {sentence}")0.93 I enjoy writing computer programs in Python.
0.82 Python is useful for building software.
0.12 The weather is rainy today.
Expect the same order from a good embedding model; the exact numbers depend on the model.# [programming, food, weather]
python = [0.9, 0.1, 0.0] # I like Python programming.
software = [0.8, 0.0, 0.1] # Python is useful for building software.
programs = [0.9, 0.2, 0.0] # I enjoy writing computer programs in Python.
weather = [0.0, 0.1, 0.9] # The weather is rainy today.
print(cosine_similarity(python, programs)) # 0.994
print(cosine_similarity(python, software)) # 0.986
print(cosine_similarity(python, weather)) # 0.012
print(cosine_similarity([1, 1], [1, 0.8])) # 0.994 - almost the same direction
print(cosine_similarity([0, 1], [1, 0])) # 0.0 - perpendicular
print(cosine_similarity([1, 0], [-1, 0])) # -1.0 - opposite
print(cosine_similarity([1, 2, 3], [2, 4, 6])) # 1.0 - length does not matterShared words, ignoring how / do / I / my / can / your / for / to / the / a / of / is / want / learn
"How can I change my password?" vs "Steps for updating your login credentials." -> none
"How do I reset my password?" vs "Changing account credentials" -> none
"How do I reset my password?" vs "Steps for updating your login password" -> password
"I want to learn how to improve database vs "PostgreSQL Database Optimization" -> database
performance."from langchain_core.embeddings import DeterministicFakeEmbedding
embeddings = DeterministicFakeEmbedding(size=768) # random-looking vectors, no meaning
query_vector = embeddings.embed_query(query)
sentence_vectors = embeddings.embed_documents(sentences)
print(len(query_vector), len(sentence_vectors)) # 768 3
# 0.10 Python is useful for building software.
# 0.02 I enjoy writing computer programs in Python.
# 0.00 The weather is rainy today.
# - the code works; the ranking means nothing. The meaning comes from the model.OllamaEmbeddings(model="llama3").embed_query("hi")
ResponseError: This server does not support embeddings. Start it with --embeddings (status code: 501)
OllamaEmbeddings(model="nomic-embed-text").embed_query("hi") # before ollama pull
ResponseError: model "nomic-embed-text" not found, try pulling it first (status code: 404)Tip: Embed the documents once and keep the vectors. Re-embedding the same text with the same model gives the same vector - recomputing it on every search is wasted work, which is exactly what a vector store saves you.
Embeddings at a glance
OllamaEmbeddingsLangChain wrapper for an Ollama embedding model.
OllamaEmbeddings(model="nomic-embed-text")
embed_queryOne vector for a search query.
embeddings.embed_query("How do I learn Python?")embed_documentsOne vector per text.
embeddings.embed_documents(texts)
Cosine similarityA · B / (|A| |B|): 1 same direction, 0 unrelated.
np.dot(a, b) / (norm(a) * norm(b))
ollama pullInstall an embedding model first.
ollama pull nomic-embed-text
DeterministicFakeEmbeddingTest the plumbing without a model.
DeterministicFakeEmbedding(size=768)
Try it yourself
The code does not change. Swap the content string and the program does something else entirely.
“Pull nomic-embed-text and run the three-sentence comparison. Is the order what you expected?”
“Score "How do I reset my password?" against "Steps for updating your login credentials" and "How to make vegetable pasta".”
“Score "The Earth orbits the Sun." against "The Sun orbits the Earth." What does the number tell you - and what does it not?”
“Rank five course titles for "I want to improve database performance".”
What usually goes wrong
Embeddings need an embedding model. llama3 and gemma3 both refused.
✗ OllamaEmbeddings(model="llama3")✓ OllamaEmbeddings(model="nomic-embed-text")Vectors from different models are in different spaces. Embed the query with the same model as the documents.
A dimension has no human meaning on its own. Only comparisons between vectors tell you anything.
"Above 0.8 is relevant" for one model can be wrong for another. Pick thresholds per model, by testing.
Close in meaning is not correct. "The Sun orbits the Earth" is about the same topic as the true sentence.
Key points
- An LLM turns text into text; an embedding model turns text into a vector.
- Similar meanings get vectors that point in similar directions.
- Cosine similarity: 1 same direction, 0 unrelated - and length does not matter.
- OllamaEmbeddings: embed_query for a question, embed_documents for texts.
- You need an embedding model such as nomic-embed-text; chat models refuse.
- Semantic search finds matches keyword search misses - with no shared words at all.
- Similarity is not truth.
Quick check before you move on
Quiz
- 1.
What is the cosine similarity of [1, 2, 3] and [2, 4, 6]?
- 2.
You call OllamaEmbeddings(model="llama3").embed_query(...). What happens?
- 3.
With DeterministicFakeEmbedding the code ran but the ranking was wrong. What does that show?
- 4.
Why might keyword search miss "Steps for updating your login credentials" for "How can I change my password?"
Interview questions
What is an embedding?
A numerical vector representation of data, designed so related inputs are close together in the vector space.
What is the difference between an LLM and an embedding model?
An LLM generates text from text. An embedding model maps text to a vector for comparison, search and clustering.
Why are embeddings useful for RAG?
They let you compare the question with document chunks by meaning and retrieve the relevant ones to give the LLM as context.
What is semantic search?
Retrieval by similarity of meaning - comparing embeddings - rather than by exact keyword matches.
Why cosine similarity rather than raw distance?
It compares direction and ignores vector length, so it measures how alike two texts are rather than how long they are.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...
AI
System Design
Backend
- GraphQL8 modules · 69 lessons planned
- Core Python13 modules · 75 lessons planned
- FastAPI5 sections · 20 lessons
- Node.js14 modules · 206 lessons planned
- Node.js Performance7 chapters · 36 topics
- Event Loop Lifecycle6 phases · 3 scenarios
- Docker & Containerization11 modules · 144 lessons planned
- AWS for Developers14 modules · 219 lessons planned
- CI/CD & DevOps Automation10 modules · 134 lessons planned