← Back to Agentic AI map
Lesson 2.9 · Building Agents with LangChain

Ollama Embeddings

Turn text into vectors with OllamaEmbeddings and compare them with cosine similarity: which of three sentences is closest in meaning?

embeddings

What you will be able to do

  • Explain what an embedding is, and how it differs from an LLM’s output
  • Picture embeddings as positions in a meaning space
  • Compute cosine similarity and read what the score means
  • Use OllamaEmbeddings: embed_query and embed_documents
  • Rank sentences by meaning - a tiny semantic search
  • Know the limits: an embedding model is required, and similarity is not truth

The idea, in plain English

An LLM turns text into text. An embedding model turns text into a list of numbers - a vector - built so that texts with similar meaning get similar vectors. "I like Python programming" and "I enjoy writing computer programs in Python" share few words but should land close together; "The weather is rainy today" should land far away.

Once meaning is a position, comparing meaning is geometry. Cosine similarity measures the angle between two vectors: near 1 when they point the same way, near 0 when they are unrelated. Rank documents by that score and you have semantic search - the foundation of vector stores (2.10), retrievers (2.11) and RAG (2.12).

In LangChain, OllamaEmbeddings(model=...) gives you embed_query() for a question and embed_documents() for a list of texts. It needs an embedding model pulled into Ollama, such as nomic-embed-text. Chat models do not do this job: asked for an embedding, our llama3 and gemma3 both returned "This server does not support embeddings."

No embedding model is installed on the machine these lessons were checked on, so the model’s similarity scores below are illustrative and labelled as such. The cosine similarity code, the hand-made vectors, the keyword comparison and the LangChain calls were all run.

Worked example: Which of these three sentences is closest in meaning?

request flowWhich sentence is closest?step 1 / 3

1 - Four texts

One query, three candidates. A person sees at once that two are about programming and one is about the weather.

query
I like Python programming.
A
Python is useful for building software.
B
I enjoy writing computer programs in Python.
C
The weather is rainy today.

A query and three sentences, in a hand-made 3-D meaning space: [programming, food, weather]. The scores are computed from these vectors.

Text in, numbers out

An embedding model reads text and returns a fixed-length list of numbers. Every text gets a vector of the same length, whatever its size, so any two can be compared. The individual numbers mean nothing to a person - dimension 17 is not "about food" - only the relationships between vectors do.

So an embedding model is not asked to answer anything. It produces a representation for your code to compare, store and search.

LLM or embedding model
LLMText -> text. "Explain Docker" -> "Docker is a platform for..."
Embedding modelText -> vector. "Docker is a container platform" -> [0.23, -0.71, ...]

A meaning space

Picture a map where similar things sit near each other: Python, JavaScript and C++ in one region, pizza, pasta and biryani in another. An embedding gives each text a position on such a map - with hundreds of dimensions instead of two.

The simulation above uses three dimensions you can read - programming, food, weather - so you can see why the scores come out as they do. A real model learns its dimensions from data, and they are not labelled.

Cosine similarity

Cosine similarity is the dot product of two vectors divided by the product of their lengths: A · B / (|A| |B|). It depends only on the angle between them. We checked the edges: two arrows pointing almost the same way scored 0.994, perpendicular arrows 0.000, opposite arrows -1.000.

Length does not matter: [1, 2, 3] and [2, 4, 6] scored exactly 1.0. That is why cosine is the usual choice for text - a longer document does not get a bigger vector score just for being longer.

Reading a cosine score
1.0Same direction - the closest possible.
around 0Unrelated directions.
-1.0Opposite directions.
In practiceCompare scores from the same model; what counts as "high" differs between models.

OllamaEmbeddings

from langchain_ollama import OllamaEmbeddings, then embeddings = OllamaEmbeddings(model="nomic-embed-text"). embed_query(text) returns one vector for a search query; embed_documents(texts) returns one vector per text. Keep the model fixed: vectors from different models live in different spaces and cannot be compared.

The model must be an embedding model, pulled first with ollama pull nomic-embed-text. With a model that is not installed, Ollama answers "model “nomic-embed-text” not found, try pulling it first". With a chat model like llama3 (Ollama 0.34), "This server does not support embeddings."

Keyword search vs semantic search

Keyword search compares words. We compared the words of "How can I change my password?" and "Steps for updating your login credentials." - after removing words like how, my and for, they share none. A keyword search finds nothing. "How do I reset my password?" against "Changing account credentials": also none.

"I want to learn how to improve database performance" and "PostgreSQL Database Optimization" share one word, database - which they would also share with "Database Backups for Beginners". Semantic search compares meaning instead, which is what an embedding model is built to capture. Keyword search is still useful - for exact names, codes and IDs it is often better - and real systems frequently use both.

The meaning comes from the model

We ran the whole lesson’s code with LangChain’s DeterministicFakeEmbedding, which turns each text into a random-looking 768-number vector. Everything worked - 768 dimensions, three document vectors, cosine scores, a sorted list. And the scores were meaningless: 0.103, 0.019 and 0.002 - the weather sentence happened to come last, but none of the three was any closer to "I like Python" than random vectors are to each other, and the paraphrase scored lower than the weaker match.

That is the useful lesson. The code is a few lines of arithmetic; the quality of semantic search is entirely the quality of the embedding model. Same text, same model gives the same vector (the fake scored exactly 1.0 against itself); a good model also puts different wordings of the same idea close together.

Similarity is not truth

"The Earth orbits the Sun." and "The Sun orbits the Earth." use the same words about the same topic, and an embedding model is likely to place them close together. One is false. Embeddings measure closeness of meaning and topic, not correctness.

In RAG this matters: retrieval finds text about the question, not text that is right. What you put in the document collection is what the model will be shown.

Watch out: High similarity means "about the same thing", not "true", and not even "agrees".

Where this goes next

Looping over every document works for three sentences. For a hundred thousand, you store the vectors once and search them with an index - a vector store, Lesson 2.10. A retriever (2.11) wraps that search, and RAG (2.12) gives the retrieved text to the LLM.

Five things not to mix up
Conversation memoryThe recent messages of one conversation (Lesson 2.8).
EmbeddingText turned into a vector.
Vector storeStores vectors and finds the nearest ones (2.10).
RetrieverReturns relevant documents for a question (2.11).
RAGRetrieved text plus the question, sent to the LLM (2.12).

Step-by-step code

Which sentence is closest in meaning?
# ollama pull nomic-embed-text import numpy as np from langchain_ollama import OllamaEmbeddings embeddings = OllamaEmbeddings(model="nomic-embed-text") query = "I like Python programming." sentences = [ "Python is useful for building software.", "I enjoy writing computer programs in Python.", "The weather is rainy today.", ] query_vector = embeddings.embed_query(query) # one vector sentence_vectors = embeddings.embed_documents(sentences) # three vectors def cosine_similarity(a, b): a = np.array(a) b = np.array(b) return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)) results = [] for sentence, vector in zip(sentences, sentence_vectors): results.append((cosine_similarity(query_vector, vector), sentence)) results.sort(reverse=True) for score, sentence in results: print(f"{score:.2f} {sentence}")
Output - illustrative, not from a run
0.93 I enjoy writing computer programs in Python. 0.82 Python is useful for building software. 0.12 The weather is rainy today. Expect the same order from a good embedding model; the exact numbers depend on the model.
Cosine similarity on vectors you can read
# [programming, food, weather] python = [0.9, 0.1, 0.0] # I like Python programming. software = [0.8, 0.0, 0.1] # Python is useful for building software. programs = [0.9, 0.2, 0.0] # I enjoy writing computer programs in Python. weather = [0.0, 0.1, 0.9] # The weather is rainy today. print(cosine_similarity(python, programs)) # 0.994 print(cosine_similarity(python, software)) # 0.986 print(cosine_similarity(python, weather)) # 0.012 print(cosine_similarity([1, 1], [1, 0.8])) # 0.994 - almost the same direction print(cosine_similarity([0, 1], [1, 0])) # 0.0 - perpendicular print(cosine_similarity([1, 0], [-1, 0])) # -1.0 - opposite print(cosine_similarity([1, 2, 3], [2, 4, 6])) # 1.0 - length does not matter
Keyword overlap - checked
Shared words, ignoring how / do / I / my / can / your / for / to / the / a / of / is / want / learn "How can I change my password?" vs "Steps for updating your login credentials." -> none "How do I reset my password?" vs "Changing account credentials" -> none "How do I reset my password?" vs "Steps for updating your login password" -> password "I want to learn how to improve database vs "PostgreSQL Database Optimization" -> database performance."
The same code with fake embeddings
from langchain_core.embeddings import DeterministicFakeEmbedding embeddings = DeterministicFakeEmbedding(size=768) # random-looking vectors, no meaning query_vector = embeddings.embed_query(query) sentence_vectors = embeddings.embed_documents(sentences) print(len(query_vector), len(sentence_vectors)) # 768 3 # 0.10 Python is useful for building software. # 0.02 I enjoy writing computer programs in Python. # 0.00 The weather is rainy today. # - the code works; the ranking means nothing. The meaning comes from the model.
Models that cannot embed - from real runs
OllamaEmbeddings(model="llama3").embed_query("hi") ResponseError: This server does not support embeddings. Start it with --embeddings (status code: 501) OllamaEmbeddings(model="nomic-embed-text").embed_query("hi") # before ollama pull ResponseError: model "nomic-embed-text" not found, try pulling it first (status code: 404)

Tip: Embed the documents once and keep the vectors. Re-embedding the same text with the same model gives the same vector - recomputing it on every search is wasted work, which is exactly what a vector store saves you.

Embeddings at a glance

OllamaEmbeddings

LangChain wrapper for an Ollama embedding model.

OllamaEmbeddings(model="nomic-embed-text")
embed_query

One vector for a search query.

embeddings.embed_query("How do I learn Python?")
embed_documents

One vector per text.

embeddings.embed_documents(texts)
Cosine similarity

A · B / (|A| |B|): 1 same direction, 0 unrelated.

np.dot(a, b) / (norm(a) * norm(b))
ollama pull

Install an embedding model first.

ollama pull nomic-embed-text
DeterministicFakeEmbedding

Test the plumbing without a model.

DeterministicFakeEmbedding(size=768)

Try it yourself

The code does not change. Swap the content string and the program does something else entirely.

Run the example

“Pull nomic-embed-text and run the three-sentence comparison. Is the order what you expected?”

Paraphrase

“Score "How do I reset my password?" against "Steps for updating your login credentials" and "How to make vegetable pasta".”

Earth and Sun

“Score "The Earth orbits the Sun." against "The Sun orbits the Earth." What does the number tell you - and what does it not?”

Course search

“Rank five course titles for "I want to improve database performance".”

What usually goes wrong

Using a chat model for embeddings

Embeddings need an embedding model. llama3 and gemma3 both refused.

✗ OllamaEmbeddings(model="llama3")
✓ OllamaEmbeddings(model="nomic-embed-text")
Mixing embedding models

Vectors from different models are in different spaces. Embed the query with the same model as the documents.

Reading meaning into single numbers

A dimension has no human meaning on its own. Only comparisons between vectors tell you anything.

Carrying a threshold between models

"Above 0.8 is relevant" for one model can be wrong for another. Pick thresholds per model, by testing.

Treating similarity as truth

Close in meaning is not correct. "The Sun orbits the Earth" is about the same topic as the true sentence.

Key points

  • An LLM turns text into text; an embedding model turns text into a vector.
  • Similar meanings get vectors that point in similar directions.
  • Cosine similarity: 1 same direction, 0 unrelated - and length does not matter.
  • OllamaEmbeddings: embed_query for a question, embed_documents for texts.
  • You need an embedding model such as nomic-embed-text; chat models refuse.
  • Semantic search finds matches keyword search misses - with no shared words at all.
  • Similarity is not truth.

Quick check before you move on

What does an embedding model produce?
A vector - a fixed-length list of numbers representing the text.
Why convert text into vectors?
So meaning can be compared with arithmetic: close vectors, similar meaning.
What is cosine similarity?
A score from the angle between two vectors: near 1 when they point the same way, near 0 when unrelated.
Which two are most similar: "How do I reset my password?", "Steps for changing login credentials", "How to make vegetable pasta"?
The first two - same meaning, almost no shared words.
Does high similarity prove a statement is true?
No. It measures closeness of meaning, not correctness.
Where do large numbers of embeddings go?
Into a vector store, which indexes them for fast nearest-neighbour search - Lesson 2.10.

Quiz

  1. 1.

    What is the cosine similarity of [1, 2, 3] and [2, 4, 6]?

  2. 2.

    You call OllamaEmbeddings(model="llama3").embed_query(...). What happens?

  3. 3.

    With DeterministicFakeEmbedding the code ran but the ranking was wrong. What does that show?

  4. 4.

    Why might keyword search miss "Steps for updating your login credentials" for "How can I change my password?"

Interview questions

What is an embedding?

A numerical vector representation of data, designed so related inputs are close together in the vector space.

What is the difference between an LLM and an embedding model?

An LLM generates text from text. An embedding model maps text to a vector for comparison, search and clustering.

Why are embeddings useful for RAG?

They let you compare the question with document chunks by meaning and retrieve the relevant ones to give the LLM as context.

What is semantic search?

Retrieval by similarity of meaning - comparing embeddings - rather than by exact keyword matches.

Why cosine similarity rather than raw distance?

It compares direction and ignores vector length, so it measures how alike two texts are rather than how long they are.

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...