Modern AI Engineering

Lesson 4.4 · 27 min

Embeddings: Encoding Meaning as Vectors

How can a computer know that “puppy” is closer to “dog” than to “invoice” when, to it, every word is just a string of characters?

In short: An embedding is a list of numbers (a vector) that represents a piece of data, such as a word, a sentence, an image or a product, so that similar things get similar numbers. We measure how close two embeddings are with cosine similarity or distance, which lets computers compare meaning. Embeddings are learned by neural networks from data, power search, RAG, recommendations and clustering, and come with real pitfalls: bias, model mismatch and domain gaps.

The problem: computers cannot compare meaning

Our running example is a help-centre search for an online shop. A customer types “my parcel never arrived”. The best help article is titled “What to do if your delivery is missing”. The two texts share no important words. A plain keyword search finds nothing useful, yet any human sees they mean the same thing.

To a computer, words are just characters or arbitrary ID numbers. If “parcel” is ID 812 and “delivery” is ID 4410, nothing about those numbers says they are related. One older trick, one-hot encoding, gives each word a vector of zeros with a single 1 in its own slot. But every one-hot vector is equally far from every other one, so “parcel” is exactly as similar to “delivery” as it is to “banana”. We need numbers that carry meaning.

Think of it like map coordinates Cities on a map have two numbers: latitude and longitude. Cities that are near each other in the real world have similar numbers, so we can compute which cities are close without knowing anything else about them. An embedding gives every word or sentence coordinates on a “map of meaning”, where nearby points mean similar things.

What are Embeddings?

An embedding is a fixed-length list of numbers, called a vector, that represents one item. The key property: items with similar meaning get vectors that are close together, and unrelated items get vectors that are far apart. Each number is one dimension; the length of the list is the embedding's dimensionality.

The word comes from mathematics: we embed (place) items into a continuous space. Embeddings are also called dense vectors because almost every number is non-zero, unlike sparse one-hot vectors.

A simple example with two numbers. Suppose we describe words by two hand-picked features: “how much is it an animal?” (0 to 1) and “how big is it?” (0 to 1). A cat might be [0.9, 0.2], an elephant [0.9, 0.95], an apple [0.1, 0.2] and a watermelon [0.1, 0.8]. Plot them and animals cluster on one side, fruits on the other, and size spreads them vertically.

Why we need many more than two numbers

Two numbers can capture “animal-ness” and “size”, but meaning has far more aspects: is it alive, is it food, is it formal, is it about money, is it positive, is it a verb, which topic does it belong to, and countless subtler ones. With only two dimensions, “bank” (money) and “bank” (river) and “bunk” would be crammed together with unrelated words. More dimensions give room to keep many kinds of similarity apart at the same time.

Typical embedding sizes (publicly documented)
ModelYearDimensions
word2vec (Google News vectors)2013300
GloVe (common releases)201450–300
BERT-base hidden states2018768
OpenAI text-embedding-3-small20241,536
OpenAI text-embedding-3-large20243,072

More dimensions is not automatically better. Each extra dimension costs storage and search time: one million 1,536-dimension float32 vectors take about 6 GB (1,000,000 × 1,536 × 4 bytes). Many current models let us shorten vectors with a small quality loss, so the right size depends on the budget and the task.

Dimensions are not human-readable In our toy examples each dimension has a name. In learned embeddings, a single dimension rarely means one clear thing; meaning is spread across many dimensions at once. We work with the vectors through distances and similarities, not by reading individual numbers.

How do we measure closeness?

Once items are vectors, “similar meaning” becomes “close vectors”. Three measures are common.

  • Cosine similarity: the angle between vectors. The most common choice for text embeddings because it ignores vector length.
  • Dot product: a · b. Equal to cosine similarity when vectors are normalised to length 1, and cheaper to compute. Many embedding APIs return normalised vectors for this reason.
  • Euclidean distance: ‖a − b‖, the straight-line distance. Smaller means more similar.

Worked example. a = [1, 2], b = [2, 3]. Dot product = 1×2 + 2×3 = 8. ‖a‖ = √5 ≈ 2.236, ‖b‖ = √13 ≈ 3.606. Cosine = 8 / (2.236 × 3.606) ≈ 0.99: almost the same direction. Euclidean distance = √((2−1)² + (3−2)²) = √2 ≈ 1.41.

Pause and think: Vector b = 2 × a (same direction, twice as long). What is their cosine similarity, and is their Euclidean distance zero?

Cosine similarity is exactly 1, because the direction is identical. The Euclidean distance is not zero; it equals ‖a‖. This is why cosine is preferred when only direction (meaning) should matter, not magnitude.

Where do embeddings come from?

Nobody types embedding numbers by hand. They are learned by a neural network during training. The guiding idea, called the distributional hypothesis, is that words used in similar contexts have similar meanings: “coffee” and “tea” both appear near “cup”, “hot” and “drink”, so a model that predicts context words will push their vectors together.

How a word embedding is learned (word2vec-style)

  1. Start random: Give every word in the vocabulary a random vector, for example 300 small random numbers.
  2. Make a prediction task: From real text, take a word and its neighbours. Ask the model to tell true neighbour pairs (coffee, cup) apart from random pairs (coffee, tractor).
  3. Measure the error: Score a pair by the dot product of their vectors. True pairs should score high and random pairs low; the loss measures how far off we are.
  4. Nudge the vectors: Gradient descent moves vectors of true pairs slightly closer and random pairs slightly apart.
  5. Repeat billions of times: Over a large corpus, words sharing contexts end up near each other. The final vectors are the embeddings.

Modern systems go further. Inside every LLM, the first layer is an embedding table: a matrix with one row per token, learned along with the rest of the model. Transformers then produce contextual embeddings, where the vector for “bank” depends on the surrounding sentence. Dedicated embedding models (for example, Sentence-BERT-style models from 2019 onward) are trained with contrastive learning to place whole sentences or documents so that a question lands near its answer.

The famous word math example

A striking finding from the word2vec work (Mikolov and colleagues, 2013) is that some relationships become directions in the space. The vector from “man” to “woman” is roughly parallel to the vector from “king” to “queen”. So king − man + woman lands near queen. We can show this with hand-made 4-dimension vectors whose dimensions we label royalty, male, female and fruit.

embedding_math.py

import numpy as np
# Hand-made 4-D embeddings. Dimensions (for teaching only):
#            royalty  male  female  fruit
emb = {
"king":  np.array([0.9, 0.8, 0.1, 0.0]),
"queen": np.array([0.9, 0.1, 0.8, 0.0]),
"man":   np.array([0.1, 0.9, 0.1, 0.0]),
"woman": np.array([0.1, 0.1, 0.9, 0.0]),
"apple": np.array([0.0, 0.1, 0.1, 0.9]),
}
def cosine(a, b):
return a @ b / (np.linalg.norm(a) * np.linalg.norm(b))
# 1. Closeness: cosine similarity between pairs
for a, b in [("king", "queen"), ("man", "woman"), ("king", "apple")]:
print(f"cos({a}, {b}) = {cosine(emb[a], emb[b]):.2f}")
# 2. Euclidean distance tells a similar story here
print("dist(king, queen) =", round(np.linalg.norm(emb["king"] - emb["queen"]), 2))
print("dist(king, apple) =", round(np.linalg.norm(emb["king"] - emb["apple"]), 2))
# 3. Word math: king - man + woman = ?
target = emb["king"] - emb["man"] + emb["woman"]
print("king - man + woman =", np.round(target, 2))
ranked = sorted(
(w for w in emb if w not in ("king", "man", "woman")),
key=lambda w: -cosine(target, emb[w]),
)
for w in ranked:
print(f"  nearest: {w:6s} cos = {cosine(target, emb[w]):.2f}")

Output:

cos(king, queen) = 0.66
cos(man, woman) = 0.23
cos(king, apple) = 0.08
dist(king, queen) = 0.99
dist(king, apple) = 1.45
king - man + woman = [0.9 0.  0.9 0. ]
nearest: queen  cos = 0.99
nearest: apple  cos = 0.08

Do not over-trust the analogy trick In real models, analogies work only approximately and for some relations. The result is usually closest to one of the input words (often “king” itself), which is why tests exclude the inputs. It is a nice illustration that structure exists, not a reliable reasoning tool.

Embeddings are not only for words

Anything a neural network can read can be embedded. The same rule holds: similar items get nearby vectors.

  • Sentences and documents: our help articles and customer questions, embedded by the same model so they can be compared.
  • Images: a vision model maps photos to vectors; similar-looking or similar-content photos are close.
  • Text and images together: models like CLIP (OpenAI, 2021) put captions and pictures in one shared space, so the text “a red sneaker” lands near photos of red sneakers.
  • Audio: speech or music clips, for example to find similar songs.
  • Users and products: recommender systems learn a vector per user and per item; a high dot product means “likely to buy”.
  • Code, molecules, graphs: specialised models embed functions, chemical structures or nodes in a network.

What we use embeddings for

Common uses in AI systems
UseHow embeddings helpIn our help centre
Semantic searchEmbed the query and all documents; return the nearest documents“parcel never arrived” finds “missing delivery”
RAGRetrieve the nearest passages and give them to an LLM as contextThe bot answers using the right help article
RecommendationsSuggest items whose vectors are near what the user liked“Customers also read…” links
ClusteringGroup nearby vectors to discover topicsFind the top themes in this week's tickets
ClassificationTrain a small classifier on top of embeddingsRoute tickets to billing, shipping or returns
DeduplicationVery high similarity means near-duplicatesMerge duplicate help articles
Anomaly detectionItems far from every cluster are unusualSpot a new kind of complaint early

Our help-centre search, end to end Offline: embed every help article once and store the vectors in a vector database. Online: embed the customer's question with the same model, find the five articles with the highest cosine similarity, and show them (or pass them to an LLM to write an answer). The question “my parcel never arrived” now finds “What to do if your delivery is missing” without a single shared keyword.

Things we must be careful about

  • Never mix models: vectors from different embedding models (or different versions of one model) live in different spaces. Comparing them gives meaningless scores. Re-embed everything when we switch.
  • Bias: embeddings absorb stereotypes in their training text. A well-known 2016 study (Bolukbasi et al.) showed word2vec analogies linking “man” to “computer programmer” and “woman” to “homemaker”.
  • Domain gap: a general model may not know that two internal product codes are related. Test on our own data and consider domain-specific or fine-tuned models.
  • Similar is not the same as correct: the nearest document may be on the right topic but answer a different question; negations (“refund allowed” vs “refund not allowed”) can be close in embedding space.
  • Length limits: embedding models have a maximum input length; longer text may be cut off silently, so we split documents into chunks.
  • Exact terms: embeddings can miss exact identifiers like order numbers or error codes, where keyword search is stronger.

The most common mistake Embedding documents with one model and queries with another (or upgrading the model for new documents only). Results quietly get worse with no error message. Store the model name and version with every vector.

Pause and think: Our search returns the right topic but misses queries that contain an exact order number like “ORD-55812”. What should we add?

Keyword search (for example BM25) alongside the embedding search, a combination called hybrid search. Embeddings capture meaning but are weak at matching exact rare strings such as IDs and codes.

Worked example, step by step

A model produces one vector per token. So how does a whole sentence become a single vector? A common recipe is mean pooling: average the token vectors. Let us do it by hand for the question “parcel never arrived”, using made-up 2-D vectors whose dimensions we call delivery and payment (illustrative numbers).

The cosine with A in full: dot product = 0.7×0.8 + 0.2×0.2 = 0.60. Lengths: √(0.49 + 0.04) ≈ 0.728 and √(0.64 + 0.04) ≈ 0.825. So 0.60 / (0.728 × 0.825) ≈ 1.00. For B: 0.32 / (0.728 × 0.922) ≈ 0.48.

What plain averaging keeps and what it loses.
PropertyAfter mean pooling
Overall topicKept: the strong delivery signal survives the average
Sentence lengthRemoved: we divide by the token count
Word orderLost if token vectors are fixed: “dog bites man” and “man bites dog” average to the same vector
One rare but vital wordDiluted: in a long text, one token is a small share of the average

The word-order problem is why modern embedding models pool contextual token vectors. Each token's vector has already been shaped by its neighbours inside the Transformer, so the average of “dog bites man” differs from “man bites dog”. The dilution problem is one more reason to split long documents into chunks before embedding them.

Practice: try it yourself

We will build the core of our help-centre search: score four articles against one question, first with a raw dot product and then with cosine similarity, and return the top two.

practice_semantic_search.py

import numpy as np
# Hand-made 3-D embeddings (illustrative). Dimensions: delivery, payment, account
titles = [
"What to do if your delivery is missing",
"How to update your card details",
"Reset your password",
"Track a parcel",
]
docs = np.array([
[0.9, 0.1, 0.1],
[0.1, 0.9, 0.2],
[0.0, 0.1, 0.9],
[1.4, 0.9, 0.1],   # also about delivery, but a much longer vector
])
query = np.array([0.8, 0.2, 0.1])   # "my parcel never arrived"
# Raw dot product rewards long vectors
print("dot    :", np.round(docs @ query, 2))
# Normalise every vector to length 1; now dot product = cosine similarity
docs_n = docs / np.linalg.norm(docs, axis=1, keepdims=True)
query_n = query / np.linalg.norm(query)
scores = docs_n @ query_n            # one matrix-vector product scores all articles
print("cosine :", np.round(scores, 2))
# Top-2 search: sort scores from high to low and keep the first two
for rank, i in enumerate(np.argsort(-scores)[:2], start=1):
print(f"{rank}. {titles[i]}  (cos = {scores[i]:.2f})")

Output:

dot    : [0.75 0.28 0.11 1.31]
cosine : [0.99 0.36 0.15 0.95]
1. What to do if your delivery is missing  (cos = 0.99)
2. Track a parcel  (cos = 0.95)

Now change it:

  • Multiply the query by 10: query = 10 * np.array([0.8, 0.2, 0.1]). Predict first: which printed line changes and which stays the same?
  • Change the query to [0.1, 0.9, 0.1] (“my card was declined”). Predict the new top-2 before running.
  • Add a fifth article “Delivery and payment FAQ” with vector [0.6, 0.6, 0.1]. Predict where it ranks for the original query.

Pause and think: The dot product ranks “Track a parcel” first (1.31 against 0.75), but cosine ranks it second (0.95 against 0.99). Which ranking should our search trust, and why?

Cosine. The dot product grew because the vector is long, not because the meaning is closer. Its direction also leans a little towards payment, which the question barely mentions. Cosine removes length and judges direction only, which is what “similar meaning” should mean here.

Pause and think: “Reset your password” still gets a score of 0.15, not 0. If we asked for the top 4, it would be returned. What does this tell us about using search results?

A similarity score is always some number, so “it was returned” does not mean “it is relevant”. Top-k always fills k slots, even with poor matches. Real systems keep k small or add a minimum score, and the right cut-off must be tuned on real questions because score ranges differ between models.

Summary

Embeddings turn items into vectors so that closeness in space reflects similarity in meaning. We compare them with cosine similarity, dot product or distance. They are learned from data, from word2vec-style context prediction to contrastively trained Transformer models, and they work for words, sentences, images, audio, users and products. They power semantic search, RAG, recommendations and clustering, as long as we keep one model per index, watch for bias and domain gaps, and pair them with keyword search for exact terms.

Key takeaways

  • An embedding is a vector where closeness means similarity in meaning.
  • Cosine similarity compares direction; dot product equals cosine on normalised vectors; Euclidean measures distance.
  • Embeddings are learned from data, from context prediction (word2vec) to contrastively trained Transformers.
  • They work for words, sentences, images, audio, users and products, and power search, RAG and recommendations.
  • Use one model per index, watch bias and domain gaps, and add keyword search for exact terms.

Key terms

  • Embedding: A dense vector of numbers representing an item so that similar items are close together.
  • Dimension: One number in a vector; the dimensionality is how many numbers the vector has.
  • One-hot vector: A vector with a single 1 marking a category and 0s elsewhere; it has no notion of similarity.
  • Cosine similarity: The cosine of the angle between two vectors, from −1 to 1, ignoring their lengths.
  • Contextual embedding: A vector for a token or text computed from its surrounding context, typically by a Transformer.
  • Distributional hypothesis: The idea that words appearing in similar contexts have similar meanings.

← 4.3 BPE Tokenization: How LLMs Split Text into Tokens · 4.5 RNNs vs Transformers: A Fundamental Architecture Shift →