Lesson 16.3 · 23 min
Image Embeddings: Encoding Visual Content as Vectors
Two photos of the same red mug can differ in almost every pixel — so how does a computer still know they show the same thing?
In short: An image embedding is a short list of numbers (a vector) that captures what an image shows, produced by a trained neural network called an encoder. Similar-looking or similar-meaning images get vectors that point in similar directions, so we can compare images with simple math such as cosine similarity. Embeddings power image search, duplicate detection, recommendations, clustering and text-to-image search.
What is an embedding?
An embedding is a way to represent a thing — a word, a sentence, a product, an image — as a vector: an ordered list of numbers such as [0.12, -0.80, 0.33, ...]. The list has a fixed length called the dimension (often 256 to 1,024 numbers for images in practice). The key property is that things with similar meaning get vectors that are close together, and unrelated things get vectors that are far apart.
We met text embeddings earlier in the course: 'refund' and 'money back' land near each other. The idea is the same for images. Each image becomes a point in a high-dimensional space, and 'how similar are these two images?' becomes 'how close are these two points?'
Think of it like a library's map Imagine a giant library where a clever librarian places every photo on a map so that similar photos sit near each other: all beaches in one corner, all mugs in another, red mugs slightly apart from blue mugs. An embedding is a photo's coordinates on that map. The map just happens to have hundreds of directions instead of two.
What is an image embedding?
An image embedding is a fixed-length vector produced by feeding an image through a trained neural network (an image encoder) and reading out one of its internal representations. Our running example is the photo archive of an appliance shop's support team: 80,000 photos of products customers have sent in. Each photo, whatever its size, becomes a vector of, say, 512 numbers.
Individual numbers in an embedding usually have no human-readable meaning — there is no 'redness' slot. Meaning lives in the pattern across all the numbers and in the relations between vectors. What matters is that two photos of cracked blender jars produce vectors that are close, and a photo of a toaster produces a vector further away.
Pause and think: One customer photo is 4000×3000 pixels and another is 640×480. Do their embeddings have different lengths?
No. The encoder resizes or crops each image to its expected input size, and the output embedding always has the same dimension (for example 512). That fixed length is what makes comparing any two images easy.
Why do we need image embeddings?
Our support team wants to ask questions like 'have we seen this kind of damage before?' or 'find all tickets showing this model of kettle'. To answer, the computer must compare images. The naive way is to compare raw pixels, and that fails badly:
- Pixels are fragile. Move the camera one centimetre, change the lighting, or zoom slightly, and almost every pixel value changes — even though the photo shows the same mug.
- Pixels are huge. A 1024×1024 colour photo is over 3 million numbers. Comparing millions of such photos number by number is slow and storage-hungry.
- Pixels carry no meaning. Two very different objects in the same colours can be closer, pixel-wise, than two photos of the same object on different backgrounds.
Embeddings fix all three: they are compact (hundreds of numbers instead of millions), and a good encoder is trained so that the vector stays nearly the same when irrelevant details change (position, lighting, background) but changes when the content changes. Once images are vectors, we can reuse all the vector tools from the RAG lessons: nearest-neighbour search, vector databases, clustering.
How does a computer see an image?
To a computer, a colour image is a 3-D array of numbers: height × width × 3 channels (red, green, blue). Each value is a brightness, typically an integer from 0 to 255, usually rescaled to 0–1 before entering a network. A tiny 2×2 image might look like this:
| Column 1 | Column 2 | |
|---|---|---|
| Row 1 | (255, 0, 0) — pure red | (250, 10, 5) — red |
| Row 2 | (0, 0, 255) — pure blue | (255, 255, 255) — white |
That is all the encoder gets: a grid of brightness numbers. Concepts like 'mug', 'crack' or 'kitchen' do not exist in the input. A neural network has to build them up layer by layer: early layers respond to edges and colour changes, middle layers to textures and simple shapes (a curve, a handle), and deep layers to whole objects and scenes. The embedding is read from the deep end, where the representation is about what is in the picture, not which pixels are bright.
How are image embeddings created?
From photo to vector
- Pre-process: Resize (and usually centre-crop) the image to the encoder's input size, for example 224×224, and normalise the pixel values the same way as during training.
- Run the encoder: Pass the image through a trained network. Common choices: a CNN such as ResNet, or a Vision Transformer (ViT) that splits the image into patches and processes them with self-attention.
- Take a summary vector: Read out one vector for the whole image: the ViT's [CLS] output, or the average of the last feature map (global average pooling) for a CNN. We drop any final classification layer — we want the features, not class labels.
- Project (optional): Some models add a learned linear layer that maps the features into a specific space, such as the joint image-text space of CLIP.
- Normalise: Divide the vector by its length so it has length 1 (L2 normalisation). Then cosine similarity is just a dot product, which vector databases compute very fast.
The quality of the embedding depends entirely on how the encoder was trained. There are three main recipes:
A simple numeric walkthrough
Real embeddings have hundreds of dimensions, but the math is identical with 4. Suppose our encoder produced these vectors for three support photos:
| Photo | Embedding |
|---|---|
| A: red mug, kitchen counter | [0.90, 0.10, 0.30, 0.20] |
| B: red mug, held in hand | [0.80, 0.20, 0.35, 0.10] |
| C: blue kettle | [0.10, 0.90, 0.20, 0.70] |
Cosine similarity of A and B by hand
- Dot product: Multiply matching positions and add: 0.90·0.80 + 0.10·0.20 + 0.30·0.35 + 0.20·0.10 = 0.72 + 0.02 + 0.105 + 0.02 = 0.865.
- Length of A: ‖A‖ = √(0.81 + 0.01 + 0.09 + 0.04) = √0.95 ≈ 0.975.
- Length of B: ‖B‖ = √(0.64 + 0.04 + 0.1225 + 0.01) = √0.8125 ≈ 0.901.
- Divide: cos(A, B) = 0.865 / (0.975 × 0.901) ≈ 0.865 / 0.879 ≈ 0.985. Almost 1: very similar.
- Compare with C: A·C = 0.09 + 0.09 + 0.06 + 0.14 = 0.38, ‖C‖ ≈ 1.162, so cos(A, C) ≈ 0.38 / (0.975 × 1.162) ≈ 0.336. Much less similar.
How do we measure similarity between two embeddings?
- Cosine similarity measures the angle between vectors, ignoring length. It ranges from −1 (opposite) through 0 (unrelated) to 1 (same direction). It is the most common choice for embeddings.
- Dot product is cosine times both lengths. If all vectors are normalised to length 1, dot product equals cosine, which is why systems normalise first.
- Euclidean (L2) distance is the straight-line distance; smaller means more similar. For unit-length vectors it ranks results in exactly the same order as cosine, because ‖a − b‖² = 2 − 2·cos(a, b).
Similarity scores are not probabilities A cosine of 0.8 does not mean '80% the same'. Typical score ranges differ between models: one model's 'very similar' may be 0.9, another's 0.3. Always pick thresholds by looking at real examples from our model and data.
A code example
A real encoder needs a deep-learning library and downloaded weights. To see the principle with only numpy, we write a tiny hand-made 'encoder' that summarises an image by its average colour, colour spread and edge strength. It is far weaker than a neural network, but it shows the two key ideas: an embedding is small and fixed-size, and it can ignore changes that should not matter (here, shifting the image).
tiny_image_embeddings.py
import numpy as np
rng = np.random.default_rng(0)
def make(base_rgb, stripes=False):
"""A 6x6 RGB image: a base colour, small noise, optional dark stripes."""
im = np.clip(np.array(base_rgb) + rng.normal(0, 0.05, (6, 6, 3)), 0, 1)
if stripes:
im[:, ::2] *= 0.3 # darken every other column
return im
imgs = {
"beach": make([0.9, 0.8, 0.5]),
"ocean": make([0.1, 0.3, 0.9]),
"ocean_2": make([0.15, 0.35, 0.85]),
"zebra": make([0.9, 0.9, 0.9], stripes=True),
}
# Same zebra, shifted one pixel to the right (np.roll wraps the edge)
imgs["zebra_moved"] = np.roll(imgs["zebra"], 1, axis=1)
def embed(im):
"""A hand-made 'encoder': mean colour + colour spread + edge strength."""
mean = im.mean(axis=(0, 1)) # 3 numbers
spread = im.std(axis=(0, 1)) # 3 numbers
edges = np.abs(np.diff(im, axis=1)).mean(keepdims=True) # 1 number
return np.concatenate([mean, spread, edges[0, 0]])
def cos(a, b):
return a @ b / (np.linalg.norm(a) * np.linalg.norm(b))
names = list(imgs)
E = {n: embed(imgs[n]) for n in names}
print("raw pixels per image:", imgs["beach"].size, "| embedding size:", E["beach"].size)
print("ocean embedding:", np.round(E["ocean"], 2))
q = "zebra"
print(f"\nquery = {q}")
print(" pixel cosine zebra vs zebra_moved:", round(cos(imgs[q].ravel(), imgs['zebra_moved'].ravel()), 3))
print(" embed cosine zebra vs zebra_moved:", round(cos(E[q], E['zebra_moved']), 3))
q = "ocean"
ranked = sorted((n for n in names if n != q), key=lambda n: -cos(E[q], E[n]))
print(f"\nnearest neighbours of {q}:")
for n in ranked:
print(f" {n:12s} cosine = {cos(E[q], E[n]):.3f}")Output:
raw pixels per image: 108 | embedding size: 7 ocean embedding: [0.1 0.3 0.9 0.05 0.05 0.05 0.05] query = zebra pixel cosine zebra vs zebra_moved: 0.552 embed cosine zebra vs zebra_moved: 1.0 nearest neighbours of ocean: ocean_2 cosine = 0.996 zebra_moved cosine = 0.664 zebra cosine = 0.664 beach cosine = 0.639
With a real model the code shape is the same: load a pre-trained encoder, run each image through it, L2-normalise the outputs, store them in a vector index, and search with cosine similarity. Only the embed function changes.
Pause and think: Why did the zebra and the beach both score only about 0.64–0.66 against the ocean, even though they look nothing alike?
Our 7 hand-made features are crude: all values are positive, so every vector points into the same 'positive' region and cosines rarely drop near 0. A trained encoder spreads images across many more dimensions, giving clearer separation. It is a reminder that scores are only meaningful relative to other scores from the same model.
Where are image embeddings used?
| Application | How embeddings help | Support-team example |
|---|---|---|
| Reverse image search | Embed the query image, return nearest neighbours | Find past tickets with the same damage |
| Text-to-image search | Embed a sentence with a CLIP-style text encoder, search image vectors | 'cracked jar base' → matching photos |
| Duplicate / near-duplicate detection | Very high similarity flags copies | Spot the same photo submitted for two claims |
| Recommendations | Suggest items whose images are close | 'Customers also viewed' similar kettles |
| Clustering and labelling | Group vectors, label each cluster once | Discover the most common failure types |
| Classification with few labels | Train a small classifier on top of frozen embeddings | Damage vs no damage with 500 labelled photos |
| Multimodal RAG | Retrieve relevant images (or pages as images) as context for an LLM | Pull the right manual diagram into the answer |
Common pitfalls Mixing embeddings from two different models (their spaces are unrelated, so comparisons are meaningless); forgetting to apply the same pre-processing at query time as at indexing time; trusting a generic encoder on a very specialised domain (X-rays, circuit boards) without evaluation; and re-embedding only part of the archive after switching models.
When not to use them: if we need exact matches (same file), a cryptographic hash is cheaper and certain; if we need to read text or numbers from an image, use OCR; and if decisions are high-stakes, embedding similarity should only shortlist candidates for a human or a stronger model to check.
Common mistakes and how to spot them
Embedding search fails quietly. There is no error message; the results are just a little worse, or strangely repetitive. So it pays to know the usual faults and the quick test for each. Start with the most common one: forgetting to normalise.
Take a query q = [1, 0] and two stored photos, A = [0.9, 0.1] and B = [3, 4]. By cosine, A wins easily: cos(q, A) = 0.9 / 0.906 ≈ 0.99, while cos(q, B) = 3 / 5 = 0.6. But the raw dot products are q·A = 0.9 and q·B = 3. Ranked by raw dot product, B comes first only because it is long. One long vector like this can show up near the top of every search.
| What we see | Likely cause | Quick test |
|---|---|---|
| The same few photos appear in almost every result list | Vectors not normalised; long ones win on dot product | Print the lengths of 100 stored vectors. They should all be ≈ 1.0 |
| Results look random after a model upgrade | Old and new model vectors mixed in one index | Embed one stored photo again and compare with its stored vector. Cosine should be ≈ 1.0 |
| Good results in tests, poor on live queries | Different resize, crop or colour scaling at query time | Run one image through both code paths and compare the two vectors |
| Everything scores between 0.6 and 0.9 | Normal for many encoders; scores are bunched | Judge by rank and by gaps, not by the absolute number |
| Duplicates found, but so are unrelated photos | Threshold copied from another model or guessed | Score 50 known duplicate pairs and 50 known different pairs, then pick a value between the two groups |
A four-step health check for a new index
- Self-match: Search with a photo that is already in the index. It must come back as result number one with a cosine very close to 1.0. If not, indexing and querying do not use the same pipeline.
- Length check: Compute the L2 norm of a sample of stored vectors. Any value far from 1.0 means normalisation was skipped somewhere.
- Eyeball test: Pick 20 real queries and look at the top 5 results for each. Wrong results that share a background or a lighting style tell us what the encoder is paying attention to.
- Labelled pairs: Collect a small set of pairs we know match and pairs we know do not. Their two score ranges show where a safe threshold sits, or that no clean threshold exists.
Practice: try it yourself
We will build a miniature search index: 12 made-up photo embeddings for mugs, kettles and toasters. One stored vector is 'broken' (far too long). We search with a new mug photo, first carelessly and then properly, and we check the link between cosine and L2 distance with real numbers.
practice_mini_index.py
import numpy as np
rng = np.random.default_rng(11)
# A tiny photo archive: 3 kinds of product, 4 photos each, 6-d embeddings.
kinds = ["mug", "kettle", "toaster"]
centres = rng.normal(size=(3, 6)) # one direction per kind
labels = [k for k in kinds for _ in range(4)]
vecs = np.repeat(centres, 4, axis=0) + rng.normal(0, 0.25, size=(12, 6))
vecs[5] *= 6.0 # one kettle vector is very long
def unit(v):
"""L2-normalise: divide each row by its length."""
return v / np.linalg.norm(v, axis=-1, keepdims=True)
query = centres[0] + rng.normal(0, 0.25, size=6) # a new mug photo
def top3(scores):
best = np.argsort(-scores)[:3]
return [f"{labels[i]}#{i}" for i in best]
raw_dot = vecs @ query # no normalisation
cosine = unit(vecs) @ unit(query) # normalise both sides first
print("top 3 by raw dot product:", top3(raw_dot))
print("top 3 by cosine :", top3(cosine))
# On unit vectors, L2 distance gives the same ranking as cosine.
dist = np.linalg.norm(unit(vecs) - unit(query), axis=1)
same = (np.argsort(dist) == np.argsort(-cosine)).all()
print("L2 ranking equals cosine ranking:", same)
print("check d^2 = 2 - 2cos on photo 0:",
round(dist[0] ** 2, 4), "vs", round(2 - 2 * cosine[0], 4))
# A threshold for "is this a mug?" must come from our own scores.
mug = cosine[:4]; other = cosine[4:]
print(f"lowest mug score {mug.min():.2f} | highest non-mug score {other.max():.2f}")Output:
top 3 by raw dot product: ['kettle#5', 'mug#3', 'mug#0'] top 3 by cosine : ['mug#2', 'mug#3', 'mug#0'] L2 ranking equals cosine ranking: True check d^2 = 2 - 2cos on photo 0: 0.2435 vs 0.2435 lowest mug score 0.74 | highest non-mug score 0.47
Now change it:
- Change the noise in the
vecsline from0.25to1.0. Predict what happens to the gap between the lowest mug score and the highest non-mug score. - Change
vecs[5] = 6.0tovecs[5] = 0.01(a very short vector). Predict whether the raw dot product ranking is now correct, and whether the cosine ranking changes at all. - Replace
querywith(centres[0] + centres[1]) / 2, a photo showing a mug next to a kettle. Predict which kinds appear in the cosine top 3.
Pause and think: All stored vectors are unit length, but we forget to normalise the query. We rank by dot product. Is the ranking wrong?
No, the ranking is still right. A longer query multiplies every score by the same number, so the order does not change. What breaks is the meaning of the scores: they are no longer cosines, so any threshold such as 'above 0.8 is a duplicate' stops working. The dangerous case is unnormalised stored vectors, because each one is scaled differently.
Pause and think: In the run above the lowest mug score was 0.74 and the highest non-mug score was 0.47. We set the threshold at 0.6 and then swap in a different encoder. Can we keep 0.6?
Not safely. The gap between 0.47 and 0.74 belongs to this encoder and this data. Another model may bunch all its scores between 0.2 and 0.4, or between 0.8 and 0.95. We must score known matching and non-matching pairs again and pick a new threshold.
Summary
- An embedding is a fixed-length vector where similar things are close together.
- An image embedding comes from running a picture through a trained encoder (CNN or ViT) and taking a summary vector.
- Raw pixels are huge, fragile and meaningless; embeddings are compact and robust to irrelevant changes.
- Encoders are trained with labels, with image-caption pairs (CLIP-style) or self-supervised; the training decides what 'similar' means.
- Compare embeddings with cosine similarity (or dot product on normalised vectors).
- Uses: image search, text-to-image search, deduplication, recommendations, clustering and multimodal RAG.
Key takeaways
- An image embedding is a compact, fixed-length vector that captures what an image shows.
- It comes from a trained encoder; how that encoder was trained defines what 'similar' means.
- Embeddings are robust to irrelevant changes (shift, lighting) where raw pixels are not.
- Cosine similarity (a dot product on normalised vectors) is the standard comparison.
- Never compare embeddings from different models, and tune thresholds on real examples.
Key terms
- Embedding: A fixed-length vector representing an item so that similar items are close together.
- Image encoder: A trained network (CNN or ViT) that turns an image into features or an embedding.
- Cosine similarity: The cosine of the angle between two vectors: dot product divided by both lengths.
- L2 normalisation: Dividing a vector by its length so it has length 1.
- Contrastive learning: Training that pulls matching pairs together and pushes non-matching pairs apart.
- Nearest-neighbour search: Finding the stored vectors most similar to a query vector.
← 16.2 Vision Transformers: Applying Self-Attention to Image Patches · 16.4 Diffusion Models: Iterative Denoising to Generate Images →