Modern AI Engineering

Lesson 16.3 · 23 min

Image Embeddings: Encoding Visual Content as Vectors

Two photos of the same red mug can differ in almost every pixel — so how does a computer still know they show the same thing?

In short: An image embedding is a short list of numbers (a vector) that captures what an image shows, produced by a trained neural network called an encoder. Similar-looking or similar-meaning images get vectors that point in similar directions, so we can compare images with simple math such as cosine similarity. Embeddings power image search, duplicate detection, recommendations, clustering and text-to-image search.

What is an embedding?

An embedding is a way to represent a thing — a word, a sentence, a product, an image — as a vector: an ordered list of numbers such as [0.12, -0.80, 0.33, ...]. The list has a fixed length called the dimension (often 256 to 1,024 numbers for images in practice). The key property is that things with similar meaning get vectors that are close together, and unrelated things get vectors that are far apart.

We met text embeddings earlier in the course: 'refund' and 'money back' land near each other. The idea is the same for images. Each image becomes a point in a high-dimensional space, and 'how similar are these two images?' becomes 'how close are these two points?'

Think of it like a library's map Imagine a giant library where a clever librarian places every photo on a map so that similar photos sit near each other: all beaches in one corner, all mugs in another, red mugs slightly apart from blue mugs. An embedding is a photo's coordinates on that map. The map just happens to have hundreds of directions instead of two.

What is an image embedding?

An image embedding is a fixed-length vector produced by feeding an image through a trained neural network (an image encoder) and reading out one of its internal representations. Our running example is the photo archive of an appliance shop's support team: 80,000 photos of products customers have sent in. Each photo, whatever its size, becomes a vector of, say, 512 numbers.

Individual numbers in an embedding usually have no human-readable meaning — there is no 'redness' slot. Meaning lives in the pattern across all the numbers and in the relations between vectors. What matters is that two photos of cracked blender jars produce vectors that are close, and a photo of a toaster produces a vector further away.

Pause and think: One customer photo is 4000×3000 pixels and another is 640×480. Do their embeddings have different lengths?

No. The encoder resizes or crops each image to its expected input size, and the output embedding always has the same dimension (for example 512). That fixed length is what makes comparing any two images easy.

Why do we need image embeddings?

Our support team wants to ask questions like 'have we seen this kind of damage before?' or 'find all tickets showing this model of kettle'. To answer, the computer must compare images. The naive way is to compare raw pixels, and that fails badly:

  • Pixels are fragile. Move the camera one centimetre, change the lighting, or zoom slightly, and almost every pixel value changes — even though the photo shows the same mug.
  • Pixels are huge. A 1024×1024 colour photo is over 3 million numbers. Comparing millions of such photos number by number is slow and storage-hungry.
  • Pixels carry no meaning. Two very different objects in the same colours can be closer, pixel-wise, than two photos of the same object on different backgrounds.

Embeddings fix all three: they are compact (hundreds of numbers instead of millions), and a good encoder is trained so that the vector stays nearly the same when irrelevant details change (position, lighting, background) but changes when the content changes. Once images are vectors, we can reuse all the vector tools from the RAG lessons: nearest-neighbour search, vector databases, clustering.

How does a computer see an image?

To a computer, a colour image is a 3-D array of numbers: height × width × 3 channels (red, green, blue). Each value is a brightness, typically an integer from 0 to 255, usually rescaled to 0–1 before entering a network. A tiny 2×2 image might look like this:

A 2×2 colour image as numbers (R, G, B per pixel, 0–255)
Column 1Column 2
Row 1(255, 0, 0) — pure red(250, 10, 5) — red
Row 2(0, 0, 255) — pure blue(255, 255, 255) — white

That is all the encoder gets: a grid of brightness numbers. Concepts like 'mug', 'crack' or 'kitchen' do not exist in the input. A neural network has to build them up layer by layer: early layers respond to edges and colour changes, middle layers to textures and simple shapes (a curve, a handle), and deep layers to whole objects and scenes. The embedding is read from the deep end, where the representation is about what is in the picture, not which pixels are bright.

How are image embeddings created?

From photo to vector

  1. Pre-process: Resize (and usually centre-crop) the image to the encoder's input size, for example 224×224, and normalise the pixel values the same way as during training.
  2. Run the encoder: Pass the image through a trained network. Common choices: a CNN such as ResNet, or a Vision Transformer (ViT) that splits the image into patches and processes them with self-attention.
  3. Take a summary vector: Read out one vector for the whole image: the ViT's [CLS] output, or the average of the last feature map (global average pooling) for a CNN. We drop any final classification layer — we want the features, not class labels.
  4. Project (optional): Some models add a learned linear layer that maps the features into a specific space, such as the joint image-text space of CLIP.
  5. Normalise: Divide the vector by its length so it has length 1 (L2 normalisation). Then cosine similarity is just a dot product, which vector databases compute very fast.

The quality of the embedding depends entirely on how the encoder was trained. There are three main recipes:

A simple numeric walkthrough

Real embeddings have hundreds of dimensions, but the math is identical with 4. Suppose our encoder produced these vectors for three support photos:

Illustrative 4-d embeddings
PhotoEmbedding
A: red mug, kitchen counter[0.90, 0.10, 0.30, 0.20]
B: red mug, held in hand[0.80, 0.20, 0.35, 0.10]
C: blue kettle[0.10, 0.90, 0.20, 0.70]

Cosine similarity of A and B by hand

  1. Dot product: Multiply matching positions and add: 0.90·0.80 + 0.10·0.20 + 0.30·0.35 + 0.20·0.10 = 0.72 + 0.02 + 0.105 + 0.02 = 0.865.
  2. Length of A: ‖A‖ = √(0.81 + 0.01 + 0.09 + 0.04) = √0.95 ≈ 0.975.
  3. Length of B: ‖B‖ = √(0.64 + 0.04 + 0.1225 + 0.01) = √0.8125 ≈ 0.901.
  4. Divide: cos(A, B) = 0.865 / (0.975 × 0.901) ≈ 0.865 / 0.879 ≈ 0.985. Almost 1: very similar.
  5. Compare with C: A·C = 0.09 + 0.09 + 0.06 + 0.14 = 0.38, ‖C‖ ≈ 1.162, so cos(A, C) ≈ 0.38 / (0.975 × 1.162) ≈ 0.336. Much less similar.

How do we measure similarity between two embeddings?

  • Cosine similarity measures the angle between vectors, ignoring length. It ranges from −1 (opposite) through 0 (unrelated) to 1 (same direction). It is the most common choice for embeddings.
  • Dot product is cosine times both lengths. If all vectors are normalised to length 1, dot product equals cosine, which is why systems normalise first.
  • Euclidean (L2) distance is the straight-line distance; smaller means more similar. For unit-length vectors it ranks results in exactly the same order as cosine, because ‖a − b‖² = 2 − 2·cos(a, b).

Similarity scores are not probabilities A cosine of 0.8 does not mean '80% the same'. Typical score ranges differ between models: one model's 'very similar' may be 0.9, another's 0.3. Always pick thresholds by looking at real examples from our model and data.

A code example

A real encoder needs a deep-learning library and downloaded weights. To see the principle with only numpy, we write a tiny hand-made 'encoder' that summarises an image by its average colour, colour spread and edge strength. It is far weaker than a neural network, but it shows the two key ideas: an embedding is small and fixed-size, and it can ignore changes that should not matter (here, shifting the image).

tiny_image_embeddings.py

import numpy as np
rng = np.random.default_rng(0)
def make(base_rgb, stripes=False):
"""A 6x6 RGB image: a base colour, small noise, optional dark stripes."""
im = np.clip(np.array(base_rgb) + rng.normal(0, 0.05, (6, 6, 3)), 0, 1)
if stripes:
im[:, ::2] *= 0.3                     # darken every other column
return im
imgs = {
"beach":        make([0.9, 0.8, 0.5]),
"ocean":        make([0.1, 0.3, 0.9]),
"ocean_2":      make([0.15, 0.35, 0.85]),
"zebra":        make([0.9, 0.9, 0.9], stripes=True),
}
# Same zebra, shifted one pixel to the right (np.roll wraps the edge)
imgs["zebra_moved"] = np.roll(imgs["zebra"], 1, axis=1)
def embed(im):
"""A hand-made 'encoder': mean colour + colour spread + edge strength."""
mean = im.mean(axis=(0, 1))                       # 3 numbers
spread = im.std(axis=(0, 1))                      # 3 numbers
edges = np.abs(np.diff(im, axis=1)).mean(keepdims=True)  # 1 number
return np.concatenate([mean, spread, edges[0, 0]])
def cos(a, b):
return a @ b / (np.linalg.norm(a) * np.linalg.norm(b))
names = list(imgs)
E = {n: embed(imgs[n]) for n in names}
print("raw pixels per image:", imgs["beach"].size, "| embedding size:", E["beach"].size)
print("ocean embedding:", np.round(E["ocean"], 2))
q = "zebra"
print(f"\nquery = {q}")
print("  pixel cosine  zebra vs zebra_moved:", round(cos(imgs[q].ravel(), imgs['zebra_moved'].ravel()), 3))
print("  embed cosine  zebra vs zebra_moved:", round(cos(E[q], E['zebra_moved']), 3))
q = "ocean"
ranked = sorted((n for n in names if n != q), key=lambda n: -cos(E[q], E[n]))
print(f"\nnearest neighbours of {q}:")
for n in ranked:
print(f"  {n:12s} cosine = {cos(E[q], E[n]):.3f}")

Output:

raw pixels per image: 108 | embedding size: 7
ocean embedding: [0.1  0.3  0.9  0.05 0.05 0.05 0.05]
query = zebra
pixel cosine  zebra vs zebra_moved: 0.552
embed cosine  zebra vs zebra_moved: 1.0
nearest neighbours of ocean:
ocean_2      cosine = 0.996
zebra_moved  cosine = 0.664
zebra        cosine = 0.664
beach        cosine = 0.639

With a real model the code shape is the same: load a pre-trained encoder, run each image through it, L2-normalise the outputs, store them in a vector index, and search with cosine similarity. Only the embed function changes.

Pause and think: Why did the zebra and the beach both score only about 0.64–0.66 against the ocean, even though they look nothing alike?

Our 7 hand-made features are crude: all values are positive, so every vector points into the same 'positive' region and cosines rarely drop near 0. A trained encoder spreads images across many more dimensions, giving clearer separation. It is a reminder that scores are only meaningful relative to other scores from the same model.

Where are image embeddings used?

Common applications
ApplicationHow embeddings helpSupport-team example
Reverse image searchEmbed the query image, return nearest neighboursFind past tickets with the same damage
Text-to-image searchEmbed a sentence with a CLIP-style text encoder, search image vectors'cracked jar base' → matching photos
Duplicate / near-duplicate detectionVery high similarity flags copiesSpot the same photo submitted for two claims
RecommendationsSuggest items whose images are close'Customers also viewed' similar kettles
Clustering and labellingGroup vectors, label each cluster onceDiscover the most common failure types
Classification with few labelsTrain a small classifier on top of frozen embeddingsDamage vs no damage with 500 labelled photos
Multimodal RAGRetrieve relevant images (or pages as images) as context for an LLMPull the right manual diagram into the answer

Common pitfalls Mixing embeddings from two different models (their spaces are unrelated, so comparisons are meaningless); forgetting to apply the same pre-processing at query time as at indexing time; trusting a generic encoder on a very specialised domain (X-rays, circuit boards) without evaluation; and re-embedding only part of the archive after switching models.

When not to use them: if we need exact matches (same file), a cryptographic hash is cheaper and certain; if we need to read text or numbers from an image, use OCR; and if decisions are high-stakes, embedding similarity should only shortlist candidates for a human or a stronger model to check.

Common mistakes and how to spot them

Embedding search fails quietly. There is no error message; the results are just a little worse, or strangely repetitive. So it pays to know the usual faults and the quick test for each. Start with the most common one: forgetting to normalise.

Take a query q = [1, 0] and two stored photos, A = [0.9, 0.1] and B = [3, 4]. By cosine, A wins easily: cos(q, A) = 0.9 / 0.906 ≈ 0.99, while cos(q, B) = 3 / 5 = 0.6. But the raw dot products are q·A = 0.9 and q·B = 3. Ranked by raw dot product, B comes first only because it is long. One long vector like this can show up near the top of every search.

Symptoms, likely causes and quick tests
What we seeLikely causeQuick test
The same few photos appear in almost every result listVectors not normalised; long ones win on dot productPrint the lengths of 100 stored vectors. They should all be ≈ 1.0
Results look random after a model upgradeOld and new model vectors mixed in one indexEmbed one stored photo again and compare with its stored vector. Cosine should be ≈ 1.0
Good results in tests, poor on live queriesDifferent resize, crop or colour scaling at query timeRun one image through both code paths and compare the two vectors
Everything scores between 0.6 and 0.9Normal for many encoders; scores are bunchedJudge by rank and by gaps, not by the absolute number
Duplicates found, but so are unrelated photosThreshold copied from another model or guessedScore 50 known duplicate pairs and 50 known different pairs, then pick a value between the two groups

A four-step health check for a new index

  1. Self-match: Search with a photo that is already in the index. It must come back as result number one with a cosine very close to 1.0. If not, indexing and querying do not use the same pipeline.
  2. Length check: Compute the L2 norm of a sample of stored vectors. Any value far from 1.0 means normalisation was skipped somewhere.
  3. Eyeball test: Pick 20 real queries and look at the top 5 results for each. Wrong results that share a background or a lighting style tell us what the encoder is paying attention to.
  4. Labelled pairs: Collect a small set of pairs we know match and pairs we know do not. Their two score ranges show where a safe threshold sits, or that no clean threshold exists.

Practice: try it yourself

We will build a miniature search index: 12 made-up photo embeddings for mugs, kettles and toasters. One stored vector is 'broken' (far too long). We search with a new mug photo, first carelessly and then properly, and we check the link between cosine and L2 distance with real numbers.

practice_mini_index.py

import numpy as np
rng = np.random.default_rng(11)
# A tiny photo archive: 3 kinds of product, 4 photos each, 6-d embeddings.
kinds = ["mug", "kettle", "toaster"]
centres = rng.normal(size=(3, 6))                # one direction per kind
labels = [k for k in kinds for _ in range(4)]
vecs = np.repeat(centres, 4, axis=0) + rng.normal(0, 0.25, size=(12, 6))
vecs[5] *= 6.0                                   # one kettle vector is very long
def unit(v):
"""L2-normalise: divide each row by its length."""
return v / np.linalg.norm(v, axis=-1, keepdims=True)
query = centres[0] + rng.normal(0, 0.25, size=6) # a new mug photo
def top3(scores):
best = np.argsort(-scores)[:3]
return [f"{labels[i]}#{i}" for i in best]
raw_dot = vecs @ query                           # no normalisation
cosine = unit(vecs) @ unit(query)                # normalise both sides first
print("top 3 by raw dot product:", top3(raw_dot))
print("top 3 by cosine         :", top3(cosine))
# On unit vectors, L2 distance gives the same ranking as cosine.
dist = np.linalg.norm(unit(vecs) - unit(query), axis=1)
same = (np.argsort(dist) == np.argsort(-cosine)).all()
print("L2 ranking equals cosine ranking:", same)
print("check  d^2 = 2 - 2cos  on photo 0:",
round(dist[0] ** 2, 4), "vs", round(2 - 2 * cosine[0], 4))
# A threshold for "is this a mug?" must come from our own scores.
mug = cosine[:4]; other = cosine[4:]
print(f"lowest mug score {mug.min():.2f} | highest non-mug score {other.max():.2f}")

Output:

top 3 by raw dot product: ['kettle#5', 'mug#3', 'mug#0']
top 3 by cosine         : ['mug#2', 'mug#3', 'mug#0']
L2 ranking equals cosine ranking: True
check  d^2 = 2 - 2cos  on photo 0: 0.2435 vs 0.2435
lowest mug score 0.74 | highest non-mug score 0.47

Now change it:

  • Change the noise in the vecs line from 0.25 to 1.0. Predict what happens to the gap between the lowest mug score and the highest non-mug score.
  • Change vecs[5] = 6.0 to vecs[5] = 0.01 (a very short vector). Predict whether the raw dot product ranking is now correct, and whether the cosine ranking changes at all.
  • Replace query with (centres[0] + centres[1]) / 2, a photo showing a mug next to a kettle. Predict which kinds appear in the cosine top 3.

Pause and think: All stored vectors are unit length, but we forget to normalise the query. We rank by dot product. Is the ranking wrong?

No, the ranking is still right. A longer query multiplies every score by the same number, so the order does not change. What breaks is the meaning of the scores: they are no longer cosines, so any threshold such as 'above 0.8 is a duplicate' stops working. The dangerous case is unnormalised stored vectors, because each one is scaled differently.

Pause and think: In the run above the lowest mug score was 0.74 and the highest non-mug score was 0.47. We set the threshold at 0.6 and then swap in a different encoder. Can we keep 0.6?

Not safely. The gap between 0.47 and 0.74 belongs to this encoder and this data. Another model may bunch all its scores between 0.2 and 0.4, or between 0.8 and 0.95. We must score known matching and non-matching pairs again and pick a new threshold.

Summary

  • An embedding is a fixed-length vector where similar things are close together.
  • An image embedding comes from running a picture through a trained encoder (CNN or ViT) and taking a summary vector.
  • Raw pixels are huge, fragile and meaningless; embeddings are compact and robust to irrelevant changes.
  • Encoders are trained with labels, with image-caption pairs (CLIP-style) or self-supervised; the training decides what 'similar' means.
  • Compare embeddings with cosine similarity (or dot product on normalised vectors).
  • Uses: image search, text-to-image search, deduplication, recommendations, clustering and multimodal RAG.

Key takeaways

  • An image embedding is a compact, fixed-length vector that captures what an image shows.
  • It comes from a trained encoder; how that encoder was trained defines what 'similar' means.
  • Embeddings are robust to irrelevant changes (shift, lighting) where raw pixels are not.
  • Cosine similarity (a dot product on normalised vectors) is the standard comparison.
  • Never compare embeddings from different models, and tune thresholds on real examples.

Key terms

  • Embedding: A fixed-length vector representing an item so that similar items are close together.
  • Image encoder: A trained network (CNN or ViT) that turns an image into features or an embedding.
  • Cosine similarity: The cosine of the angle between two vectors: dot product divided by both lengths.
  • L2 normalisation: Dividing a vector by its length so it has length 1.
  • Contrastive learning: Training that pulls matching pairs together and pushes non-matching pairs apart.
  • Nearest-neighbour search: Finding the stored vectors most similar to a query vector.

← 16.2 Vision Transformers: Applying Self-Attention to Image Patches · 16.4 Diffusion Models: Iterative Denoising to Generate Images →