Lesson 2.9 · 27 min
Contrastive Learning: Training by Comparison
How can a model learn that two photos show the same dog, or that a caption describes an image, when nobody has labelled a single example?
In short: Contrastive learning trains an encoder to map inputs to embeddings so that related pairs (positives) end up close together and unrelated pairs (negatives) end up far apart. Positive pairs usually come for free, from two augmented views of the same image or from an image and its caption, so huge unlabelled datasets can be used. Losses such as InfoNCE turn this into a "pick the true partner out of the batch" task, and methods like SimCLR, MoCo and CLIP use it to build the embeddings behind semantic search, RAG and multimodal AI.
What is contrastive learning?
Contrastive learning is a way of training a model to produce useful embeddings (lists of numbers that represent an input) by comparing examples. We show the model pairs of things and tell it only one fact about each pair: are these two related, or not? The model learns to place related things close together in embedding space and unrelated things far apart.
The model being trained is called an encoder: a network that turns an input (an image, a sentence, an audio clip) into a vector. After training, we usually throw away the training task and keep the encoder, because its embeddings capture meaning: similar inputs get similar vectors.
Spot the same person Show a child two photos of their aunt, one in sunlight and one at night with a hat, and a photo of a stranger. The child learns that lighting and hats do not matter, while face shape does. Nobody explains "face shape" in words; the child learns it by contrasting same-person pairs with different-person pairs. Contrastive learning teaches an encoder the same way: what stays the same across positives is what matters.
Contrastive learning is usually a form of self-supervised learning (Lesson 2.2): the training signal comes from the structure of the data itself, not from human labels. It can also be supervised, when we use labels to decide which pairs are positive.
Why do we need contrastive learning?
Supervised learning needs labels, and labels are expensive. Meanwhile the world is full of unlabelled data: billions of images, web pages, and image-caption pairs. We want a way to learn good representations from all of it.
We also need representations that capture meaning. Early approaches described data with hand-made features (Lesson 2.4) or with one-hot encoding, where each word or category gets its own 0/1 column. One-hot vectors treat every pair of items as equally different: "cat" is as far from "kitten" as it is from "carburettor". Learned embeddings fix this by placing similar things nearby, and contrastive learning is one of the most effective ways to learn them.
- Uses unlabelled data: positives can be created automatically.
- Produces general-purpose embeddings: one encoder can serve search, clustering, classification and recommendation.
- Needs few labels downstream: a small classifier on top of good embeddings (a "linear probe") can work well with little labelled data.
- Connects different modalities: it can put images and text in the same space, so a sentence can find a photo.
Background For the background ideas mentioned here, feature engineering and one-hot encoding, see Lesson 2.4.
The key idea: pull together, push apart
Every contrastive method follows one principle: similar pairs should have similar embeddings; dissimilar pairs should have dissimilar embeddings. We measure similarity between two embeddings, usually with cosine similarity (the cosine of the angle between the vectors, from −1 to 1), after normalising each embedding to length 1.
Why do we need the "push apart" half? Without it, the encoder can cheat: map every input to the same vector. Then every positive pair is perfectly similar and the loss looks great, but the embeddings are useless. This failure is called representation collapse. Negatives prevent it by requiring different things to stay different.
There is also a deeper effect: by deciding what counts as a positive, we decide what the model should ignore. If two crops of the same photo with different colours are positives, the model learns that crop position and colour shifts do not change meaning. These are called invariances.
Positive pairs and negative pairs
A positive pair is two inputs that should be close in embedding space. A negative pair is two inputs that should be far apart. In a typical setup we pick an anchor example, one positive for it, and many negatives.
| Setting | Positive pair | Negatives |
|---|---|---|
| Images, self-supervised (e.g. SimCLR) | Two random augmentations (crop, flip, colour jitter, blur) of the same photo | Augmentations of other photos in the batch |
| Image and text (e.g. CLIP) | An image and its own caption | The other captions in the batch |
| Text search embeddings | A question and a passage that answers it | Other passages in the batch, plus "hard" passages that look similar but do not answer |
| Sentences (e.g. SimCSE) | The same sentence encoded twice with different dropout noise | Other sentences in the batch |
| Face recognition | Two photos of the same person | Photos of other people |
| Supervised contrastive | Two examples with the same class label | Examples with different labels |
A very common trick is in-batch negatives: with a batch of N positive pairs, each anchor uses its own partner as the positive and the other N − 1 partners as negatives. Negatives come for free, which is why bigger batches often help.
Hard negatives are negatives that look similar to the anchor but are not related, such as a passage about Python the snake for a question about Python the language. They teach fine distinctions. Random negatives are often too easy to teach much.
False negatives With in-batch negatives, two different examples might actually be related, for example two different photos of golden retrievers in the same batch. The loss then pushes apart things that should be close. This is usually tolerable in large, diverse datasets, but it hurts with small or repetitive ones, and some methods try to detect and remove such false negatives.
Pause and think: Our batch has 64 image-caption pairs and we use in-batch negatives. How many negatives does each image get?
63: every caption in the batch except its own. With a batch of 32,768 (the size used for CLIP), each image would get 32,767 negatives.
How contrastive learning works, step by step
One training step (SimCLR-style, for images)
- Sample a batch: Take N unlabelled images, for example N = 256.
- Create two views: Apply two random augmentations to each image (random crop + colour change, etc.). Now we have 2N images; the two views of the same photo form a positive pair.
- Encode: Pass every view through the encoder (e.g. a ResNet or Vision Transformer) to get a representation, then through a small projection head to get an embedding.
- Normalise and compare: Normalise each embedding to length 1 and compute cosine similarities between all pairs: a big similarity matrix.
- Compute the loss: For each view, apply a softmax over its similarities (divided by a temperature) and ask the model to give high probability to its true partner. This is the InfoNCE loss.
- Update: Backpropagate and update the encoder so positive similarities rise and negative similarities fall. Repeat for many batches.
- Keep the encoder: After training, discard the projection head and use the encoder's representations for downstream tasks.
Loss functions used in contrastive learning
Three loss families dominate. Let d be a distance between embeddings and sim a similarity.
Code: training a tiny encoder with InfoNCE
Eight items each have 4 "content" numbers that identify them, plus 12 "nuisance" numbers that change wildly between views (think lighting and background). We train a linear encoder from 16 numbers to a 4-D embedding with InfoNCE. Nobody tells it which dimensions matter; it must discover that content is shared across views and nuisance is not.
tiny_infonce.py
import numpy as np
rng = np.random.default_rng(0)
n, tau = 8, 0.2
content = rng.normal(size=(n, 4)) # what makes each item unique
noise_scale = np.r_[np.full(4, 0.2), np.full(12, 2.0)] # 12 "nuisance" dims vary a lot
def views(): # two random augmentations of the same 8 items (like crops/colour jitter)
base = np.hstack([content, np.zeros((n, 12))])
return [base + rng.normal(size=base.shape) * noise_scale for _ in range(2)]
W = rng.normal(scale=0.1, size=(16, 4)) # the encoder we train: 16 -> 4
def embed(X): # linear encoder + L2-normalise
U = X @ W
norm = np.linalg.norm(U, axis=1, keepdims=True)
return U / norm, norm
def info_nce(Z1, Z2): # row i of Z1 should pick row i of Z2
logits = Z1 @ Z2.T / tau # cosine similarity / temperature
P = np.exp(logits - logits.max(axis=1, keepdims=True))
P /= P.sum(axis=1, keepdims=True) # softmax over the 8 candidates
return -np.log(np.diag(P)).mean(), P
def back_norm(dZ, Z, norm): # gradient through U / |U|
return (dZ - Z * (dZ * Z).sum(axis=1, keepdims=True)) / norm
for step in range(401):
X1, X2 = views()
(Z1, n1), (Z2, n2) = embed(X1), embed(X2)
loss, P = info_nce(Z1, Z2)
if step % 100 == 0:
S = Z1 @ Z2.T
print(f"step {step:3d} loss {loss:.3f} correct matches {(P.argmax(1) == np.arange(n)).sum()}/8"
f" positive sim {np.diag(S).mean():+.2f} negative sim {S[~np.eye(n, dtype=bool)].mean():+.2f}")
dL = (P - np.eye(n)) / (n * tau) # gradient of the loss w.r.t. Z1 @ Z2.T
W -= 0.1 * (X1.T @ back_norm(dL @ Z2, Z1, n1) + X2.T @ back_norm(dL.T @ Z1, Z2, n2))
print(f"weight on content dims {np.abs(W[:4]).mean():.3f}, on nuisance dims {np.abs(W[4:]).mean():.3f}")
print(f"random-guess loss with 8 candidates = ln(8) = {np.log(n):.3f}")Output:
step 0 loss 4.488 correct matches 1/8 positive sim +0.02 negative sim -0.04 step 100 loss 1.111 correct matches 5/8 positive sim +0.82 negative sim -0.10 step 200 loss 0.704 correct matches 7/8 positive sim +0.85 negative sim -0.02 step 300 loss 0.621 correct matches 7/8 positive sim +0.89 negative sim -0.06 step 400 loss 0.535 correct matches 7/8 positive sim +0.87 negative sim -0.09 weight on content dims 1.186, on nuisance dims 0.060 random-guess loss with 8 candidates = ln(8) = 2.079
Pause and think: If we removed the negatives and only maximised the similarity of positive pairs, what could the encoder learn instead?
It could collapse: map every input to the same vector, making every positive pair perfectly similar while the embeddings carry no information. The negatives (the denominator of InfoNCE) are what force different items apart. Methods without negatives, like BYOL, need other tricks to avoid collapse.
Popular contrastive learning methods
Key methods
- Contrastive loss: Hadsell, Chopra and LeCun introduce a margin-based pairwise loss for learning embeddings.
- FaceNet: Google trains face embeddings with triplet loss; same-person faces cluster together.
- CPC and InfoNCE: Contrastive Predictive Coding introduces the InfoNCE loss.
- MoCo: Momentum Contrast (Facebook AI) keeps a large queue of negatives encoded by a slowly updated "momentum" encoder, so it does not need huge batches.
- SimCLR: A simple framework (Google): strong augmentations, a projection head, NT-Xent loss and large batches.
- BYOL: DeepMind shows good representations can be learned without negatives, using an online and a target network.
- CLIP and SimCSE: CLIP (OpenAI) aligns images and text from about 400 million pairs. SimCSE uses dropout noise as augmentation for sentence embeddings.
Real-world use cases, and limits
- Semantic search and RAG: text embedding models are commonly trained contrastively on (query, relevant passage) pairs with in-batch and hard negatives. These embeddings power vector databases and retrieval in RAG (Module 10).
- Multimodal search: CLIP-style models let users search photos with words ("red sneakers on a beach") and are used as components in many text-to-image and vision-language systems (Module 16).
- Face verification: unlocking a phone or matching ID photos compares face embeddings learned with contrastive or triplet-style losses.
- Recommendation: users and items are embedded so that a user is near items they engaged with.
- Pre-training with few labels: in medical imaging or industrial inspection, contrastive pre-training on unlabelled images followed by a small labelled fine-tune can beat training from scratch.
- Duplicate and near-duplicate detection: finding repeated support tickets, copied images or near-identical documents.
Common pitfalls Weak or wrong augmentations teach the wrong invariances (if colour matters for your task, do not use colour jitter as an augmentation). Too few or too easy negatives give weak embeddings. False negatives in small or repetitive datasets push related items apart. Temperature is a sensitive setting. And contrastive embeddings reflect the data they were trained on, including its biases.
When not to use it: if you have plenty of labelled data for one narrow task, plain supervised training is simpler. If an off-the-shelf embedding model already works for your domain, use it rather than training your own; fine-tune it contrastively on your own (query, passage) pairs only when retrieval quality on your data is not good enough.
Common mistakes and how to spot them
A contrastive training run can fail quietly. The loss goes down, nothing crashes, and the embeddings are still poor. The good news is that a few cheap numbers reveal most problems. After each evaluation we print three things: the mean positive similarity, the mean negative similarity, and the loss compared with ln(N), the loss of a blind guess among N candidates.
| What we see | Likely cause | What to try |
|---|---|---|
| Positive and negative similarity are both near 1; loss sits at ln(N) | Collapse: every input maps to almost the same vector | Check that negatives are really in the loss; check the normalisation step; lower the learning rate |
| Loss falls fast, but search quality on real queries stays poor | Negatives are too easy, so the task is solved without learning fine detail | Add hard negatives; use a larger batch |
| Loss stalls well above zero; some true pairs never match | False negatives: related items are being pushed apart | Remove duplicates from each batch; skip negatives that score suspiciously high |
| Good on training pairs, poor on a new kind of input | The positives taught the wrong invariances | Rethink the augmentations or the way pairs are built |
| Training is unstable, or a few pairs dominate every update | Temperature too low | Raise τ a little and compare |
| Positives and negatives stay close together for a long time | Temperature too high, so the push is weak and spread thin | Lower τ a little and compare |
Why does temperature matter so much? In InfoNCE, each negative is pushed away in proportion to the probability the softmax gives it. A low τ puts almost all of that probability on the negatives closest to the anchor. That is useful when those are true hard negatives. It is harmful when one of them is a false negative, because nearly the whole push then lands on an item that should have stayed close. The practice code below shows this with four numbers.
A five-minute sanity check before a long run
- Look at ten pairs by eye: Print ten positives and a few negatives for each. If we cannot tell why a pair is positive, the model cannot either.
- Check the starting loss: With a fresh encoder it should be near ln(N), or above it at a low temperature. A value far below means the pairs leak an easy shortcut.
- Overfit one small batch: Train on a single fixed batch. The loss should drop close to zero. If it cannot, there is a bug in the loss or the gradient.
- Track the two similarities: Positive similarity should rise while negative similarity stays low. If both rise together, collapse has begun.
- Test on the real task: Measure retrieval on held-out queries, not just the loss. The loss depends on batch size and τ, so it is not comparable across runs.
Practice: try it yourself
We will compute InfoNCE by hand for one anchor with four candidates: its positive, one hard negative and two easy negatives. No training, just the loss. We try two temperatures and also print how the push on the negatives is shared out. Then we feed in a collapsed encoder, where every similarity is 1.
practice_infonce_temperature.py
import numpy as np
def info_nce(sims, tau):
"""sims[0] is the positive; the rest are negatives. Returns the loss
and the softmax probability given to every candidate."""
logits = np.array(sims) / tau
p = np.exp(logits - logits.max())
p /= p.sum()
return -np.log(p[0]), p
cases = {
"trained encoder": [0.8, 0.7, 0.1, 0.0], # one negative is almost as close
"collapsed encoder": [1.0, 1.0, 1.0, 1.0], # every input maps to one vector
}
for label, sims in cases.items():
print(f"{label}: similarities {sims}")
for tau in (1.0, 0.1):
loss, p = info_nce(sims, tau)
# The gradient pushes each negative away in proportion to its probability
push = p[1:] / p[1:].sum()
print(f" tau={tau:<4} loss={loss:.3f} P(positive)={p[0]:.3f}"
f" share of push on negatives: {np.round(push, 3).tolist()}")
print(f"ln(4) = {np.log(4):.3f} (loss when all 4 candidates look the same)")Output:
trained encoder: similarities [0.8, 0.7, 0.1, 0.0] tau=1.0 loss=1.048 P(positive)=0.351 share of push on negatives: [0.489, 0.268, 0.243] tau=0.1 loss=0.314 P(positive)=0.730 share of push on negatives: [0.997, 0.002, 0.001] collapsed encoder: similarities [1.0, 1.0, 1.0, 1.0] tau=1.0 loss=1.386 P(positive)=0.250 share of push on negatives: [0.333, 0.333, 0.333] tau=0.1 loss=1.386 P(positive)=0.250 share of push on negatives: [0.333, 0.333, 0.333] ln(4) = 1.386 (loss when all 4 candidates look the same)
Now change it:
- Make the hard negative even harder: change
0.7to0.79. Predict whether the loss at τ = 0.1 goes up or down, and roughly what P(positive) becomes when two candidates are almost tied. - Add
0.05to the temperatures. Predict what happens to the hard negative's share of the push, and to the loss of the collapsed encoder. - Add ten more easy negatives with similarity
0.0to the trained case. Predict which temperature's loss changes more, and what the collapsed loss would be with the same 14 candidates.
Pause and think: At τ = 0.1 the hard negative receives 99.7% of the push, compared with 48.9% at τ = 1.0. Why is this a strength and also a risk?
A strength, because the two easy negatives are already far away and teach nothing, so focusing the update on the one confusing candidate is efficient. A risk, because if that "hard negative" is actually a related item (a false negative), almost the entire update goes into pushing apart two things that belong together. Low temperatures need clean negatives.
Pause and think: For the collapsed encoder, the loss is 1.386 at both temperatures and the push is shared equally. Why can temperature not help here?
All four similarities are equal, so dividing them by any τ still gives four equal logits, and the softmax gives each candidate 0.25. The loss is −ln(0.25) = ln(4) ≈ 1.386 whatever τ is. A loss stuck at ln(N) with similarities near 1 is the fingerprint of collapse; the fix lies in the encoder and training setup, not in the temperature.
Key takeaways
- Contrastive learning trains an encoder so positives are close and negatives are far apart in embedding space.
- Positives usually come for free: two augmentations of one input, or naturally paired data like an image and its caption.
- Negatives prevent representation collapse; in-batch and hard negatives are common sources.
- InfoNCE turns learning into "pick the true partner among N candidates"; temperature τ controls its sharpness.
- SimCLR, MoCo and CLIP are landmark methods; contrastive embeddings power semantic search, RAG and multimodal AI.
Key terms
- Contrastive learning: Learning embeddings by pulling related pairs together and pushing unrelated pairs apart.
- Positive pair: Two inputs that should have similar embeddings, such as two views of the same image.
- Negative pair: Two inputs that should have dissimilar embeddings.
- InfoNCE: A softmax-based contrastive loss that asks the model to identify the positive among many negatives.
- Temperature (τ): A scale applied to similarities before the softmax; smaller values make the distribution sharper.
- Representation collapse: A failure where the encoder maps all inputs to nearly the same embedding.
- Hard negative: A negative example that looks similar to the anchor and is therefore informative to train on.
← 2.8 Reinforcement Learning: Teaching Agents Through Reward · 3.1 Neural Network Bias: What It Is and Why It Matters →