Lesson 16.6 · 26 min
Variational Autoencoders: Learning a Compressed Latent Space
A normal autoencoder can squeeze a photo into a few numbers and rebuild it — so why can it not invent a new photo when we hand it a few random numbers?
In short: An autoencoder compresses data into a small latent code and reconstructs it, but its latent space has gaps, so random codes decode to garbage. A Variational Autoencoder (VAE) fixes this by making the encoder output a small probability cloud (a mean and a spread) instead of a single point, and by adding a KL penalty that pulls all clouds toward a standard normal distribution. The result is a smooth, well-organised latent space we can sample from to generate new data; the reparameterization trick makes this trainable with backpropagation.
What is an autoencoder?
An autoencoder is a neural network trained to copy its input to its output through a narrow middle. It has two halves:
- The encoder squeezes the input (say a 64×64 product photo, 12,288 numbers) into a short vector called the latent code z (say 16 numbers).
- The decoder takes z and tries to rebuild the original input, producing a reconstruction x̂.
Training minimises the reconstruction error, for example the mean squared difference between x and x̂. Because z is much smaller than x, the network cannot simply copy; it must learn the most important structure of the data. The middle layer is called the bottleneck, and the space of all possible z vectors is the latent space.
Think of it like describing a photo over the phone We look at a kettle photo and tell a friend: 'steel, tall, black handle, side view, white background'. The friend draws it from that description. The description is the latent code; we are the encoder; the friend is the decoder. A short description forces us to mention only what matters.
Running example: our appliance shop has 20,000 product photos. An autoencoder trained on them learns a compact code for each kettle, toaster and blender. Such codes are useful for compression, denoising and anomaly detection (a damaged product reconstructs poorly because the network has never seen damage).
The problem with a normal autoencoder
The marketing team now asks: 'Can we generate new kettle designs?' A natural idea: pick a random latent code and run the decoder. With a plain autoencoder this usually produces garbage. Why?
- The latent space has holes. The encoder is only trained to give each training image a code that decodes well. Nothing tells it what codes between or around those points should mean. Large regions of the latent space are never visited, and the decoder has no idea what to do there.
- No known shape. The codes might cluster in strange, scattered patches far from zero. If we do not know where valid codes live, we cannot sample a good one.
- No smoothness. Two codes that are close together may decode to completely different images, so interpolating between two kettles can pass through nonsense.
Pause and think: Our plain autoencoder maps every kettle to codes between 40 and 60 in the first latent dimension. We sample a code with first dimension 0. What do we expect from the decoder?
Probably garbage. The decoder has never been trained on codes near 0 in that dimension, so its output there is undefined. That is exactly the 'holes in the latent space' problem a VAE is designed to fix.
What is a Variational Autoencoder?
The Variational Autoencoder (VAE), introduced by Kingma and Welling (2013) and independently by Rezende, Mohamed and Wierstra (2014), changes two things:
- The encoder outputs a distribution, not a point. For each input it gives a mean vector μ and a spread σ (standard deviation) for every latent dimension. The code z is then sampled from the Gaussian N(μ, σ²). So each image owns a small fuzzy cloud in latent space rather than a single dot.
- A regulariser organises the space. The loss includes a penalty — the KL divergence — that pulls every cloud toward a fixed prior, the standard normal N(0, I) (mean 0, spread 1 in every dimension).
Clouds fill the holes: because each image is decoded from many nearby samples during training, the decoder learns that the whole neighbourhood should look like that image. The KL penalty packs all clouds around the origin with overlapping edges, so the space between two kettles decodes to something kettle-like too. And to generate, we simply sample z ~ N(0, I) — we know exactly where valid codes live.
The encoder, the latent space, and the decoder
Two practical details. First, encoders usually output the log-variance log σ² rather than σ, because a network output can be any real number, while σ must be positive; σ = exp(0.5 · log σ²) is always positive. Second, the dimensions are treated as independent (a 'diagonal' Gaussian), which keeps the math cheap.
Once trained, the parts are used separately. To generate, throw the encoder away: sample z ~ N(0, I) and decode. To encode a real image (for search or editing), use only μ — the centre of its cloud. To edit, move z in a direction and decode: if one direction happens to correspond to 'handle colour', sliding along it changes just that.
The reparameterization trick
Here is a puzzle. Training uses backpropagation, which needs a gradient for every step from the loss back to the encoder's weights. But in the middle we sample z from N(μ, σ²). Sampling is random: there is no derivative of 'roll a die' with respect to μ. Gradients cannot pass through a random draw.
The trick is to move the randomness out of the way. Draw a fixed-distribution noise ε ~ N(0, I) that does not depend on any weights, then build z with ordinary arithmetic:
Think of it like a recipe with a dice roll on the side Instead of 'the chef randomly decides how much salt to add' (impossible to blame or tune), we say 'roll a die first; then add μ grams plus σ grams per pip'. The randomness is still there, but now we can ask exactly how changing μ or σ changes the dish — and that is all backpropagation needs.
Pause and think: μ = 0.8, σ = 0.5 and the noise draw is ε = −1.2. What is z, and what is ∂z/∂σ for this draw?
z = 0.8 + 0.5 × (−1.2) = 0.8 − 0.6 = 0.2. Since z = μ + σε, ∂z/∂σ = ε = −1.2.
The loss function of a Variational Autoencoder
The VAE loss for one input has two parts that pull in opposite directions:
- Reconstruction wants each cloud small and well separated, so the decoder knows exactly which image it came from.
- KL wants every cloud to look like N(0, 1): centred at 0 (penalising μ² ) with spread 1 (penalising σ² − log σ² − 1, which is 0 only at σ = 1).
- The balance gives clouds that are informative yet overlapping, centred near the origin: a smooth, sample-able space.
A simple example walk-through, in code
One training example, by the numbers
- Encode: The encoder looks at one 4-pixel image and outputs, for a 2-d latent, μ = [0.8, −0.3] and log σ² = [−1, −2], so σ ≈ [0.607, 0.368].
- Sample with the trick: Draw ε ~ N(0, I) and set z = μ + σ ⊙ ε. Every draw gives a slightly different z near μ.
- Decode: The decoder turns z into a reconstruction, here x̂ = [0.8, 0.2, 0.4, 0.6] for the input x = [0.9, 0.1, 0.4, 0.7].
- Score: Reconstruction error = 0.01 + 0.01 + 0 + 0.01 = 0.03. KL = ½[(0.64 + 0.368 + 1 − 1) + (0.09 + 0.135 + 2 − 1)] ≈ 1.117. Total ≈ 1.147 with β = 1.
- Update: Backpropagate through the decoder, through z = μ + σε, into the encoder, and adjust all weights to lower the total.
vae_pieces.py
import numpy as np
rng = np.random.default_rng(0)
# Pretend the encoder looked at one image and output these for a 2-d latent:
mu = np.array([0.8, -0.3]) # centre of the cloud for this image
log_var = np.array([-1.0, -2.0]) # log of the variance (can be any real number)
sigma = np.exp(0.5 * log_var)
print("mu =", mu, " sigma =", np.round(sigma, 3))
# Reparameterization trick: z = mu + sigma * eps, with eps ~ N(0, I)
eps = rng.normal(size=(3, 2))
z = mu + sigma * eps
print("three sampled z:\n", np.round(z, 3))
# KL divergence between N(mu, sigma^2) and the prior N(0, 1), closed form
def kl(mu, log_var):
return 0.5 * np.sum(mu**2 + np.exp(log_var) - log_var - 1)
print("KL to prior =", round(kl(mu, log_var), 3))
print("KL if mu=0, sigma=1 =", round(kl(np.zeros(2), np.zeros(2)), 3))
# Why the trick matters: the gradient flows through z = mu + sigma*eps.
# Check: estimate d/dmu E[z1^2] by averaging 2*z1 * dz1/dmu1 (=1) over samples
eps = rng.normal(size=100_000)
z1 = mu[0] + sigma[0] * eps
print("reparam gradient estimate:", round(np.mean(2 * z1), 3),
"| exact 2*mu1 =", 2 * mu[0])
# Full loss for one example = reconstruction error + beta * KL
x = np.array([0.9, 0.1, 0.4, 0.7]) # the input (4 pixels)
x_hat = np.array([0.8, 0.2, 0.4, 0.6]) # decoder output from a sampled z
recon = np.sum((x - x_hat) ** 2)
for beta in (1.0, 4.0):
print(f"beta={beta}: recon={recon:.3f} + KL term={beta * kl(mu, log_var):.3f}"
f" = loss {recon + beta * kl(mu, log_var):.3f}")Output:
mu = [ 0.8 -0.3] sigma = [0.607 0.368] three sampled z: [[ 0.876 -0.349] [ 1.188 -0.261] [ 0.475 -0.167]] KL to prior = 1.117 KL if mu=0, sigma=1 = 0.0 reparam gradient estimate: 1.599 | exact 2*mu1 = 1.6 beta=1.0: recon=0.030 + KL term=1.117 = loss 1.147 beta=4.0: recon=0.030 + KL term=4.466 = loss 4.496
In a full implementation (for example in PyTorch), the encoder and decoder are neural networks, the batch loss is averaged over many images, and an optimiser like Adam updates both networks together. The three key lines are always the same: z = mu + exp(0.5log_var) randn_like(mu), the KL formula above, and loss = recon + beta * kl.
Advantages of Variational Autoencoders
- Stable, simple training. One network pair and one loss; no adversarial game as in GANs.
- A smooth, structured latent space. Interpolations look sensible, and directions in latent space can correspond to meaningful attributes.
- Fast generation. One decoder pass per sample, unlike the many steps of diffusion.
- An encoder for free. Unlike a GAN, a VAE can map a real image into the latent space, which enables editing, search and anomaly detection.
- A principled probabilistic model. The ELBO gives a lower bound on the data likelihood, useful for comparing models and for anomaly scores.
Limits and common mistakes Samples from a plain VAE are often blurry, partly because a squared-error loss rewards averaging over plausible details. If the decoder is very powerful, the model may ignore z altogether ('posterior collapse': KL drops to 0 and all clouds equal the prior); KL warm-up or annealing helps. Forgetting the KL term turns the model back into a plain autoencoder; setting β too high gives blurry, generic outputs. And for top-quality open-ended image generation, diffusion models usually win.
Where Variational Autoencoders are used
| Use | How the VAE helps |
|---|---|
| Latent diffusion (e.g. Stable Diffusion) | A VAE-style autoencoder compresses images to a small latent grid; diffusion runs there, and the decoder turns the result back into pixels |
| Anomaly detection | Inputs that reconstruct badly or have low ELBO are flagged (damaged products, faulty sensor readings, fraud) |
| Data generation and augmentation | Sample new examples, especially for tabular, time-series or molecular data |
| Drug and molecule design | Search a smooth latent space of molecules for ones with desired properties |
| Representation learning | Compact, smooth features for clustering and downstream models |
| Discrete tokenizers (VQ-VAE family) | Vector-quantised autoencoders turn images or audio into discrete tokens that transformers can model |
Back to the shop We could train a VAE on photos of intact products only. At inspection time, a photo of a cracked jar reconstructs poorly and gets a high loss, so it is flagged for a human — without ever collecting labelled examples of every kind of damage.
Common mistakes and how to spot them
A VAE that trains without errors can still be broken. The most useful habit is to log the two loss parts separately, and to log the KL per latent dimension, averaged over a batch. One total number hides nearly everything.
A few unused dimensions are normal: the model only keeps as many as it needs. All dimensions at zero is posterior collapse. Here is a second, very common bug, with numbers. Our photos have 64 × 64 × 3 = 12,288 values and the latent has 16. The loss formula sums squared error over all 12,288 values and sums KL over the 16 dimensions. If our code instead takes the mean over pixels but still sums the KL, the reconstruction term becomes 12,288 times smaller. That is the same as training with β = 12,288. The KL term wins, and the clouds collapse onto the prior.
| What we see | Likely cause | What to check |
|---|---|---|
| KL ≈ 0 in every dimension; all samples look alike | Posterior collapse | KL per dimension; is β too large, or is one term averaged and the other summed? |
| Sharp reconstructions, but random samples are garbage | KL weight far too small; the model behaves like a plain autoencoder | Mean and spread of the μ values over the dataset: far from 0 and 1? |
| Everything blurry, even reconstructions | β too high, or the latent is too small | Reconstruction loss on its own; try a lower β |
| σ blows up or the loss becomes NaN | exp of a large log-variance overflowed | Range of log σ² values; clamp them to a sensible range |
| Search results change every time we embed the same photo | Using a sampled z instead of μ at inference | Encode twice and compare the two codes |
A quick health check after training
- Reconstruct: Encode and decode a few held-out photos. This tests the encoder and decoder together.
- Sample: Decode z ~ N(0, I) a few dozen times. This tests whether the prior and the learned clouds actually line up.
- Interpolate: Walk in a straight line between the μ of two photos and decode along the way. Sudden jumps or nonsense in the middle mean holes.
- Count active dimensions: Average the KL per dimension over a batch. That tells us how many latent numbers the model really uses.
Practice: try it yourself
We will measure the 'holes' directly. Five product photos live in a 1-d latent space, once as a plain autoencoder would place them and once as a VAE would. We then draw 20,000 codes from the prior N(0, 1), exactly as we do when generating, and count how many land somewhere the decoder has seen before.
practice_latent_holes.py
import numpy as np
rng = np.random.default_rng(5)
# A 1-d latent space holding five product photos.
# Plain autoencoder: five sharp points, placed wherever training left them.
# VAE: five wide clouds, pulled toward the prior N(0, 1) by the KL term.
models = {
"plain AE": (np.array([-6.0, -2.5, 0.5, 3.0, 7.0]), 0.05),
"VAE": (np.array([-1.2, -0.6, 0.0, 0.6, 1.2]), 0.50),
}
def kl(mu, sigma):
"""KL( N(mu, sigma^2) || N(0, 1) ) for one latent dimension."""
return 0.5 * (mu**2 + sigma**2 - np.log(sigma**2) - 1)
z_prior = rng.normal(size=20_000) # what we sample at generation
for name, (mu, sigma) in models.items():
# A prior sample is "known" to the decoder if it lies inside some cloud
# (within 2 sigma of a centre): the decoder saw codes like it in training.
known = (np.abs(z_prior[:, None] - mu) < 2 * sigma).any(axis=1).mean()
# The codes seen in training: pick an image, then sample z with the trick
pick = rng.integers(0, 5, size=20_000)
z_train = mu[pick] + sigma * rng.normal(size=20_000)
mid = (mu[1] + mu[2]) / 2 # halfway between two photos
mid_known = bool((np.abs(mid - mu) < 2 * sigma).any())
print(f"{name}")
print(f" prior samples the decoder knows : {known:.0%}")
print(f" training codes: mean {z_train.mean():+.2f}, std {z_train.std():.2f}")
print(f" midpoint of photos 2 and 3 known: {mid_known}")
print(f" average KL per photo : {kl(mu, sigma).mean():.2f}")Output:
plain AE prior samples the decoder knows : 7% training codes: mean +0.38, std 4.46 midpoint of photos 2 and 3 known: False average KL per photo : 12.55 VAE prior samples the decoder knows : 97% training codes: mean -0.01, std 0.98 midpoint of photos 2 and 3 known: True average KL per photo : 0.68
Now change it:
- Give the VAE a spread of
0.05instead of0.50, keeping its centres. Predict the 'known' percentage and whether the average KL goes up or down. - Give the plain AE the VAE's centres but keep its spread of
0.05. Predict: is putting the points in the right place enough to fill the holes? - Set all five VAE centres to
0.0and the spread to1.0. Predict the KL and the 'known' percentage. Then explain why this 'perfect' score is actually a failure.
Pause and think: In the run above, the VAE's training codes had mean −0.01 and std 0.98. Why does that matter for generating new photos?
At generation time we feed the decoder codes drawn from N(0, 1). The decoder only works well on codes like the ones it saw in training. If the training codes, taken all together, also look like N(0, 1), then prior samples are familiar territory. The plain autoencoder's codes had std 4.46 with empty gaps between them, so most prior samples fall where the decoder was never trained.
Pause and think: A teammate's VAE code computes the reconstruction loss as a mean over all 12,288 pixel values and the KL as a sum over 16 latent dimensions. Training runs smoothly, KL falls to almost 0, and every sample is the same grey blur. What went wrong?
The two terms are on different scales. Taking the mean divides the reconstruction term by 12,288, so the KL term is thousands of times heavier than intended, like a very large β. The cheapest way to lower the loss is to make every cloud equal the prior, which removes all information from z: posterior collapse. The fix is to reduce both terms the same way (sum both per image), or to lower β to compensate.
Key takeaways
- An autoencoder compresses data to a latent code and reconstructs it, but its latent space has holes.
- A VAE encodes each input as a Gaussian cloud N(μ, σ²) and pulls all clouds toward N(0, I) with a KL penalty.
- The reparameterization trick, z = μ + σ ⊙ ε, lets gradients flow through the sampling step.
- Loss = reconstruction + β·KL; β trades reconstruction detail for a more organised latent space.
- VAEs give fast generation and an encoder; they power latent diffusion, anomaly detection and more.
Key terms
- Autoencoder: A network that compresses input to a latent code and reconstructs it.
- Latent space: The space of codes z produced by the encoder and consumed by the decoder.
- Prior: The distribution we want latent codes to follow, usually the standard normal N(0, I).
- KL divergence: A measure of how different one probability distribution is from another; zero when they are equal.
- Reparameterization trick: Writing z = μ + σ ⊙ ε with ε ~ N(0, I) so that sampling becomes differentiable.
- ELBO: Evidence Lower Bound: the training objective of a VAE, a lower bound on the log-likelihood.
- Posterior collapse: A failure where the decoder ignores z and the encoder's output equals the prior.
← 16.5 GANs: A Generator and Discriminator in Constant Competition · 17.1 GPUs for Deep Learning: Parallelism at the Core →