Modern AI Engineering

Lesson 3.5 · 23 min

Dropout: Controlled Forgetting as Regularization

Why would switching off random neurons on purpose, every single training step, make a neural network smarter?

In short: Dropout is a regularisation technique: during training, each neuron's output is set to zero with probability p, and the survivors are scaled up by 1/(1 − p). This stops neurons from relying on specific partners and acts like training many thinned networks at once, which reduces overfitting. At test time dropout is switched off and the full network is used.

What is dropout?

Dropout is a simple trick used while training neural networks. On every training step, we pick a random subset of neurons in a layer and temporarily set their outputs to zero, as if they had been removed from the network. On the next step we pick a different random subset. The fraction we drop is a hyperparameter called the dropout rate, written p (for example p = 0.5 drops about half).

Dropout was introduced by Geoffrey Hinton's group at the University of Toronto around 2012 and described in detail by Srivastava and colleagues in 2014. It quickly became one of the standard tools for making networks generalise better.

Think of it like a team where anyone might be absent Imagine a restaurant kitchen where, every day, a random half of the cooks call in sick. No dish can depend on one specific cook knowing one secret step. Every cook has to be broadly capable, and important knowledge gets spread across the team. On the day everyone shows up (test time), the kitchen is extra robust. Dropout forces neurons to work the same way.

The problem of overfitting

Overfitting happens when a model learns the training data too specifically, including its noise and accidents, instead of the general pattern. The symptom is a gap: the loss on the training set keeps falling, while the loss on a held-out validation set (data the model never trains on) stops improving and starts rising.

Our running example: a network that classifies customer support messages into topics (billing, shipping, login problems) using only 2,000 labelled messages. A big network can simply memorise all 2,000. It might learn that any message containing a particular customer's name is about billing, because that customer happened to write three billing messages. That rule is useless on new messages.

Techniques that fight overfitting are called regularisation. Others include getting more data, data augmentation, weight decay (L2), early stopping and smaller models. Dropout is a regulariser designed specifically for neural networks.

Why do we need dropout?

There are two complementary explanations for why dropout helps.

1. It breaks co-adaptation. In a normal network, neurons can form fragile partnerships: neuron A only produces something useful if neuron B fixes up its mistakes. These partnerships often encode quirks of the training data. With dropout, B might be missing at any step, so A must learn features that are useful on their own, in many different contexts.

2. It is a cheap ensemble. An ensemble combines the predictions of several separately trained models, which usually beats any single one. With n neurons that can each be dropped, there are 2ⁿ possible "thinned" sub-networks. Each training step trains one random sub-network, and they all share weights. At test time, using the full network with proper scaling approximates averaging the predictions of all these sub-networks, without training them separately.

Pause and think: A hidden layer has 10 neurons with dropout. How many different thinned versions of that layer can training sample?

2¹⁰ = 1,024, because each neuron is independently either kept or dropped. For a layer of 1,000 neurons the number is astronomically large, which is why we call it an implicit ensemble rather than literally training each one.

How does dropout work?

Dropout is applied to the outputs (activations) of a layer, usually after the activation function. For each activation, we draw a random number; if it falls below p, that activation becomes 0. The random pattern of kept and dropped units is called the mask. A fresh mask is drawn for every example and every training step.

Why scale by 1/(1 − p)? If we drop half the neurons, the next layer receives on average only half the total signal. At test time, with nobody dropped, it would suddenly receive twice as much, and its learned weights would be miscalibrated. Scaling survivors by 1/(1 − p) during training makes the expected value of each activation equal to its original value: E[h̃ᵢ] = (1 − p) · hᵢ / (1 − p) = hᵢ. So the full network at test time sees signals of the same size it was trained on.

A step-by-step example

Dropout with p = 0.5 on six activations

  1. Start with activations: A hidden layer outputs h = [1, 2, 3, 4, 5, 6]. The sum is 21.
  2. Draw a mask: Each unit is kept with probability 0.5. Suppose the mask is [1, 0, 0, 0, 1, 1].
  3. Zero the dropped units: h · m = [1, 0, 0, 0, 5, 6]. Sum is now 12.
  4. Scale the survivors: Multiply by 1/(1 − 0.5) = 2: [2, 0, 0, 0, 10, 12]. Sum 24, close to 21. On average across many masks the sum is exactly 21.
  5. Backward pass: Gradients flow only through the kept units (multiplied by the same factor 2). Dropped neurons get no update from this example on this step.
  6. Next step, new mask: Another mask, another sub-network. Over thousands of steps, every neuron is trained in many different company.

The code below runs exactly this, then checks the expectation argument with 100,000 random masks.

dropout_demo.py

import numpy as np
rng = np.random.default_rng(0)
def dropout(x, p, training):
"""Inverted dropout: drop with probability p, scale survivors by 1/(1-p)."""
if not training or p == 0:
return x                              # evaluation: identity, no scaling
keep = rng.random(x.shape) >= p           # True = neuron survives
return x * keep / (1 - p)
h = np.array([1.0, 2.0, 3.0, 4.0, 5.0, 6.0])  # activations of one hidden layer
print("input        ", h)
for trial in range(3):
print(f"train pass {trial}", dropout(h, p=0.5, training=True))
print("eval pass    ", dropout(h, p=0.5, training=False))
# Inverted scaling keeps the expected value the same as at evaluation time
many = np.stack([dropout(h, 0.5, True) for _ in range(100_000)])
print("mean over 100k train passes", many.mean(axis=0).round(2))
# What goes wrong without the 1/(1-p) scaling
no_scale = np.stack([h * (rng.random(h.shape) >= 0.5) for _ in range(100_000)])
print("mean without scaling       ", no_scale.mean(axis=0).round(2))

Output:

input         [1. 2. 3. 4. 5. 6.]
train pass 0 [ 2.  0.  0.  0. 10. 12.]
train pass 1 [ 2.  4.  6.  8. 10.  0.]
train pass 2 [ 2.  0.  6.  0. 10. 12.]
eval pass     [1. 2. 3. 4. 5. 6.]
mean over 100k train passes [1.   1.99 3.   4.01 4.98 6.02]
mean without scaling        [0.5  1.   1.5  2.01 2.5  2.99]

Dropout during training vs testing

Common mistake: evaluating in training mode If you forget model.eval() in PyTorch before validation or inference, dropout stays on. Predictions become random, accuracy looks worse than it is, and the same input gives different answers each time. Equally, forgetting to switch back with model.train() silently turns regularisation off for the rest of training.

A deliberate exception is Monte Carlo dropout: keeping dropout on at test time, running the same input many times and looking at the spread of the predictions as an estimate of the model's uncertainty. It is a known technique, but it is a conscious choice, not the default.

Pause and think: You validate your model twice on the same data and get 87.2% then 86.5% accuracy. Nothing changed in between. What is the most likely cause?

Dropout (or another random layer) is still active because the model is in training mode. In evaluation mode the network is deterministic, so the same data must give the same accuracy. Call model.eval() before validating.

Dropout in code with a framework

In practice we never write the mask ourselves; frameworks provide a dropout layer. Here is the usual PyTorch pattern for our support-ticket classifier (PyTorch is not installed in this course sandbox, so this snippet is shown without output; the numpy demo above shows the same mechanism running).

classifier.py (PyTorch)

import torch.nn as nn
model = nn.Sequential(
nn.Linear(300, 128),
nn.ReLU(),
nn.Dropout(p=0.5),       # drops 50% of the 128 activations in training
nn.Linear(128, 3),       # 3 topics: billing, shipping, login
)
model.train()   # dropout active  -> use while fitting
model.eval()    # dropout is a no-op -> use for validation and inference

Watch the convention PyTorch nn.Dropout(p) and Keras Dropout(rate) both take the probability of dropping. Some older papers and code (including the original TensorFlow 1 keep_prob argument) used the probability of keeping. A value of 0.9 means opposite things under the two conventions.

Variants of dropout

Common variants and what they drop
VariantWhat gets droppedTypical use
Standard dropoutIndividual activationsFully connected layers, Transformer sub-layers
DropConnectIndividual weights instead of activationsResearch on dense layers; less common in practice
Spatial dropout (Dropout2d)Entire feature maps (channels) of a CNNConvolutional networks, where neighbouring pixels are strongly correlated
Variational / recurrent dropoutThe same mask reused at every time step of a sequenceRNNs and LSTMs
Attention dropoutEntries of the attention-weight matrixTransformers
DropPath / stochastic depthWhole residual branches of a blockDeep residual networks and vision Transformers
Monte Carlo dropoutStandard dropout kept on at inferenceUncertainty estimation

They all share the core idea: inject structured random noise during training so the network cannot depend on any single path, then remove the noise for inference.

Advantages, where dropout is used, and when not to use it

  • Simple: one line of code and one hyperparameter.
  • Cheap: almost no extra compute, and zero cost at inference with inverted dropout.
  • Effective: often noticeably reduces overfitting, especially for large fully connected layers and small datasets.
  • Works alongside other regularisers such as weight decay and data augmentation.

Where dropout is used Classic image networks such as AlexNet and VGG used dropout with p = 0.5 in their large fully connected layers. The original Transformer applied dropout of 0.1 to the outputs of each sub-layer and to the embeddings, and Transformer-based models such as BERT use dropout of about 0.1 as well. Smaller models fine-tuned on modest datasets commonly keep dropout on. Many very large language models are pretrained with dropout set very low or turned off, because each training example is seen only about once and overfitting is less of a concern; exact settings vary by model and are not always published.

Typical rates: 0.1–0.3 for Transformers and convolutional layers, up to 0.5 for big dense layers. Rates above about 0.5 usually hurt because too little signal gets through.

When not to use it, or to use less: when the model is underfitting (training loss itself is high), dropout makes things worse. With huge datasets seen once, it may add little. Mixing dropout with Batch Normalization in the same block can cause a mismatch between training and test statistics, so many CNNs that rely on BatchNorm use little or no dropout in convolutional layers. Never apply dropout at inference unless you want Monte Carlo dropout on purpose.

Going one level deeper

We showed that inverted dropout keeps the average of each activation unchanged. But an average hides how much a single pass jumps around. Let us measure that jump, because it explains why high drop rates hurt.

Take one activation h. After dropout it is either 0 (with probability p) or h / (1 − p) (with probability 1 − p). The average is h. The typical distance from that average, the standard deviation, works out to:

Example with h = 4 and p = 0.5: the value is either 0 or 8, always 4 away from the average 4. So the noise is exactly as large as the signal. With p = 0.9 the value is 0 nine times out of ten and 40 once: the noise is three times the signal.

What one activation of 4.0 looks like under different rates
Drop rate pValue if keptChance it is keptNoise (std)
0.14.4490%1.33
0.58.0050%4.00
0.940.0010%12.00

Two things calm this noise down. Width: the next layer adds up many activations, each with its own independent mask, so their noise partly cancels. A wide layer feels less noise per unit than a narrow one, which is one reason big dense layers tolerate p = 0.5 while small ones do not. Averaging: if we average n independent passes, the noise shrinks by √n. That is what Monte Carlo dropout relies on, and it is also what training does over many steps.

Practice: try it yourself

We test the "cheap ensemble" claim. We take our six activations, feed them into one neuron of the next layer, and compare two things: the single answer of the full network in evaluation mode, and the average answer of 20,000 random thinned networks.

practice_dropout_ensemble.py

import numpy as np
rng = np.random.default_rng(1)
h = np.array([1.0, 2.0, 3.0, 4.0, 5.0, 6.0])     # hidden activations
w = np.array([0.5, -0.4, 0.3, 0.2, -0.3, 0.2])   # weights of one next-layer neuron
b = 0.1
def pre_activation(h_in):
return h_in @ w + b                           # z = w.h + b
# Evaluation mode: no dropout, one fixed answer
z_full = pre_activation(h)
print(f"eval mode       z={z_full:.3f}  ReLU(z)={max(0.0, z_full):.3f}")
# Training mode: 20,000 random thinned networks for each drop rate
for p in (0.1, 0.5, 0.8):
keep = rng.random((20000, 6)) >= p            # one mask per row
thinned = h * keep / (1 - p)                  # inverted dropout
z = pre_activation(thinned)
out = np.maximum(0.0, z)                      # ReLU of each thinned network
print(f"p={p}  mean z={z.mean():.3f}  std z={z.std():.3f}  "
f"mean ReLU(z)={out.mean():.3f}")

Output:

eval mode       z=1.200  ReLU(z)=1.200
p=0.1  mean z=1.202  std z=0.824  mean ReLU(z)=1.231
p=0.5  mean z=1.212  std z=2.474  mean ReLU(z)=1.725
p=0.8  mean z=1.229  std z=4.899  mean ReLU(z)=2.515

Now change it:

  • Remove the / (1 - p) on line 18. Before running, predict the mean of z for p = 0.5. (Hint: the weighted sum without the bias is 1.1.)
  • Make the layer ten times wider with the same total signal: h = np.tile(h, 10), w = np.tile(w, 10) / 10, and change the mask shape to (20000, 60). Predict what happens to std z, and to the gap between the eval answer and mean ReLU(z).
  • Change the bias b from 0.1 to 10.0. Predict whether mean ReLU(z) now stays close to the eval answer, and explain why the ReLU no longer matters.

Pause and think: The mean of z stays near 1.2 at every drop rate, yet the mean of ReLU(z) climbs to 2.515 at p = 0.8. Where does the extra come from?

From the ReLU being non-linear. With a wide spread, many thinned networks give a negative z, which ReLU lifts to 0, while the large positive ones pass through untouched. Clipping only the low side raises the average. So "the full network equals the average of all thinned networks" is exact only for a linear layer. Through non-linearities it is an approximation, and the approximation gets worse as the noise grows.

Pause and think: At p = 0.5 the spread of z (2.474) is about twice the signal itself (1.2). How can a network learn anything from a signal that noisy?

Because training never relies on a single pass. Each step draws a new mask, and gradient descent adds up small updates over thousands of steps and examples, so the random part averages out while the consistent part (the true pattern) accumulates. The noise is the point: it stops the network from fitting details that only survive under one exact combination of neurons.

Key takeaways

  • Dropout randomly zeroes activations during training with probability p; a fresh mask is drawn each step.
  • Inverted dropout scales survivors by 1/(1 − p) so the expected signal matches inference, where dropout is off.
  • It reduces overfitting by breaking co-adaptation and acting like an ensemble of sub-networks.
  • Always switch to evaluation mode for validation and deployment.
  • Typical rates are 0.1–0.5; use less (or none) when underfitting or when training on huge data once.

Key terms

  • Dropout: A regularisation method that randomly sets a fraction of activations to zero during training.
  • Dropout rate (p): The probability that any given unit is dropped on a training step.
  • Overfitting: When a model fits the training data, including noise, so well that it performs worse on new data.
  • Inverted dropout: Dropout that scales kept activations by 1/(1 − p) during training so nothing changes at inference.
  • Co-adaptation: Neurons relying on specific other neurons to correct them, forming fragile features.
  • Monte Carlo dropout: Keeping dropout on at inference and averaging many runs to estimate uncertainty.

← 3.4 Cross-Entropy Loss: Scoring Probability Predictions · 3.6 Batch Norm vs Layer Norm: When to Use Each →