Lesson 3.3 · 27 min
Backpropagation: How Neural Networks Learn from Mistakes
Gradient descent needs the slope of the loss for every single weight, so how does a network with millions of weights get all of them from one backward sweep?
In short: Backpropagation is the algorithm that computes the gradient of the loss with respect to every weight in a neural network. It runs a forward pass to get the prediction and loss, then walks backwards layer by layer, using the chain rule to pass an error signal from the output towards the input. Gradient descent then uses those gradients to update the weights.
What is backpropagation?
In the previous lesson, gradient descent updated each parameter with θ ← θ − η · ∂L/∂θ. For a two-parameter line we could write the gradients by hand. A neural network, though, is a long chain of layers, and a weight in the first layer affects the loss only indirectly, through every layer after it. We need a systematic way to get ∂L/∂w for every weight.
Backpropagation (short for "backward propagation of errors") is that way. It is not a learning rule by itself. It is an efficient method for computing gradients. The learning happens when gradient descent (or Adam, etc.) uses those gradients. The two together are what people usually mean by "training a neural network".
Think of it like tracing blame in a relay team A relay team finishes 4 seconds late. The coach starts at the finish line and works backwards: the last runner was 1 second slow; they were also handed the baton late, so some blame passes to the third runner, and so on. Each runner's share of the blame depends on how much their part affected the next runner. Backpropagation assigns blame for the loss to every weight in the same backward way.
The key reason it works efficiently: the gradients of early layers reuse the gradients already computed for later layers. One backward sweep costs roughly as much as one or two forward passes, no matter how many weights there are. Without this reuse, training today's large networks would be hopeless.
The chain rule of calculus
Backpropagation is the chain rule applied over and over. The chain rule says how to differentiate a function of a function. If y depends on u, and u depends on x, then a small change in x changes u, which changes y. The rates multiply:
Small example: y = (3x + 1)². Let u = 3x + 1, so y = u². Then dy/du = 2u and du/dx = 3, which gives dy/dx = 2u · 3 = 6(3x + 1). At x = 1: u = 4, so dy/dx = 24. Check: nudging x from 1 to 1.001 changes y from 16 to about 16.024, a change of 0.024 for 0.001, i.e. a rate of 24.
A neural network is just a longer chain: input → weighted sum → activation → weighted sum → activation → ... → loss. So ∂L/∂w for any weight is a product of local derivatives along the path from that weight to the loss. When a value feeds into several later values, we add the contributions from each path (the multivariable chain rule).
Pause and think: If L = u² and u = 5w, what is dL/dw when w = 2?
u = 10, dL/du = 2u = 20, du/dw = 5, so dL/dw = 20 × 5 = 100.
Forward pass
We will use a deliberately tiny network so every number fits on screen: one input, one hidden neuron with a sigmoid activation, and one linear output neuron. The sigmoid is σ(z) = 1 / (1 + e⁻ᶻ), an S-shaped function that squashes any number into the range 0 to 1.
The forward pass computes these from left to right. Our numbers: x = 1, target y = 1, and starting parameters w₁ = 0.5, b₁ = 0, w₂ = 1.0, b₂ = 0.
z₁ = 0.5 × 1 + 0 = 0.5h = σ(0.5) = 0.6225ŷ = 1.0 × 0.6225 + 0 = 0.6225
Crucially, the forward pass stores the intermediate values z₁, h and ŷ. The backward pass will need them. This is why training uses more memory than just running a model: all those activations must be kept until the gradients are computed.
Loss calculation
The loss compares the prediction with the truth. With squared error: L = (0.6225 − 1)² = (−0.3775)² = 0.1425. In a real network we average this over a mini-batch; for classification we would use cross-entropy instead (next lesson). Backpropagation works the same way for any differentiable loss; only the very first derivative ∂L/∂ŷ changes.
The derivative of the loss with respect to the prediction is the starting point of the backward pass: ∂L/∂ŷ = 2(ŷ − y) = 2 × (−0.3775) = −0.7551. It is negative, which tells us that increasing ŷ would lower the loss. That makes sense: the prediction 0.62 is too low compared with the target 1.
Backward pass (backpropagation)
Now we walk from the loss back to the inputs. At each node we multiply the gradient arriving from above by the node's local derivative: how its output changes with its own input. We call the gradient arriving at a value its error signal, often written δ (delta).
The backward pass for our tiny network
- Start at the loss:
∂L/∂ŷ = 2(ŷ − y). This is the error signal at the output. - Output layer parameters: Since
ŷ = w₂·h + b₂:∂ŷ/∂w₂ = hand∂ŷ/∂b₂ = 1. So∂L/∂w₂ = ∂L/∂ŷ · hand∂L/∂b₂ = ∂L/∂ŷ. - Pass the error to the hidden layer:
∂ŷ/∂h = w₂, so∂L/∂h = ∂L/∂ŷ · w₂. The weight that carried the signal forward now carries the blame backward. - Through the activation: The sigmoid has the handy derivative
σ'(z) = σ(z)(1 − σ(z)) = h(1 − h). So∂L/∂z₁ = ∂L/∂h · h(1 − h). - Hidden layer parameters:
∂z₁/∂w₁ = xand∂z₁/∂b₁ = 1. So∂L/∂w₁ = ∂L/∂z₁ · xand∂L/∂b₁ = ∂L/∂z₁.
Notice the pattern. Every gradient for a weight is (error signal at the layer's output) × (the input that weight multiplied). And the error signal for the previous layer is (error signal) × (weight) × (activation derivative). In matrix form for a whole layer, that becomes ∂L/∂W = inputᵀ · δ and δ_prev = (δ · Wᵀ) ⊙ f'(z_prev), where ⊙ means element-wise multiplication.
Step-by-step numeric example
Plugging in our numbers (x = 1, y = 1, h = ŷ = 0.6225, w₂ = 1.0):
| Quantity | Formula | Value |
|---|---|---|
| ∂L/∂ŷ | 2(ŷ − y) = 2(0.6225 − 1) | −0.7551 |
| ∂L/∂w₂ | ∂L/∂ŷ · h = −0.7551 × 0.6225 | −0.4700 |
| ∂L/∂b₂ | ∂L/∂ŷ · 1 | −0.7551 |
| ∂L/∂h | ∂L/∂ŷ · w₂ = −0.7551 × 1.0 | −0.7551 |
| σ'(z₁) | h(1 − h) = 0.6225 × 0.3775 | 0.2350 |
| ∂L/∂z₁ | ∂L/∂h · σ'(z₁) = −0.7551 × 0.2350 | −0.1774 |
| ∂L/∂w₁ | ∂L/∂z₁ · x = −0.1774 × 1 | −0.1774 |
| ∂L/∂b₁ | ∂L/∂z₁ · 1 | −0.1774 |
All gradients are negative, so gradient descent will increase all four parameters, pushing the prediction up towards 1. Also notice that the hidden-layer gradients (−0.18) are much smaller than the output-layer ones. The sigmoid's derivative is at most 0.25, so each sigmoid layer shrinks the error signal by at least 4×. Stack many of them and early layers get almost no gradient: the vanishing gradient problem, one big reason modern networks prefer ReLU-style activations, normalisation and residual connections.
Pause and think: If w₂ had been 0 instead of 1.0, what would ∂L/∂w₁ be? Why?
Zero. ∂L/∂h = ∂L/∂ŷ · w₂ = 0, so no error signal reaches the hidden layer. If the output ignores h, changing w₁ cannot affect the loss. This is also why we never initialise all weights to zero.
Weight update using gradient descent
With the gradients in hand, we apply the gradient descent rule to every parameter at once. Using a learning rate η = 0.5:
w₁ = 0.5 − 0.5 × (−0.1774) = 0.5887b₁ = 0 − 0.5 × (−0.1774) = 0.0887w₂ = 1.0 − 0.5 × (−0.4700) = 1.2350b₂ = 0 − 0.5 × (−0.7551) = 0.3775
Running the forward pass again with these new values gives ŷ ≈ 1.1966 and a loss of about 0.0386, down from 0.1425. The prediction has actually overshot past 1 (this learning rate is large for one example), but the loss has still dropped by almost 4×. The next step would pull it back. Training is simply this cycle, forward, loss, backward, update, repeated thousands or millions of times.
Backpropagation in Python
Now a real (still small) network: 2 inputs, 4 hidden tanh neurons, 1 sigmoid output, trained on XOR: output 1 when exactly one input is 1. XOR is the classic problem a single neuron cannot solve, so the hidden layer and therefore backprop are essential. We also do a gradient check: compare one backprop gradient with a numerical estimate from nudging the weight.
backprop_xor.py
import numpy as np
rng = np.random.default_rng(42)
X = np.array([[0., 0.], [0., 1.], [1., 0.], [1., 1.]]) # 4 examples, 2 features
y = np.array([[0.], [1.], [1.], [0.]]) # XOR targets
W1, b1 = rng.normal(0, 1, (2, 4)), np.zeros(4) # input -> 4 hidden
W2, b2 = rng.normal(0, 1, (4, 1)), np.zeros(1) # hidden -> 1 output
sigmoid = lambda z: 1 / (1 + np.exp(-z))
def forward(W1, b1, W2, b2):
h = np.tanh(X @ W1 + b1) # hidden layer
y_hat = sigmoid(h @ W2 + b2) # output probability
return h, y_hat, np.mean((y_hat - y) ** 2)
lr = 2.0
for step in range(3001):
h, y_hat, loss = forward(W1, b1, W2, b2)
# ---- backward pass: chain rule, from the loss back to each weight ----
d_yhat = 2 * (y_hat - y) / len(X) # dL/dŷ
d_z2 = d_yhat * y_hat * (1 - y_hat) # through sigmoid: σ' = σ(1-σ)
dW2, db2 = h.T @ d_z2, d_z2.sum(0)
d_h = d_z2 @ W2.T # send the error back to hidden
d_z1 = d_h * (1 - h ** 2) # through tanh: tanh' = 1 - tanh²
dW1, db1 = X.T @ d_z1, d_z1.sum(0)
if step == 0: # check one gradient numerically: (L(w+ε) - L(w-ε)) / 2ε
eps, Wp, Wm = 1e-5, W1.copy(), W1.copy()
Wp[1, 2] += eps; Wm[1, 2] -= eps
num = (forward(Wp, b1, W2, b2)[2] - forward(Wm, b1, W2, b2)[2]) / (2 * eps)
print(f"backprop dL/dW1[1,2]={dW1[1, 2]:.6f} numerical={num:.6f}")
W1 -= lr * dW1; b1 -= lr * db1; W2 -= lr * dW2; b2 -= lr * db2
if step % 1000 == 0:
print(f"step {step:4d} loss={loss:.4f}")
print("predictions:", forward(W1, b1, W2, b2)[1].ravel().round(3))Output:
backprop dL/dW1[1,2]=-0.012556 numerical=-0.012556 step 0 loss=0.2874 step 1000 loss=0.0004 step 2000 loss=0.0002 step 3000 loss=0.0001 predictions: [0.005 0.988 0.99 0.013]
Gradient checking When you write backprop by hand, always compare a few gradients against the numerical estimate (L(w + ε) − L(w − ε)) / 2ε. It is far too slow for training (two forward passes per weight) but perfect for catching bugs. Frameworks like PyTorch do the backward pass automatically, so you rarely need this outside custom code.
Alternatives, real-world use and pitfalls
Real-world use Every model you will meet in this course, from image classifiers to large language models, was trained with backpropagation. In PyTorch, loss.backward() runs exactly the backward pass we did by hand, over every operation in the network, and stores each parameter's gradient in its .grad attribute.
Common mistakes Initialising all weights to the same value (every hidden neuron then gets identical gradients and they never become different). Forgetting that sigmoid/tanh derivatives shrink the signal, causing vanishing gradients in deep stacks. Letting gradients explode in very deep or recurrent networks (fixed with gradient clipping, normalisation and careful initialisation). And confusing backprop (computing gradients) with the optimiser (using them).
Worked example, step by step
Our tiny network was a single chain: each value fed exactly one later value. We said that when a value feeds several later values we add the contributions. Let us work one such case by hand, because this is where hand-written backprop most often goes wrong.
Take one weight w = 3 and one input x = 2. The weight is used twice: u = w·x and v = w². The loss is their product, L = u·v. So w reaches the loss along two paths, one through u and one through v.
Backprop when a weight is used twice
- Forward pass:
u = 3·2 = 6,v = 3² = 9,L = 6·9 = 54. We storeuandv. - Start at the loss: For a product, each factor's gradient is the other factor:
∂L/∂u = v = 9and∂L/∂v = u = 6. - Path through u:
∂u/∂w = x = 2, so this path gives9 · 2 = 18. - Path through v:
∂v/∂w = 2w = 6, so this path gives6 · 6 = 36. - Add the paths:
∂L/∂w = 18 + 36 = 54. Check with plain calculus:L = w³·x, sodL/dw = 3w²·x = 3·9·2 = 54. They agree.
| What we do with the two paths | Result | Correct? |
|---|---|---|
| Add them | 18 + 36 = 54 | Yes |
| Keep only the last one computed | 36 | No: the path through u is lost |
| Multiply them | 18 × 36 = 648 | No: we multiply along a path, never across paths |
The rule to remember: multiply along a path, add across paths. This matters far beyond toy examples. A recurrent network uses the same weights at every time step, so each weight's gradient is a sum over all the steps. A residual connection sends a value down two routes that meet again. In all these cases the gradients must be accumulated with +=, not overwritten with =. A bug of this kind is silent: the code runs, the loss may even fall a little, and only a gradient check reveals it.
Practice: try it yourself
We will measure the vanishing gradient ourselves. We build a chain of one-neuron layers, run backprop by hand, and print the gradient that reaches the first weight as the chain gets deeper. We do it once with sigmoid and once with ReLU.
practice_gradient_depth.py
import math
def sigmoid(z):
return 1 / (1 + math.exp(-z))
ACTS = {
"sigmoid": (sigmoid, lambda z: sigmoid(z) * (1 - sigmoid(z))),
"relu": (lambda z: max(0.0, z), lambda z: 1.0 if z > 0 else 0.0),
}
def first_weight_gradient(depth, name, x=1.0, w=1.0):
act, d_act = ACTS[name]
# Forward pass: h_k = act(w * h_(k-1)), and the loss is the last h.
zs, hs = [], [x]
for _ in range(depth):
zs.append(w * hs[-1])
hs.append(act(zs[-1]))
# Backward pass: start with dL/dh_last = 1 and walk towards the input.
delta = 1.0
for k in reversed(range(depth)):
delta *= d_act(zs[k]) # through the activation
if k == 0:
return delta * hs[0] # dL/dw of the first layer
delta *= w # through the weight to the layer below
print("depth sigmoid chain relu chain")
for depth in (1, 2, 4, 8, 16):
gs = first_weight_gradient(depth, "sigmoid")
gr = first_weight_gradient(depth, "relu")
print(f"{depth:5d} {gs:13.2e} {gr:10.2f}")Output:
depth sigmoid chain relu chain 1 1.97e-01 1.00 2 4.31e-02 1.00 4 2.16e-03 1.00 8 5.52e-06 1.00 16 3.58e-11 1.00
Now change it:
- Call the function with
w=4.0for both chains. Larger weights multiply the signal on the way back. Predict whether that rescues the sigmoid chain, and what it does to the ReLU chain at depth 16. - Call it with
x=-1.0. Predict the ReLU gradient at every depth before you run it, and name the problem this shows. - Add a third activation, tanh, whose derivative is
1 - math.tanh(z) ** 2. Predict whether its gradients at depth 16 land closer to the sigmoid column or to the ReLU column.
Pause and think: At depth 2 the sigmoid gradient is 0.0431, yet the lesson said each sigmoid layer can pass on up to 0.25 of the signal, which would allow 0.25 × 0.25 = 0.0625. Why is the real number smaller?
The derivative σ'(z) equals 0.25 only at z = 0. Here the pre-activations are 1.0 and about 0.73, where the derivative is about 0.197 and 0.219. Their product is 0.0431. So 0.25 is a best case; real layers usually pass on less, and the further z is from zero, the less gets through.
Pause and think: The ReLU column is exactly 1.00 at every depth. Does that mean ReLU networks can never have gradient problems?
No. It is 1.00 here because every pre-activation is positive (derivative 1) and every weight is 1. If any layer's pre-activation is negative, its derivative is 0 and the whole product becomes 0: a dead path. And if the weights are larger than 1, the product grows with depth instead: an exploding gradient. ReLU removes the shrinking caused by the activation, not the effect of the weights.
Key takeaways
- Backpropagation computes gradients; the optimiser (gradient descent, Adam) uses them to update weights.
- It is the chain rule applied backwards: multiply local derivatives along the path from the loss to each weight.
- Forward pass computes and stores activations; backward pass reuses them to get every gradient in one sweep.
- For a weight: gradient = error signal at its output × the input it multiplied.
- Small activation derivatives shrink gradients layer by layer (vanishing gradients); always sanity-check hand-written gradients numerically.
Key terms
- Backpropagation: An algorithm that computes the gradient of the loss for every parameter by applying the chain rule backwards through the network.
- Chain rule: The calculus rule that the derivative of a composed function is the product of the derivatives of its parts.
- Forward pass: Running inputs through the network to compute and store activations, the prediction and the loss.
- Error signal (δ): The gradient of the loss with respect to a node's pre-activation, passed backwards layer by layer.
- Vanishing gradient: When gradients become tiny in early layers because many small derivatives are multiplied together.
- Gradient check: Comparing backprop gradients against finite-difference estimates to detect bugs.
← 3.2 Gradient Descent: Rolling Downhill to the Optimum · 3.4 Cross-Entropy Loss: Scoring Probability Predictions →