Lesson 3.2 · 25 min
Gradient Descent: Rolling Downhill to the Optimum
A model with billions of numbers has no map of the best settings, so how does it find them by only ever looking at the ground under its feet?
In short: Gradient descent is the algorithm that trains almost every neural network. It measures how wrong the model is with a loss function, computes the gradient (the direction in which the loss rises fastest), and nudges every parameter a small step the opposite way: w ← w − η · ∂L/∂w. Repeating this many times walks the parameters downhill to a low loss.
The big picture
Training a model means finding good values for its parameters: the weights and biases inside it. A small linear model has two. A large language model has billions. We cannot try every combination, and for neural networks there is no formula that solves for the best values directly. We need a method that improves the parameters a little at a time.
That method is gradient descent. Our running example is the house-price model from the previous lesson: price = w · size + b. We want to find the w and b that make predictions closest to real prices. Gradient descent will get there by repeatedly asking one question: if I nudge each parameter a tiny bit, does the error go up or down, and how fast?
Think of it like walking down a foggy hill You are on a hillside in thick fog and want to reach the valley floor. You cannot see the valley. But you can feel the slope under your feet. So you take a step in the steepest downhill direction, feel the slope again, and repeat. The height is the loss, your position is the parameters, the slope you feel is the gradient, and the length of each step is the learning rate.
What is a loss function?
Before we can go "downhill" we need a single number that says how bad the model is. That number is the loss (also called cost or error), and the function that computes it is the loss function. Lower is better. A loss of zero means perfect predictions on the training data.
For predicting numbers, the most common choice is mean squared error (MSE): take each prediction ŷᵢ, subtract the true value yᵢ, square the difference, and average over all n examples.
Example: if the model predicts 6 for a house worth 5 and predicts 9 for a house worth 9, the errors are 1 and 0, so MSE = (1² + 0²)/2 = 0.5. For classification we use a different loss, cross-entropy, which gets its own lesson. Gradient descent works with any loss that we can differentiate.
The key idea is that the loss is a function of the parameters. Fix the data, change w and b, and the loss changes. If we plot the loss for every possible w, we get a curve (or, with two parameters, a surface) called the loss landscape. Training means finding a low point on that landscape.
What gradient descent is, and the intuition
A derivative dL/dw tells us how fast the loss changes when w changes a tiny bit. If it is positive, increasing w increases the loss, so we should decrease w. If it is negative, increasing w lowers the loss, so we should increase w. Either way, we move against the sign of the derivative.
With many parameters, we take one derivative per parameter while holding the others fixed. These are called partial derivatives, written ∂L/∂w. Collected into a vector they form the gradient, written ∇L. The gradient points in the direction of steepest increase of the loss, so we step in the direction of −∇L. That is where the name comes from: we descend along the gradient.
The size of the gradient matters too. Far from the minimum the slope is usually steep, so the steps are big. Near the minimum the slope flattens, so steps shrink automatically. That is why gradient descent slows down gracefully as it approaches a good solution.
The math: the update rule
Everything in gradient descent comes down to one line, applied to every parameter at every step:
For our house model with MSE loss, calculus gives the two gradients. Using the chain rule on (w·xᵢ + b − yᵢ)²:
The gradient descent loop
- Initialise: Start the parameters somewhere, e.g. small random weights and zero biases.
- Predict: Run the model on the training data to get predictions ŷ (the forward pass).
- Measure loss: Compute L from the predictions and true values.
- Compute gradients: Find ∂L/∂θ for every parameter. For neural networks this is done by backpropagation (next lesson).
- Update: Apply θ ← θ − η·∂L/∂θ to every parameter at the same time.
- Repeat: Go back to Predict. Stop after a fixed number of steps or when the loss stops improving.
Step-by-step numeric example
Let us do it by hand on the simplest possible loss: L(w) = (w − 3)². Its minimum is obviously at w = 3, which lets us check our work. The derivative is dL/dw = 2(w − 3). We start at w = 0 with learning rate η = 0.1.
| Step | w before | Gradient 2(w − 3) | Update w − 0.1 × grad | Loss after |
|---|---|---|---|---|
| 1 | 0.000 | −6.000 | 0 + 0.600 = 0.600 | 5.760 |
| 2 | 0.600 | −4.800 | 0.600 + 0.480 = 1.080 | 3.686 |
| 3 | 1.080 | −3.840 | 1.080 + 0.384 = 1.464 | 2.359 |
The gradient is negative (the loss falls as w grows), so each update increases w. Each step closes 20% of the remaining gap to 3, because w − 3 gets multiplied by 1 − 2η = 0.8. The steps get smaller as the slope flattens. Starting from a loss of 9, the loss is multiplied by 0.64 each step.
Pause and think: Using the same loss and η = 0.1, suppose we started at w = 5 instead. What is the first gradient, and does w go up or down?
Gradient = 2(5 − 3) = +4. The update is w − 0.1 × 4 = 4.6, so w goes down, towards 3. A positive gradient always pushes the parameter down; a negative one pushes it up.
Gradient descent with multiple parameters
Real models have many parameters, but nothing new is needed. We compute one partial derivative per parameter and update all of them simultaneously, using gradients computed from the same old values. Updating w first and then using the new w to compute the gradient for b would be a different (and usually worse) algorithm.
Worked example with our house model: one house, size x = 2, price y = 7. Start at w = 0, b = 0, η = 0.05. Prediction ŷ = 0, error ŷ − y = −7. Gradients: ∂L/∂w = 2 · (−7) · 2 = −28 and ∂L/∂b = 2 · (−7) = −14. Updates: w = 0 − 0.05 · (−28) = 1.4 and b = 0 − 0.05 · (−14) = 0.7. New prediction: 1.4 · 2 + 0.7 = 3.5, already halfway to 7.
Notice that w moved more than b. Its gradient is multiplied by the input x = 2, so parameters attached to larger inputs get larger gradients. This is one reason we usually scale input features to similar ranges: otherwise one direction of the landscape is much steeper than another and the descent zig-zags.
The role of the learning rate
The learning rate η is a hyperparameter: a setting we choose rather than something the model learns. It is the single most important knob in gradient descent.
- Too small (e.g. 0.01 on our toy loss): every step is tiny, so training is slow and can stall on flat regions.
- Just right (0.1 here): steady, fast progress. On this particular loss, 0.5 even jumps to the exact minimum in one step.
- Too large (above 1.0 here): each step overshoots the valley by more than it started, so the loss grows and the parameters fly off to infinity. This is called divergence.
In practice, people try values spread over powers of ten (0.1, 0.01, 0.001...) and often use a learning-rate schedule that changes η over training: a short warm-up that increases it from near zero, followed by a slow decay. Optimizers such as Momentum, RMSProp and Adam build on plain gradient descent: momentum keeps a running average of past gradients to smooth the path, and Adam also scales each parameter's step by an estimate of its recent gradient size. Adam and its variant AdamW are the usual default for training Transformers.
Common mistake: blaming the model for a bad learning rate If the loss becomes NaN or shoots up in the first few hundred steps, the learning rate is the first suspect, not the architecture. If the loss falls painfully slowly from the very start, try a larger one. Always plot the loss curve.
Types of gradient descent
The formula for the gradient averages over training examples. The three classic types differ only in how many examples we use to compute each gradient before taking a step. One full pass over the training data is called an epoch.
Gradient descent in Python
This script reproduces the hand calculation, fits the house-price model with all three types of gradient descent for 50 epochs, and shows what different learning rates do.
gradient_descent.py
import numpy as np
# Part 1: one parameter. Loss L(w) = (w - 3)², gradient dL/dw = 2(w - 3)
w, lr = 0.0, 0.1
for step in range(1, 4):
grad = 2 * (w - 3)
w = w - lr * grad
print(f"step {step}: grad={grad:+.3f} w={w:.3f} loss={(w - 3) ** 2:.3f}")
# Part 2: two parameters (w, b) on house data, three flavours of gradient descent
rng = np.random.default_rng(0)
x = rng.uniform(0, 5, 200) # size in 100s of m²
y = 2 * x + 3 + rng.normal(0, 0.5, 200) # price in $100k, with noise
def fit(batch_size, lr=0.02, epochs=50):
w, b = 0.0, 0.0
for _ in range(epochs):
idx = rng.permutation(len(x)) # shuffle each epoch
for start in range(0, len(x), batch_size):
j = idx[start:start + batch_size]
err = (w * x[j] + b) - y[j]
w -= lr * 2 * np.mean(err * x[j]) # ∂L/∂w
b -= lr * 2 * np.mean(err) # ∂L/∂b
loss = np.mean((w * x + b - y) ** 2)
return w, b, loss
for name, bs in [("batch", 200), ("mini-batch", 32), ("stochastic", 1)]:
w, b, loss = fit(bs)
updates = 50 * int(np.ceil(200 / bs))
print(f"{name:11s} updates={updates:5d} w={w:.3f} b={b:.3f} MSE={loss:.3f}")
# Part 3: learning rate too large on L(w) = (w - 3)²
for lr in (0.01, 0.5, 1.1):
w = 0.0
for _ in range(10):
w -= lr * 2 * (w - 3)
print(f"lr={lr:<4} w after 10 steps = {w:.3f}")Output:
step 1: grad=-6.000 w=0.600 loss=5.760 step 2: grad=-4.800 w=1.080 loss=3.686 step 3: grad=-3.840 w=1.464 loss=2.359 batch updates= 50 w=2.396 b=1.568 MSE=0.765 mini-batch updates= 350 w=1.995 b=2.917 MSE=0.264 stochastic updates=10000 w=2.015 b=3.016 MSE=0.274 lr=0.01 w after 10 steps = 0.549 lr=0.5 w after 10 steps = 3.000 lr=1.1 w after 10 steps = -15.575
Look at the second block of output. With the same number of epochs, full-batch gradient descent is still far off (b = 1.57) because it took only 50 steps. Mini-batch reaches almost the noise floor with 350 steps. SGD gets there too, but with 10,000 tiny noisy steps that would be slow on real hardware. That trade-off is why mini-batch is the standard.
Pause and think: For L(w) = (w − 3)², each step multiplies the gap (w − 3) by (1 − 2η). Using that, why does η = 1.1 diverge?
1 − 2 × 1.1 = −1.2. The gap flips sign and grows by 20% every step, so w jumps back and forth across 3 with ever bigger swings. Any η above 1.0 makes |1 − 2η| > 1 on this loss, which means divergence.
Going one level deeper
Our one-parameter loss had a single slope. Real landscapes are steep in some directions and gentle in others. Let us see what that does with a two-parameter bowl: L(a, b) = a² + 10·b². The gradients are ∂L/∂a = 2a and ∂L/∂b = 20b. Direction b is ten times steeper than direction a.
Each update multiplies a by (1 − 2η) and b by (1 − 20η). Both factors must stay between −1 and 1, or that parameter grows. For a any η below 1.0 is safe. For b the limit is η < 0.1. The steepest direction sets the speed limit for everyone. Here is η = 0.09, starting from a = 5, b = 1:
| Step | a (× 0.82 each step) | b (× −0.8 each step) | Loss |
|---|---|---|---|
| 0 | 5.000 | 1.000 | 35.00 |
| 1 | 4.100 | −0.800 | 23.21 |
| 2 | 3.362 | 0.640 | 15.40 |
| 3 | 2.757 | −0.512 | 10.22 |
| 4 | 2.261 | 0.410 | 6.79 |
Read the two columns. b jumps from one side of the valley to the other on every step: that is the zig-zag. a moves in a steady line, but slowly, because we cannot raise η without breaking b.
Momentum is the standard fix. We keep a velocity v for each parameter: v ← β·v + gradient, then θ ← θ − η·v, with β around 0.8 to 0.9. In the zig-zag direction, gradients keep changing sign, so they partly cancel inside v. In the steady direction they all point the same way, so they add up and the step grows.
Practice: try it yourself
We now run gradient descent on the steep-and-gentle bowl ourselves. We count how many steps it needs for three learning rates, then switch on momentum and compare.
practice_momentum_bowl.py
# A bowl with one gentle and one steep direction: L(a, b) = a*a + 10*b*b
def loss(a, b):
return a * a + 10 * b * b
def run(lr, beta, max_steps=500):
a, b = 5.0, 1.0 # starting point
va, vb = 0.0, 0.0 # velocity: running sum of past gradients
for step in range(1, max_steps + 1):
ga, gb = 2 * a, 20 * b # the two partial derivatives
va = beta * va + ga # beta = 0 gives plain descent
vb = beta * vb + gb
a -= lr * va
b -= lr * vb
if loss(a, b) > 1e6:
return f"diverged at step {step}"
if loss(a, b) < 1e-4:
return f"reached loss < 0.0001 in {step} steps"
return f"still at loss {loss(a, b):.4f} after {max_steps} steps"
for lr, beta in [(0.02, 0.0), (0.09, 0.0), (0.11, 0.0), (0.02, 0.8)]:
print(f"lr={lr:<5} momentum={beta:<4} -> {run(lr, beta)}")
# Watch the steep direction b zig-zag when lr is close to its limit
b, trail = 1.0, []
for _ in range(5):
b -= 0.09 * 20 * b
trail.append(round(b, 3))
print("b with lr=0.09:", trail)Output:
lr=0.02 momentum=0.0 -> reached loss < 0.0001 in 153 steps lr=0.09 momentum=0.0 -> reached loss < 0.0001 in 32 steps lr=0.11 momentum=0.0 -> diverged at step 32 lr=0.02 momentum=0.8 -> reached loss < 0.0001 in 57 steps b with lr=0.09: [-0.8, 0.64, -0.512, 0.41, -0.328]
Now change it:
- Add the pair
(0.1, 0.0)to the list of runs. The factor forbbecomes exactly −1. Before running, predict which of the three messages it prints and what loss it reports. - Make the bowl steeper: change
10 b bto100 b band20 bto200 b. Work out the new speed limit forηfirst, then predict which of the four runs still converge. - Raise momentum in the last run from 0.8 to 0.99. Predict whether it needs fewer or more than 57 steps, then explain the result.
Pause and think: With lr = 0.11, direction a is perfectly stable (its factor is 1 − 0.22 = 0.78). Why does the whole run still diverge?
Because the loss adds both directions. The factor for b is 1 − 20·0.11 = −1.2, so b grows by 20% per step while flipping sign. However well a behaves, 10·b² grows without limit. One unstable direction is enough to blow up the loss, so the steepest direction decides the largest safe learning rate.
Pause and think: Suppose we rescale the problem so the loss is a² + b² instead of a² + 10·b². What learning rate reaches the minimum in one step, and what does that tell us about feature scaling?
Both gradients are now 2a and 2b, so both factors are (1 − 2η). With η = 0.5 both become 0 and we land on the minimum in a single step. When all directions have the same steepness, one learning rate is ideal for all of them. That is why we scale features: it makes the bowl rounder, so we no longer have to pick a rate that is too slow for most directions just to keep one of them stable.
Putting it all together
Where you meet this in the real world Every time a framework such as PyTorch runs loss.backward() followed by optimizer.step(), it is doing one iteration of this loop: backward computes the gradients, step applies an update rule (plain SGD, or a smarter variant such as Adam). Training a large language model is this same loop repeated over a huge number of mini-batches.
- Local minima and saddle points: on non-convex landscapes (all neural networks), gradient descent can stop at a point that is not the global best, or slow down on flat saddle regions. In practice, for large networks, the noise of mini-batches and good optimizers usually find solutions that work well.
- Gradient descent needs gradients. It does not work directly on things you cannot differentiate, such as accuracy counts or discrete choices; we optimise a smooth stand-in loss instead.
- Scale your features so no single parameter has a much steeper slope than the others.
Key takeaways
- A loss function turns "how wrong is the model" into one number that depends on the parameters.
- The gradient points uphill; gradient descent steps the opposite way: θ ← θ − η·∂L/∂θ.
- All parameters are updated together, each by its own partial derivative.
- The learning rate decides step size: too small is slow, too large overshoots or diverges.
- Mini-batch gradient descent is the standard: many fairly accurate, GPU-friendly updates per epoch.
Key terms
- Loss function: A function that turns the model's predictions and the true answers into one number measuring error; lower is better.
- Gradient: The vector of partial derivatives of the loss with respect to every parameter; it points in the direction of steepest increase.
- Learning rate (η): A hyperparameter setting how big a step gradient descent takes each update.
- Epoch: One full pass through the training dataset.
- Mini-batch: A small random subset of training examples used to compute one gradient update.
- Divergence: When updates overshoot so badly that the loss grows instead of shrinking, often ending in NaN.
← 3.1 Neural Network Bias: What It Is and Why It Matters · 3.3 Backpropagation: How Neural Networks Learn from Mistakes →