Modern AI Engineering

Lesson 2.6 · 21 min

L1 vs L2 Loss: Choosing Your Error Penalty

One late delivery out of six can make up 99.7% of a model's loss, or only 89% of it, depending on a single choice: do we take the absolute value of the error, or square it?

In short: A loss function turns prediction errors into one number that training tries to minimise. L1 loss (mean absolute error) averages |error| and treats every unit of error equally, so it is robust to outliers and aims at the median. L2 loss (mean squared error) averages error², punishing big errors much more, so it is smooth and easy to optimise but sensitive to outliers and aims at the mean. Pick based on how much big errors should matter and how much you trust your data.

First, what is a loss function?

When a model makes a prediction, it is usually a bit wrong. The error (also called the residual) for one example is the difference between the true value and the prediction: e = y − ŷ. A loss function turns all those individual errors into a single number that says how bad the model is overall. Training (Lesson 2.1) is the search for parameters that make this number as small as possible.

The choice of loss is not a detail. It is how we tell the model what kind of mistakes we care about. Two models trained on the same data with different losses can end up making noticeably different predictions.

Two strict teachers Imagine two teachers grading how late students are. The L1 teacher gives one penalty point per minute late: 2 minutes late = 2 points, 20 minutes late = 20 points. The L2 teacher squares the minutes: 2 minutes = 4 points, 20 minutes = 400 points. Under the L2 teacher, one very late student dominates the whole class's penalty.

Running example: we predict food delivery times in minutes. Five deliveries arrived roughly on time; one hit a road closure and took 90 minutes instead of the predicted 34.

L1 loss: mean absolute error

L1 loss uses the absolute value of each error. Averaged over n examples, it is called Mean Absolute Error (MAE). The name "L1" comes from the L1 norm, the sum of absolute values of a vector.

Computing MAE for our deliveries

  1. List the errors: Actual minus predicted: −1, 2, −1, 2, −1 and 56 (the road closure).
  2. Take absolute values: 1, 2, 1, 2, 1, 56. Over-predicting and under-predicting count the same.
  3. Sum them: 1 + 2 + 1 + 2 + 1 + 56 = 63.
  4. Average: 63 / 6 = 10.5 minutes. Without the outlier, MAE is 7 / 5 = 1.4 minutes.

Key properties of L1:

  • Linear penalty. An error of 10 costs exactly ten times an error of 1. Big errors matter, but not disproportionately.
  • Robust to outliers. A single wild value cannot dominate the loss as much as under L2.
  • Easy to interpret. MAE is in the same units as the target: "on average we are 10.5 minutes off".
  • Constant gradient. The slope of |e| is +1 or −1 (and undefined exactly at 0). Every error pushes the model with the same strength regardless of size, so near the optimum it does not naturally slow down, and optimisers may need a decaying learning rate.
  • Aims at the median. If a model could only output one constant, the value that minimises L1 is the median of the targets.

L2 loss: mean squared error

L2 loss squares each error. Averaged, it is the Mean Squared Error (MSE), the default loss for regression and the one we used in Lessons 2.1 and 2.3. Its square root, RMSE, brings it back to the original units.

  • Quadratic penalty. An error of 10 costs 100 times an error of 1. The model is pushed hard to avoid big misses.
  • Sensitive to outliers. Our single 56-minute error makes up 3136 of the 3147 total, 99.7% of the loss. The model will bend itself to reduce that one error, even if that makes typical predictions worse.
  • Smooth gradient. The slope of e² is 2e: large for large errors and shrinking to zero as the error shrinks. This makes gradient descent stable and naturally slows down near the optimum.
  • Aims at the mean. The single constant that minimises L2 is the mean of the targets.
  • Statistical meaning. Minimising MSE is equivalent to maximum likelihood when the noise is Gaussian (normally distributed), one reason it is the default.

Pause and think: An error of 0.5 contributes 0.5 to L1 loss. What does it contribute to L2 loss, and what does this tell us about small errors?

0.5² = 0.25, less than L1. Squaring shrinks errors below 1, so L2 cares relatively little about small errors and a lot about large ones. (Note this depends on units: an error of 0.5 hours is 30 minutes.)

Code: the same errors under both losses

l1_vs_l2.py

import numpy as np
# Delivery times (minutes); the last one hit a road closure
actual = np.array([30, 32, 29, 35, 31, 90], dtype=float)
predicted = np.array([31, 30, 30, 33, 32, 34], dtype=float)
err = actual - predicted
def l1(e): return np.abs(e).mean()      # mean absolute error (MAE)
def l2(e): return (e ** 2).mean()       # mean squared error (MSE)
print("errors:", err.tolist())
print(f"without the outlier: L1 = {l1(err[:5]):.2f}   L2 = {l2(err[:5]):.2f}")
print(f"with the outlier:    L1 = {l1(err):.2f}   L2 = {l2(err):.2f}")
print(f"outlier's share of total loss: L1 {abs(err[5]) / np.abs(err).sum():.1%}, "
f"L2 {err[5]**2 / (err**2).sum():.1%}")
# If a model could only predict ONE constant c, which c minimises each loss?
grid = np.linspace(25, 95, 7001)              # try c = 25.00, 25.01, ..., 95.00
best_l1 = grid[np.argmin([l1(actual - c) for c in grid])]
best_l2 = grid[np.argmin([l2(actual - c) for c in grid])]
print(f"best constant under L1: {best_l1:.2f}  (median = {np.median(actual):.2f})")
print(f"best constant under L2: {best_l2:.2f}  (mean   = {actual.mean():.2f})")
# Gradients: how hard each loss pushes on an error of 1 vs 56
for e in (1.0, 56.0):
print(f"error {e:>4}: L1 gradient = {np.sign(e):.0f}, L2 gradient = {2 * e:.0f}")

Output:

errors: [-1.0, 2.0, -1.0, 2.0, -1.0, 56.0]
without the outlier: L1 = 1.40   L2 = 2.20
with the outlier:    L1 = 10.50   L2 = 524.50
outlier's share of total loss: L1 88.9%, L2 99.7%
best constant under L1: 31.00  (median = 31.50)
best constant under L2: 41.17  (mean   = 41.17)
error  1.0: L1 gradient = 1, L2 gradient = 2
error 56.0: L1 gradient = 1, L2 gradient = 112

Pause and think: If the predictions above were used to set a promised delivery time for customers, which loss gives a more useful "typical" time: L1 (≈31) or L2 (≈41)?

For "typical" customer experience, L1's ≈31 minutes is more representative: five of six deliveries took 29–35 minutes. L2's 41 minutes is pulled up by a single road closure. (If late deliveries are very costly to the business, though, we might deliberately prefer a loss that weighs them heavily.)

How to decide between L1 and L2 loss

A practical set of questions:

  • Are there outliers, and are they real or errors? If they are data-entry mistakes or rare freak events we do not want to model, prefer L1 (or clean the data). If they are real and important, L2 makes sure the model takes them seriously.
  • How costly is a big error compared with several small ones? If one 50-minute miss is much worse than ten 5-minute misses, L2 matches that preference.
  • What will we report to people? MAE ("10 minutes off on average") is easier to explain than MSE. Many teams train with one loss and report several metrics.
  • Does training need to be smooth? L2's gradient is friendlier for gradient descent, especially near the optimum.

Where these losses show up

SettingTypical lossWhy
Linear regression on clean dataL2 (MSE)Closed-form solution, Gaussian noise assumption
Demand or delivery-time forecasting with occasional spikesL1 (MAE) or HuberSpikes should not distort typical predictions
Bounding-box regression in object detectionSmooth L1 / Huber, or L1 in some modelsRobust to badly wrong early predictions
Diffusion models predicting noise (Module 16)Usually L2 (MSE) on the noiseSmooth objective with a clear probabilistic meaning
Image-to-image models where blur is a problemOften L1 on pixelsL2 tends to average possibilities into blurry images
ClassificationNeither: cross-entropy (Lesson 3.4)Targets are classes, not numbers

Do not confuse with L1/L2 regularisation L1 and L2 losses measure prediction errors. L1 and L2 regularisation (next lesson) add a penalty on the model's weights to prevent overfitting. Both use the same two norms, |·| and (·)², which is why the names match, but they serve different purposes and are often used together, for example MSE loss + L2 weight penalty.

Common mistakes and limits

Comparing MSE across different units MSE is in squared units: an MSE of 100 minutes² is an RMSE of 10 minutes. Comparing raw MSE values between datasets or units is meaningless. Report RMSE or MAE for humans.

  • Using L2 on dirty data without looking. A few mislabelled rows can drag the whole model. Plot the errors first.
  • Dropping outliers automatically. Outliers may be the most important cases (fraud, failures). Decide on purpose whether to model them.
  • Assuming L1 has no downsides. Its gradient does not shrink near the optimum, and the median ignores how far the extreme values are, which is wrong if extremes are what matter.
  • Using a regression loss for classification. For class probabilities, use log loss or cross-entropy (Lesson 2.3 and Module 3).

Worked example, step by step

So far we scored one model with two losses. A sharper test is to score two models and ask which one wins. The answer can flip depending on the loss. Here are two made-up delivery-time models, each tested on the same four deliveries.

Illustrative errors in minutes on four deliveries
ModelError 1Error 2Error 3Error 4Character
Model A3333Always a little off
Model B00010Usually perfect, once badly off

Scoring both models by hand

  1. MAE of model A: (3 + 3 + 3 + 3) / 4 = 3.0 minutes.
  2. MAE of model B: (0 + 0 + 0 + 10) / 4 = 2.5 minutes. Under L1, B wins.
  3. MSE of model A: (9 + 9 + 9 + 9) / 4 = 9.0.
  4. MSE of model B: (0 + 0 + 0 + 100) / 4 = 25.0. Under L2, A wins, and by a wide margin.
  5. Back to minutes: RMSE is √9 = 3.0 for A and √25 = 5.0 for B. For A, RMSE equals MAE because all its errors are the same size. For B, RMSE is twice its MAE.

Neither metric is wrong. They answer different questions. MAE asks "how far off are we in total?" and treats the 10-minute miss as worth ten 1-minute misses. MSE asks "how bad are our worst misses?" and treats it as worth a hundred.

A quick diagnostic: compare RMSE with MAE RMSE can never be smaller than MAE. When the two are close, the errors are all of similar size. When RMSE is much larger than MAE, a few big errors are hiding behind a decent average. That gap is a cheap signal to go and look at the worst cases before trusting either number.

Practice: try it yourself

This time we do not just measure with each loss, we train with it. We fit a one-number model, minutes = w × km, by gradient descent three times: with the L2 gradient, the L1 gradient and the Huber gradient. Five deliveries follow the rule of 5 minutes per km exactly. One got stuck in traffic.

practice_train_with_losses.py

import numpy as np
# Delivery distance (km) and time (minutes). The rule is 5 minutes per km,
# but the last delivery got stuck in traffic: 60 minutes instead of 30.
km = np.array([1, 2, 3, 4, 5, 6], dtype=float)
minutes = np.array([5, 10, 15, 20, 25, 60], dtype=float)
def grad_l2(e):                 # slope of e^2: grows with the error
return 2 * e
def grad_l1(e):                 # slope of |e|: always +1 or -1
return np.sign(e)
def grad_huber(e, delta=2.0):   # like L2 for small errors, like L1 for big ones
return np.clip(e, -delta, delta)
def train(grad, steps=4000):
"""Fit minutes = w * km by gradient descent with a shrinking step size."""
w = 0.0
for t in range(steps):
e = w * km - minutes                 # prediction minus truth
lr = 0.02 / (1 + 0.01 * t)           # step size decays over time
w -= lr * (grad(e) * km).mean()      # chain rule: d(error)/dw = km
return w
for name, grad in (("L2", grad_l2), ("L1", grad_l1), ("Huber", grad_huber)):
w = train(grad)
e = np.round(w * km - minutes, 2) + 0.0   # final errors, tidy for printing
print(f"{name:5s}  w = {w:.2f} min/km   error on normal 5 km trip = {e[4]:+5.2f}"
f"   error on traffic trip = {e[5]:+6.2f}")

Output:

L2     w = 6.98 min/km   error on normal 5 km trip = +9.89   error on traffic trip = -18.13
L1     w = 5.00 min/km   error on normal 5 km trip = +0.00   error on traffic trip = -30.00
Huber  w = 5.22 min/km   error on normal 5 km trip = +1.09   error on traffic trip = -28.69

Now change it:

  • Remove the outlier: change 60 to 30. Predict the three values of w before you run it.
  • Change Huber's delta=2.0 to delta=40.0, then to delta=0.1. Predict which of the other two results it moves towards each time.
  • Use a fixed step size: replace the lr = ... line with lr = 0.02, and print w with four decimals. Predict which loss no longer lands exactly on its answer, and why.

Pause and think: The L1 model leaves a 30-minute error on the traffic trip, the largest of the three. Does that make it the worst model here?

Not necessarily. It is exactly right on all five normal trips, because it settles on the rate that most of the data agrees on. The L2 model cuts the traffic error to about 18 minutes, but pays for it by being almost 10 minutes too high on a normal 5 km trip. If traffic jams are rare events we do not want to model, L1 is the better fit. If big misses are very costly, we may prefer L2 or Huber.

Pause and think: Why does the training loop shrink the step size over time, and which loss needs that most?

L1 needs it most. Its gradient is always +1 or −1 per example, however close we are, so with a fixed step size w keeps hopping back and forth around the best value instead of settling. The L2 gradient shrinks by itself as the errors shrink, so it slows down naturally even with a fixed step size.

Key takeaways

  • A loss function turns errors into one number to minimise, and its shape decides which mistakes the model cares about.
  • L1 (MAE) = average |error|: linear penalty, robust to outliers, targets the median.
  • L2 (MSE) = average error²: quadratic penalty, smooth gradient, outlier-sensitive, targets the mean.
  • Choose L2 for clean data where big errors are costly; L1 when outliers or noisy labels are expected; Huber to combine both.
  • L1/L2 losses are about errors; L1/L2 regularisation is about weights.

Key terms

  • Loss function: A function that measures how wrong a model's predictions are, which training minimises.
  • Residual: The error for one example: true value minus predicted value.
  • L1 loss (MAE): The mean of the absolute errors.
  • L2 loss (MSE): The mean of the squared errors.
  • RMSE: The square root of MSE, in the same units as the target.
  • Outlier: A data point far from the others, which can heavily influence squared-error losses.
  • Huber loss: A loss that is quadratic for small errors and linear for large errors.

← 2.5 Precision and Recall: Picking the Right Metric · 2.7 Regularization: Stopping Overfitting with L1 and L2 →