Modern AI Engineering

Lesson 2.5 · 22 min

Precision and Recall: Picking the Right Metric

A spam filter that marks nothing as spam can still score 99% accuracy. So how do we actually tell whether a classifier is good?

In short: Precision asks: of everything the model flagged, how much was really positive? Recall asks: of everything that was really positive, how much did the model catch? Raising the decision threshold usually raises precision and lowers recall, so we choose based on which mistake costs more: false alarms (favour precision) or misses (favour recall). F1 combines both into one number.

The problem we are trying to solve

We have built a spam filter: a classifier that reads each email and outputs a spam score between 0 and 1. If the score is above a threshold (say 0.5), the email goes to the spam folder. How do we judge whether it is any good?

The obvious answer is accuracy: the fraction of emails classified correctly. But imagine an inbox of 1,000 emails where only 10 are spam. A lazy "filter" that never marks anything as spam is right on 990 emails: 99% accuracy, while catching zero spam. Accuracy hides the failure because the positive class (spam) is rare. This is called class imbalance, and it is the normal situation for fraud, disease, defects, and many other things we care about.

A fishing net Think of the model as a fishing net, and spam as the fish we want. Precision asks: of everything in the net, how much is fish and how much is old boots? Recall asks: of all the fish in the lake, how many ended up in the net? A tiny net gives a clean catch (high precision) but misses most fish (low recall). A huge net catches every fish (high recall) and lots of boots too (low precision).

To measure those two ideas precisely, we first need names for the four ways a single prediction can turn out.

The four possible outcomes

For a yes/no classifier, "positive" means the class we are looking for (spam) and "negative" means the other class (a normal email, often called "ham"). Each prediction is either right or wrong, and either positive or negative, giving four outcomes:

OutcomeModel saidTruthIn our spam filter
True Positive (TP)SpamSpamSpam correctly sent to the spam folder
False Positive (FP)SpamNot spamA real email wrongly hidden in spam (a false alarm)
False Negative (FN)Not spamSpamSpam that slipped into the inbox (a miss)
True Negative (TN)Not spamNot spamA real email correctly left in the inbox

A handy way to read the names: the second word (Positive or Negative) is what the model predicted; the first word (True or False) says whether that prediction was correct. Arranged in a 2 × 2 grid, these counts form the confusion matrix.

In statistics, a false positive is also called a Type I error and a false negative a Type II error. Accuracy is (TP + TN) / total, which here is (7 + 8) / 20 = 0.75.

What is precision?

Precision is the fraction of the model's positive predictions that were actually positive. It answers: "When the model says spam, how often is it right?"

High precision means few false alarms. Notice what precision ignores: the spam the model never flagged (false negatives) does not appear in the formula at all. A filter that flags only the single most obvious spam email can have perfect precision of 1.0 while missing everything else.

What is recall?

Recall is the fraction of actual positives that the model caught. It answers: "Of all the real spam, how much did we find?" Recall is also called sensitivity or the true positive rate.

High recall means few misses. Recall ignores false alarms: a filter that sends every email to spam has perfect recall of 1.0, and is useless.

Pause and think: Pause and predict: a model flags 50 emails as spam. 40 of them really are spam. There are 80 spam emails in total. What are its precision and recall?

TP = 40, FP = 50 − 40 = 10, FN = 80 − 40 = 40. Precision = 40 / 50 = 0.80. Recall = 40 / 80 = 0.50. It is usually right when it flags, but it misses half of the spam.

Precision vs recall: the trade-off

Precision and recall usually pull against each other, and the threshold is the lever. Raise the threshold and the model flags only emails it is very sure about: fewer false alarms (precision up) but more misses (recall down). Lower it and it flags more: more spam caught (recall up) but more good emails caught too (precision down).

Watching the threshold move

  1. Threshold 0.9: Only 2 emails flagged, both spam. Precision 1.00, recall 0.25. Very safe, misses most spam.
  2. Threshold 0.7: 7 flagged, 5 are spam. Precision 0.71, recall 0.625.
  3. Threshold 0.5: 11 flagged, 7 are spam. Precision 0.64, recall 0.875.
  4. Threshold 0.3: 15 flagged, all 8 spam caught. Precision 0.53, recall 1.00. Nothing missed, many false alarms.

When we want a single number that balances both, we use the F1 score, the harmonic mean of precision and recall:

Code: computing everything from scratch

precision_recall.py

import numpy as np
# 20 emails: model's spam score (0-1) and the truth (1 = spam, 0 = not spam)
scores = np.array([0.95, 0.91, 0.88, 0.85, 0.80, 0.74, 0.70, 0.66, 0.62, 0.55,
0.52, 0.47, 0.41, 0.38, 0.33, 0.27, 0.21, 0.15, 0.09, 0.04])
truth  = np.array([1,    1,    1,    0,    1,    1,    0,    1,    0,    1,
0,    0,    1,    0,    0,    0,    0,    0,    0,    0])
def evaluate(threshold):
pred = (scores >= threshold).astype(int)      # flag as spam if score >= threshold
tp = int(((pred == 1) & (truth == 1)).sum())  # spam caught
fp = int(((pred == 1) & (truth == 0)).sum())  # good email wrongly flagged
fn = int(((pred == 0) & (truth == 1)).sum())  # spam missed
tn = int(((pred == 0) & (truth == 0)).sum())  # good email left alone
precision = tp / (tp + fp) if tp + fp else 0.0
recall = tp / (tp + fn)
f1 = 2 * precision * recall / (precision + recall) if precision + recall else 0.0
return tp, fp, fn, tn, precision, recall, f1
print("thresh  TP FP FN TN  precision recall   F1")
for t in (0.9, 0.7, 0.5, 0.3):
tp, fp, fn, tn, p, r, f1 = evaluate(t)
print(f"  {t:.1f}   {tp:2d} {fp:2d} {fn:2d} {tn:2d}    {p:.2f}     {r:.2f}   {f1:.2f}")
# Accuracy can mislead: a model that NEVER flags spam
acc_lazy = (truth == 0).mean()
print(f"'never spam' model: accuracy {acc_lazy:.2f}, recall 0.00")

Output:

thresh  TP FP FN TN  precision recall   F1
0.9    2  0  6 12    1.00     0.25   0.40
0.7    5  2  3 10    0.71     0.62   0.67
0.5    7  4  1  8    0.64     0.88   0.74
0.3    8  7  0  5    0.53     1.00   0.70
'never spam' model: accuracy 0.60, recall 0.00

When to use which one?

The right metric depends on which mistake is more expensive in our product.

Precision and recall in AI engineering In a RAG system (Module 10), retrieval recall asks: did the retriever find the documents that contain the answer? If not, the LLM cannot answer correctly however good it is, so retrieval often favours recall, and a reranker then improves precision by pushing the best passages to the top. Guardrails (Module 15) face the same trade-off: blocking too eagerly annoys users (low precision); blocking too little lets harmful content through (low recall).

Pause and think: A hospital screening test flags patients for a follow-up scan. Missing a sick patient is far worse than an extra scan. Which metric should we prioritise, and should the threshold go up or down?

Prioritise recall and lower the threshold, so almost every sick patient is flagged. The extra false positives are handled by the follow-up scan.

Going one level deeper

Here is a surprise that catches many teams. We test a filter in the lab, get a good precision, ship it, and precision collapses. The model did not change. The share of positives in the data did. That share is called the base rate (or prevalence).

To see why, we describe the model by two numbers that do not depend on the base rate. Recall is the share of real positives it flags. The false positive rate is the share of real negatives it wrongly flags: FPR = FP / (FP + TN). Take an illustrative filter with recall 0.90 and FPR 0.05, and run it on 10,000 emails.

Same filter, three different inboxes

  1. Half the emails are spam: 5,000 spam and 5,000 normal. TP = 0.90 × 5,000 = 4,500. FP = 0.05 × 5,000 = 250. Precision = 4,500 / 4,750 ≈ 0.95.
  2. One in ten is spam: 1,000 spam and 9,000 normal. TP = 900. FP = 0.05 × 9,000 = 450. Precision = 900 / 1,350 ≈ 0.67.
  3. One in a hundred is spam: 100 spam and 9,900 normal. TP = 90. FP = 0.05 × 9,900 = 495. Precision = 90 / 585 ≈ 0.15.
  4. What changed: Recall stayed at 0.90 every time. But the pile of normal emails grew, and 5% of a big pile is a lot of false alarms. They swamp the few true positives.
  • How to spot it: precision in production is far below precision on the test set, while recall looks about the same. Compare the share of positives in the two datasets.
  • How to avoid it: build the test set with the same base rate we expect in production. A test set that was balanced to 50/50 for convenience will overstate precision.
  • How to fix it: when positives are rare, we need a much lower false positive rate. That means a higher threshold, a better model, or a second-stage check on the flagged items.

Practice: try it yourself

We will pick a threshold the way a product team should: by adding up what the mistakes cost. We switch to a fraud detector with 16 scored card payments, put a price on a false alarm and on a miss, and let the code find the cheapest threshold. Then we flip the prices and watch the answer move.

practice_cost_threshold.py

# Pick a threshold by total cost, not by habit.
# 16 card payments: the model's fraud score and the truth (1 = fraud).
scores = [0.97, 0.92, 0.86, 0.81, 0.77, 0.69, 0.63, 0.58,
0.51, 0.44, 0.39, 0.31, 0.26, 0.18, 0.12, 0.05]
truth = [1, 1, 0, 1, 0, 1, 0, 0,
1, 0, 0, 0, 1, 0, 0, 0]
def counts(threshold):
"""Return (false positives, false negatives) at this threshold."""
flagged = [s >= threshold for s in scores]
fp = sum(f and t == 0 for f, t in zip(flagged, truth))
fn = sum((not f) and t == 1 for f, t in zip(flagged, truth))
return fp, fn
def best_threshold(cost_fp, cost_fn):
"""Try each threshold and keep the one with the lowest total cost."""
best = None
for t in (0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9):
fp, fn = counts(t)
cost = fp * cost_fp + fn * cost_fn
print(f"  t={t:.1f}  FP={fp}  FN={fn}  cost={cost:3d}")
if best is None or cost < best[1]:
best = (t, cost)
return best
for cost_fp, cost_fn in ((1, 20), (20, 1)):
print(f"a false alarm costs {cost_fp}, a miss costs {cost_fn}:")
t, cost = best_threshold(cost_fp, cost_fn)
print(f"  -> best threshold {t:.1f} with total cost {cost}")

Output:

a false alarm costs 1, a miss costs 20:
t=0.1  FP=9  FN=0  cost=  9
t=0.2  FP=7  FN=0  cost=  7
t=0.3  FP=7  FN=1  cost= 27
t=0.4  FP=5  FN=1  cost= 25
t=0.5  FP=4  FN=1  cost= 24
t=0.6  FP=3  FN=2  cost= 43
t=0.7  FP=2  FN=3  cost= 62
t=0.8  FP=1  FN=3  cost= 61
t=0.9  FP=0  FN=4  cost= 80
-> best threshold 0.2 with total cost 7
a false alarm costs 20, a miss costs 1:
t=0.1  FP=9  FN=0  cost=180
t=0.2  FP=7  FN=0  cost=140
t=0.3  FP=7  FN=1  cost=141
t=0.4  FP=5  FN=1  cost=101
t=0.5  FP=4  FN=1  cost= 81
t=0.6  FP=3  FN=2  cost= 62
t=0.7  FP=2  FN=3  cost= 43
t=0.8  FP=1  FN=3  cost= 23
t=0.9  FP=0  FN=4  cost=  4
-> best threshold 0.9 with total cost 4

Now change it:

  • Make both mistakes cost the same: change the price pairs to ((5, 5),). Predict which threshold wins before you run it. (Hint: now only the total number of mistakes matters.)
  • Change the truth of the payment scored 0.26 from 1 to 0. Predict the new best threshold when a miss costs 20.
  • Inside best_threshold, also print precision and recall for each threshold (there are 6 frauds in total). Predict which of the two rises as the threshold goes up.

Pause and think: When a miss costs 20, thresholds 0.1 and 0.2 both miss nothing (FN = 0). Why does the code prefer 0.2?

Because 0.2 has fewer false alarms: 7 instead of 9. Lowering the threshold from 0.2 to 0.1 flags two more good payments and catches no extra fraud, so it only adds cost. Once recall is already 1.0, going lower can only hurt precision.

Pause and think: Between thresholds 0.2 and 0.3, FP stays at 7 but FN goes from 0 to 1, and the cost jumps from 7 to 27. What does this tell us about the data, without looking at it?

Exactly one payment has a score between 0.2 and 0.3, and it is fraud (it is the one scored 0.26). Raising the threshold past it turned a true positive into a miss and removed no false alarm. A single low-scoring positive like this is what forces a recall-first system to use a very low threshold.

A quick recap of the formulas, and summary

All the formulas in one place
MetricFormulaQuestion it answers
Accuracy(TP + TN) / (TP + TN + FP + FN)How often is the model right overall?
PrecisionTP / (TP + FP)When it says positive, how often is it right?
Recall (sensitivity)TP / (TP + FN)Of all real positives, how many did it find?
SpecificityTN / (TN + FP)Of all real negatives, how many did it correctly leave alone?
F1 score2 · P · R / (P + R)One number that is high only when both P and R are high

Common mistakes Reporting accuracy on imbalanced data; reporting precision without recall (or the reverse), since each can be made perfect by a silly model; mixing up which class is "positive"; and leaving the threshold at 0.5 by habit instead of choosing it for the business cost of each error.

Summary: precision measures how trustworthy the model's positive predictions are; recall measures how complete they are. The threshold trades one for the other. Choose based on the cost of false alarms vs misses, use F1 (or F-beta) when you need one number, and never rely on accuracy alone when one class is rare.

Key takeaways

  • Every prediction is a TP, FP, FN or TN; together they form the confusion matrix.
  • Precision = TP / (TP + FP): how trustworthy positive predictions are.
  • Recall = TP / (TP + FN): how many real positives we found.
  • Raising the threshold usually raises precision and lowers recall; lowering it does the opposite.
  • Favour precision when false alarms are costly and recall when misses are costly; use F1 to balance both.
  • Accuracy misleads on imbalanced data.

Key terms

  • Confusion matrix: A table of counts of true positives, false positives, false negatives and true negatives.
  • Precision: The fraction of predicted positives that are actually positive: TP / (TP + FP).
  • Recall: The fraction of actual positives that the model found: TP / (TP + FN).
  • F1 score: The harmonic mean of precision and recall.
  • Threshold: The score above which a classifier predicts the positive class.
  • Class imbalance: When one class is much rarer than the other, making accuracy misleading.

← 2.4 Feature Engineering: Turning Raw Data into Signal · 2.6 L1 vs L2 Loss: Choosing Your Error Penalty →