Lesson 2.3 · 23 min
Predicting Numbers vs Categories: Regression Compared
Despite the shared word "regression", one of these models predicts numbers and the other answers yes-or-no questions. How can one small change, a sigmoid, turn the first into the second?
In short: Linear regression predicts a continuous number by fitting a straight line (or plane) that minimises squared error. Logistic regression predicts the probability of a yes/no outcome by passing the same kind of linear score through a sigmoid that squashes it into 0 to 1, and it is trained with log loss. Use linear regression for "how much?" and logistic regression for "which class?".
Two kinds of questions
Imagine a teacher with data on students: how many hours each one studied, the score they got, and whether they passed. Two natural questions arise:
- "If a student studies 5.5 hours, what score will they probably get?" The answer is a number on a continuous scale.
- "If a student studies 5.5 hours, will they pass?" The answer is yes or no, and ideally a probability like "50% chance".
The first is a regression problem (predict a number). The second is a classification problem (predict a category). Both are supervised learning, which we met in Lesson 2.2. The two simplest, most widely used models for them are linear regression and logistic regression.
A ruler and a dimmer switch Linear regression is like a ruler laid through the data: it can extend to any value, as high or low as the line goes. Logistic regression is like a dimmer switch that is fully off at one end and fully on at the other, with a smooth fade in between. It always stays between 0 (off, "fail") and 1 (on, "pass").
About the confusing name Logistic regression is a classification method, despite its name. The name is historical: it regresses (fits) the log-odds of the outcome with a linear function, which we unpack below.
Linear regression
Linear regression models the target as a weighted sum of the features plus a constant:
With one feature this is just the school formula for a line, ŷ = w·x + b. Training means finding the w and b that make predictions as close as possible to the true values. "Close" is usually measured by mean squared error (MSE): the average of the squared differences between prediction and truth.
Fitting a line, step by step
- Collect pairs: For each student we have (hours, score), e.g. (1, 35), (2, 41), … (10, 90).
- Propose a line: Pick any w and b, e.g. w = 5, b = 30, so 4 hours predicts 50.
- Measure error: Compare each prediction with the true score, square the differences, and average them to get the MSE.
- Improve the line: Adjust w and b to reduce the MSE, either with gradient descent or directly with a closed-form formula called ordinary least squares.
- Use it: Our data gives about score = 6.04 × hours + 29.47, so each extra hour adds about 6 points.
A big advantage of linear regression is interpretability: each weight has a plain meaning. Its main assumption is that the relationship is roughly a straight line (in the features we give it). If the real relationship is curved, we can still use linear regression by adding engineered features such as x², which we explore in Lesson 2.4.
Logistic regression
For pass/fail we want a probability, a number between 0 and 1. A straight line cannot give that: it keeps rising forever, so with enough study hours it would predict a "probability" of 1.8 or more. Logistic regression fixes this with one extra step. It computes the same linear score z = w·x + b, then squashes it with the sigmoid function:
To turn the probability into a class we pick a threshold, usually 0.5: predict "pass" if σ(z) ≥ 0.5. Since σ(z) = 0.5 exactly when z = 0, the decision boundary is where w·x + b = 0. With one feature that is a single cut-off point; with two features it is a straight line; in general it is a flat plane. That is why logistic regression is called a linear classifier.
Why "log-odds"? The odds of passing are p / (1 − p). Taking the log of the odds gives exactly z: ln(p / (1 − p)) = w·x + b. So each weight says how much one unit of a feature changes the log-odds. For example, a weight of 1.30 on hours means each extra hour multiplies the odds of passing by e¹·³⁰ ≈ 3.7.
Logistic regression is trained with log loss (also called binary cross-entropy), not MSE:
Code: both models on the same students
Ten students, their study hours, their scores (for linear regression) and whether they passed (for logistic regression). Notice the student who studied 6 hours but failed: real data is noisy, which is why probabilities are useful.
linear_vs_logistic.py
import numpy as np
# Data: hours studied -> exam score (number) and passed? (0 or 1)
hours = np.array([1, 2, 3, 4, 5, 6, 7, 8, 9, 10], dtype=float)
score = np.array([35, 41, 50, 52, 61, 64, 72, 79, 83, 90], dtype=float)
passed = np.array([0, 0, 0, 0, 1, 0, 1, 1, 1, 1], dtype=float)
# --- Linear regression: least squares line for the score ---
X = np.column_stack([hours, np.ones_like(hours)]) # [x, 1] for slope and intercept
w, b = np.linalg.lstsq(X, score, rcond=None)[0]
print(f"linear: score = {w:.2f} * hours + {b:.2f}")
print(f" predicted score for 5.5 h: {w * 5.5 + b:.1f}")
# --- Logistic regression: sigmoid(w*x + b) gives P(pass) ---
sigmoid = lambda z: 1 / (1 + np.exp(-z))
w2, b2, lr = 0.0, 0.0, 0.1
for _ in range(20000):
p = sigmoid(w2 * hours + b2)
w2 -= lr * ((p - passed) * hours).mean() # gradient of log loss
b2 -= lr * (p - passed).mean()
print(f"logistic: P(pass) = sigmoid({w2:.2f} * hours + {b2:.2f})")
for h in (2, 5, 5.5, 9):
print(f" {h:>4} h -> P(pass) = {sigmoid(w2 * h + b2):.2f}")
# Why not just fit a straight line to the 0/1 labels?
wl, bl = np.linalg.lstsq(X, passed, rcond=None)[0]
print(f"line on 0/1 labels at 15 h: {wl * 15 + bl:.2f} (not a valid probability)")Output:
linear: score = 6.04 * hours + 29.47 predicted score for 5.5 h: 62.7 logistic: P(pass) = sigmoid(1.30 * hours + -7.16) 2 h -> P(pass) = 0.01 5 h -> P(pass) = 0.34 5.5 h -> P(pass) = 0.50 9 h -> P(pass) = 0.99 line on 0/1 labels at 15 h: 1.82 (not a valid probability)
Pause and think: Using the learned model, what is the log-odds z for a student who studies 7 hours, and is the predicted class pass or fail?
z = 1.30 × 7 − 7.16 = 1.94. Since z > 0, σ(z) > 0.5 (about 0.87), so we predict pass.
Differences between linear and logistic regression
More than two classes Logistic regression extends to many classes with softmax regression (multinomial logistic regression): one linear score per class, turned into probabilities that sum to 1 by the softmax function. This is exactly what the final layer of an LLM does when it picks the next token from its vocabulary.
Real-world use
| Question | Model | Target |
|---|---|---|
| What will this flat rent for? | Linear regression | Monthly rent in dollars |
| How many support tickets will arrive tomorrow? | Linear regression (often with time features) | Ticket count |
| Will this customer cancel next month? | Logistic regression | Churn: yes/no |
| Is this transaction fraud? | Logistic regression (as a baseline) | Fraud: yes/no |
| Will the user click this ad? | Logistic regression has long been a standard choice in ad click prediction | Click: yes/no |
Why simple models still matter in the LLM era Both models train in milliseconds, need little data, and can be explained to a regulator or a manager weight by weight. Teams routinely use them as a baseline: if a giant model cannot beat logistic regression on a task, it is not worth its cost. Logistic regression is also a common "probe" for checking whether an embedding contains some piece of information.
Common mistakes and limits
Using linear regression for a yes/no target It can predict values below 0 or above 1, and its squared-error loss treats classification poorly. Use logistic regression for binary outcomes.
- Thinking logistic regression can draw curved boundaries on its own. Its boundary is linear in the features. For curved boundaries, add engineered features (like x² or interactions) or use a non-linear model such as a tree ensemble or neural network.
- Treating 0.5 as a sacred threshold. The threshold is a business choice. For cancer screening we might flag anything above 0.1; for auto-blocking payments we might need 0.95. Lesson 2.5 explains this trade-off.
- Forgetting outliers. One extreme house price can tilt a linear regression line a lot because errors are squared (Lesson 2.6).
- Reading weights when features are on different scales or strongly correlated. A weight's size depends on the feature's units, and correlated features can split or swap weight between them, so scale features before comparing weights.
- Extrapolating far outside the training range. A line fitted on 1–10 hours says 20 hours gives a score of 150 out of 100.
Pause and think: A model predicts whether an email is spam. On some emails, it outputs 1.3 and on others −0.2. Which model was probably used, and what should we use instead?
Linear regression, because only an unbounded line can output values outside 0–1. Use logistic regression, whose sigmoid output is always a valid probability.
Worked example, step by step
The code above ran 20,000 updates in a blink. Let us slow down and do the very first update of a logistic regression by hand. We use four made-up students so every number fits on one line.
| Hours | Passed (y) | p at start (w = 0, b = 0) | p − y | p after one update |
|---|---|---|---|---|
| 2 | 0 | 0.50 | +0.50 | 0.55 |
| 4 | 0 | 0.50 | +0.50 | 0.60 |
| 6 | 1 | 0.50 | −0.50 | 0.65 |
| 8 | 1 | 0.50 | −0.50 | 0.69 |
One update of w and b, by hand
- Predict: With
w = 0andb = 0, every score is z = 0, so every probability is σ(0) = 0.5. The model knows nothing yet. - Measure the loss: Each student costs −ln(0.5) ≈ 0.693, so the log loss is 0.693. This is the loss of a coin flip, a useful number to remember.
- Gradient for w: Average (p − y) × hours: (0.5×2 + 0.5×4 − 0.5×6 − 0.5×8) / 4 = (1 + 2 − 3 − 4) / 4 = −1.0. Negative, so
wshould go up. - Gradient for b: Average (p − y): (0.5 + 0.5 − 0.5 − 0.5) / 4 = 0. The classes are balanced, so
bdoes not move on this step. - Update: w = 0 − 0.1 × (−1.0) = 0.1 and b = 0. New scores are z = 0.2, 0.4, 0.6, 0.8, which give p ≈ 0.55, 0.60, 0.65, 0.69.
- Check it helped: The loss falls from 0.693 to about 0.630. The passing students moved the right way. The failing ones moved the wrong way, from 0.50 up to 0.55 and 0.60.
That last point is the interesting one. With b = 0, a positive w pushes every student above 0.5, because hours are never negative. The next update fixes it: the gradient for b is now the average of (p − y) ≈ (0.55 + 0.60 − 0.35 − 0.31) / 4 ≈ +0.12, which is positive, so b moves down. Step by step, w grows and b falls until the boundary −b / w sits between 4 and 6 hours.
A quick health check for any classifier Compare the log loss with 0.693. A two-class model that scores about 0.693 on balanced data has learned nothing yet. A model far above 0.693 is worse than a coin flip, which usually means confident wrong answers or a bug such as swapped labels.
Practice: try it yourself
We will train logistic regression on data that is perfectly separable: one cut-off splits every fail from every pass with no exceptions. Then we keep training for longer and longer and watch the weight, the boundary and the loss. One of them never settles.
practice_separable.py
import numpy as np
# Eight students. The labels are perfectly separable: everyone at or
# below 4 hours failed and everyone at or above 5 hours passed.
hours = np.array([1, 2, 3, 4, 5, 6, 7, 8], dtype=float)
passed = np.array([0, 0, 0, 0, 1, 1, 1, 1], dtype=float)
def sigmoid(z):
return 1 / (1 + np.exp(-z))
w, b, lr, done = 0.0, 0.0, 0.1, 0
print(" steps w b boundary P(pass | 5 h) log loss")
for target in (100, 1_000, 10_000, 100_000):
while done < target: # keep training from where we stopped
p = sigmoid(w * hours + b)
w -= lr * ((p - passed) * hours).mean() # gradient of log loss
b -= lr * (p - passed).mean()
done += 1
p = sigmoid(w * hours + b)
loss = -(passed * np.log(p) + (1 - passed) * np.log(1 - p)).mean()
print(f"{done:7d} {w:5.2f} {b:6.2f} {-b / w:8.2f} {sigmoid(w * 5 + b):13.3f} {loss:8.4f}")Output:
steps w b boundary P(pass | 5 h) log loss 100 0.43 -1.38 3.20 0.685 0.4081 1000 1.32 -5.68 4.29 0.718 0.1512 10000 3.18 -14.16 4.46 0.849 0.0490 100000 6.82 -30.58 4.48 0.971 0.0082
Now change it:
- Swap two labels so the data is no longer separable: set
passedto[0, 0, 0, 1, 0, 1, 1, 1]. Predict: doeswstill keep growing between 10,000 and 100,000 steps? - Add a small pull towards zero inside the loop, right after the
wupdate:w -= lr 0.01 w. Predict whetherwand the loss still change much between the last two rows. - Change
lrfrom0.1to1.0. Predict whetherwat 100 steps is larger or smaller than before, and whether the boundary gets close to 4.5 sooner.
Pause and think: The boundary barely moves after 10,000 steps (4.46, then 4.48), but w more than doubles. Why does training keep making w bigger?
Because the data is perfectly separable. Once the boundary sits between 4 and 5 hours, making w and b larger keeps the boundary in place but makes the curve steeper, so every probability moves closer to its label and the log loss falls a little more. There is no finite best value. In practice we stop early or add regularisation (Lesson 2.7) to keep the weights sensible.
Pause and think: Use the row for 100 steps. Is a student with 4 hours classified as pass or fail, and is that right?
z = 0.43 × 4 − 1.38 = 0.34, which is above 0, so the model says pass. That is wrong: this student failed. The boundary is still at 3.20 hours, below 4. The model has not trained long enough, and the high log loss (0.41) is the warning sign.
Key takeaways
- Linear regression predicts a number: ŷ = w·x + b, trained by minimising mean squared error.
- Logistic regression predicts a probability: p = σ(w·x + b), trained by minimising log loss.
- The sigmoid keeps outputs between 0 and 1; the decision boundary is where w·x + b = 0.
- Logistic regression is a classifier despite its name, and its boundary is linear in the features.
- Both are fast, interpretable baselines that every bigger model should beat.
Key terms
- Linear regression: A model that predicts a continuous value as a weighted sum of features plus a bias.
- Logistic regression: A classification model that outputs a probability by applying the sigmoid to a linear score.
- Sigmoid: The function σ(z) = 1 / (1 + e⁻ᶻ), which maps any real number into the range 0 to 1.
- Log-odds (logit): ln(p / (1 − p)); logistic regression models this as a linear function of the features.
- Log loss: Binary cross-entropy: the loss that heavily penalises confident wrong probability predictions.
- Decision boundary: The set of inputs where the model switches between predicted classes.
← 2.2 Labeled vs Unlabeled: Two Ways Machines Learn · 2.4 Feature Engineering: Turning Raw Data into Signal →