Modern AI Engineering

Lesson 2.1 · 22 min

Machine Learning from First Principles

How can a program get good at a task that nobody ever wrote the rules for, like spotting spam or pricing a house?

In short: Machine learning is a way of building software where, instead of writing rules by hand, we show the computer many examples and let an algorithm adjust a model's numbers until its predictions match the examples. Training is a loop: predict, measure the error with a loss, nudge the parameters to reduce it, repeat. The real test is how well the model does on new data it has never seen.

Why we need machine learning

Traditional programming works like a recipe: a programmer writes exact rules, the computer follows them. That is perfect for things like calculating tax or sorting a list, where we know the rules precisely.

But try writing rules for "is this email spam?" We might start with "if it contains free money, mark spam". Spammers then write "fr3e m0ney". We add more rules, they adapt, and soon we have thousands of brittle rules that still miss things. The same problem appears with recognising faces, understanding speech, or predicting house prices: the patterns are real, but too many and too subtle to write down.

Teaching a child to recognise dogs We never give a child a rulebook like "four legs, fur, tail, barks". We just point at many dogs and say "dog", and at cats and say "not a dog". After enough examples, the child recognises dogs they have never seen before, even odd-looking ones. Machine learning works the same way: learn from examples, then generalise to new cases.

Machine learning (ML) is the field of building programs that improve at a task by learning from data rather than from hand-written rules. A classic, often-quoted definition from Tom Mitchell (1997) says a program learns from experience E with respect to a task T and performance measure P if its performance on T, measured by P, improves with E. For a spam filter: T = classify emails, P = percent classified correctly, E = a pile of emails already labelled spam or not spam.

The core idea: data in, rules out

The cleanest way to see the shift is to compare what goes in and what comes out.

Let us define the vocabulary we will use for the rest of the course, using a running example: predicting a house's price.

  • Example (or sample): one row of data, such as one house.
  • Feature: an input measurement describing the example, such as size in square feet or number of bedrooms.
  • Label (or target): the answer we want to predict, such as the sale price.
  • Model: a mathematical function that maps features to a prediction, for example price = w × size + b.
  • Parameters (weights): the adjustable numbers inside the model, here w and b. Learning means finding good values for them.
  • Training: the process of adjusting parameters using example data.
  • Inference (prediction): using the trained model on new inputs.

The main kinds of machine learning

ML methods are grouped by what kind of feedback the learner gets. We meet each in detail later in this module.

KindWhat the data looks likeExample task
Supervised learningInputs paired with correct answers (labels)Predict house price; classify email as spam
Unsupervised learningInputs only, no answersGroup customers into segments; detect unusual transactions
Reinforcement learningAn agent acts and receives rewards or penaltiesLearn to play a game; control a robot
Self-supervised learningLabels are created automatically from the data itselfPredict the next word in a sentence (how LLMs are pre-trained)

Supervised tasks are further split by the type of label. Regression predicts a number (a price, a temperature). Classification predicts a category (spam or not spam, cat or dog). Lesson 2.3 contrasts the simplest model for each.

AI, ML, deep learning and LLMs These terms nest inside each other. Artificial intelligence is the broad goal of machines doing tasks that seem to need intelligence (it also includes non-learning methods like search and rule systems). Machine learning is the part of AI that learns from data. Deep learning is ML using neural networks with many layers (Module 3). LLMs are very large deep learning models trained on text (Module 4).

How a model learns, step by step

Almost every ML model, from a two-parameter line to a giant LLM, learns with the same loop.

One training step on a tiny example

  1. Start with a guess: Our model is price = w × size + b, with size in hundreds of square feet and price in thousands of dollars. We start with w = 0, b = 0.
  2. Predict: For a house of size 10 (1,000 sq ft) that sold for 250, we predict 0 × 10 + 0 = 0.
  3. Measure: The error is 0 − 250 = −250. Squared, that is 62,500. A huge loss: our guess is terrible.
  4. Find the direction: The prediction is too low, and w multiplies a positive size, so increasing w (and b) would raise the prediction and reduce the error.
  5. Update: Increase w and b by a small amount proportional to the error. Repeat over all houses, thousands of times, and the line settles where the total error is smallest.

This downhill-walking procedure is called gradient descent. The learning rate controls the step size: too small and learning crawls, too large and it overshoots and can diverge. Try it below; Lesson 3.2 covers it in depth.

Code: learning house prices from examples

We generate 50 houses from a hidden rule (price = 20 × size + 50, plus random noise) and pretend we do not know that rule. The model sees only the examples and must discover it. We keep 10 houses aside to test it.

learn_house_prices.py

import numpy as np
# Training data: house size (100s of sq ft) -> price ($1000s)
rng = np.random.default_rng(42)
size = rng.uniform(5, 30, 50)                      # 50 houses
price = 20 * size + 50 + rng.normal(0, 25, 50)     # hidden rule + noise
# Split: learn from 40 houses, test on 10 the model never saw
x_tr, y_tr, x_te, y_te = size[:40], price[:40], size[40:], price[40:]
w, b, lr = 0.0, 0.0, 0.002                         # start knowing nothing
for epoch in range(20001):
pred = w * x_tr + b                            # 1. predict
err = pred - y_tr                              # 2. measure error
loss = (err ** 2).mean()                       #    mean squared error
w -= lr * 2 * (err * x_tr).mean()              # 3. nudge w and b
b -= lr * 2 * err.mean()                       #    downhill on the loss
if epoch in (0, 10, 1000, 20000):
print(f"epoch {epoch:4d}  loss {loss:9.1f}  w={w:6.2f}  b={b:6.2f}")
test_mae = np.abs(w * x_te + b - y_te).mean()
print(f"learned rule: price = {w:.1f} * size + {b:.1f}")
print(f"average error on 10 unseen houses: ${test_mae:.1f}k")
print(f"prediction for a 2,000 sq ft house: ${w * 20 + b:.0f}k")

Output:

epoch    0  loss  192172.1  w= 34.45  b=  1.66
epoch   10  loss     639.6  w= 22.21  b=  1.33
epoch 1000  loss     437.1  w= 21.24  b= 21.04
epoch 20000  loss     327.2  w= 19.87  b= 49.92
learned rule: price = 19.9 * size + 49.9
average error on 10 unseen houses: $18.6k
prediction for a 2,000 sq ft house: $447k

Pause and think: Why is the average test error about $18.6k and not zero, even though the model found almost exactly the true rule?

Because the prices include random noise (standard deviation 25) that no rule based on size alone can predict. Real data always has factors the features do not capture, so some error is irreducible.

The real goal: generalising to new data

A model is only useful if it works on data it has never seen. This ability is called generalisation. To measure it honestly, we split our data: a training set to learn from, a validation set to tune choices like the learning rate, and a test set that we only look at at the very end.

Two failure modes appear again and again. Underfitting: the model is too simple to capture the pattern (a straight line for a curved relationship), so it does badly on both training and test data. Overfitting: the model is so flexible that it memorises the training examples, including their noise, so it looks great on training data and does badly on new data. Lesson 2.7 shows how regularisation fights overfitting.

The most common beginner mistake Evaluating on the same data used for training, or letting test data leak into training (for example, normalising using statistics from the whole dataset, or having duplicate rows in both splits). This gives impressive numbers that collapse in production.

A short history

Milestones on the road to modern AI

  1. The term "machine learning": Arthur Samuel popularises the term while building a checkers program that improved by playing against itself.
  2. Backpropagation popularised: Rumelhart, Hinton and Williams show how to train multi-layer neural networks efficiently (Lesson 3.3).
  3. Statistical ML: Methods like support vector machines, decision trees and random forests power spam filters, search and recommendations.
  4. Deep learning breakthrough: AlexNet, a deep convolutional network trained on GPUs, wins the ImageNet image-recognition challenge by a wide margin.
  5. The Transformer: The architecture behind modern LLMs is introduced in "Attention Is All You Need" (Module 4).
  6. LLMs go mainstream: ChatGPT and similar assistants bring large language models to hundreds of millions of users.

Where ML is used, and when not to use it

ML you used today Email spam filters, product and video recommendations, card fraud alerts, voice assistants, photo search ("show me beach pictures"), map arrival-time estimates, machine translation, and of course AI chat assistants all rely on machine learning.

A typical ML project follows a lifecycle: define the problem and the success metric, collect and clean data, engineer features (Lesson 2.4), train a model, evaluate on held-out data, deploy, and then monitor, because the world changes and models drift out of date.

  • Do not use ML when simple rules work. A shipping-cost calculator should be code, not a model.
  • Do not use ML without enough representative data. A model can only learn patterns present in its examples.
  • Be careful when errors are costly and must be explained (some legal or medical decisions); prefer simpler, interpretable models or keep a human in the loop.
  • Watch for bias in data. If historical decisions were unfair, a model trained on them learns the same unfairness.

Pause and think: A bank wants to compute monthly loan interest from a fixed published formula. Should it use machine learning?

No. The rule is known exactly, so ordinary code is simpler, exact, and explainable. ML is for patterns we cannot write down, such as predicting which loans are likely to default.

Worked example, step by step

Let us do one full gradient descent step by hand, on numbers small enough to check on paper. We use the simplest model, price = w × size (no b, to keep the sums short), and three made-up houses.

Three illustrative houses and our predictions with w = 10
SizeTrue pricePrediction (w = 10)Error (prediction − true)
12010−10
24020−20
36030−30

One update of w, by hand

  1. Measure the loss: Square each error and average: (100 + 400 + 900) / 3 ≈ 466.7. That is our mean squared error with w = 10.
  2. Find the slope: The gradient for w is 2 × mean(error × size). The products are −10×1, −20×2, −30×3 = −10, −40, −90. Their mean is −46.67, so the gradient is about −93.3.
  3. Read the sign: A negative gradient means the loss goes down when w goes up. That matches common sense: all three predictions are too low.
  4. Take the step: With learning rate 0.05: new w = 10 − 0.05 × (−93.3) ≈ 14.67.
  5. Check it helped: New predictions are 14.67, 29.33, 44.0. Errors are −5.33, −10.67, −16.0. The new loss is about 132.7, down from 466.7.
  6. Repeat: Here each step removes the same share of the remaining gap to the best value, w = 20. The gap shrinks from 10 to 5.33, then to about 2.84, and so on.

What a too-large step looks like Try learning rate 0.25 on the same numbers: new w = 10 + 0.25 × 93.3 ≈ 33.3. We jumped past 20 and landed further away than we started (a gap of 13.3 instead of 10). The loss goes up. If the loss rises step after step, the first thing to try is a smaller learning rate.

Practice: try it yourself

We will see underfitting and overfitting with real numbers. We fit curves of different flexibility (polynomials of degree 1, 2, 5 and 12) to 20 training points, then score each curve on 10 points it never saw.

practice_overfitting.py

import numpy as np
# 30 points from a gently curved hidden rule, plus noise
rng = np.random.default_rng(2)
x = np.sort(rng.uniform(-1, 1, 30))
y = 1.5 * x**2 - x + rng.normal(0, 0.15, 30)
# Every 3rd point is held out for testing; the model never trains on it
test = np.arange(30) % 3 == 0
x_tr, y_tr, x_te, y_te = x[~test], y[~test], x[test], y[test]
def mse(coefs, xs, ys):
"""Mean squared error of a polynomial on some points."""
return ((np.polyval(coefs, xs) - ys) ** 2).mean()
print(f"train points: {len(x_tr)}, test points: {len(x_te)}")
print("degree  train MSE  test MSE")
for degree in (1, 2, 5, 12):
coefs = np.polyfit(x_tr, y_tr, degree)   # fit on training data only
print(f"{degree:6d}  {mse(coefs, x_tr, y_tr):9.4f}  {mse(coefs, x_te, y_te):8.4f}")

Output:

train points: 20, test points: 10
degree  train MSE  test MSE
1     0.1558    0.2682
2     0.0189    0.0316
5     0.0121    0.0358
12     0.0065    0.3034

Now change it:

  • Add degree 0 to the list (a flat line). Predict: will its training error be higher or lower than degree 1, and is that underfitting or overfitting?
  • Change the noise from 0.15 to 0.0. Predict what happens to the test error of degree 2, and whether degree 12 still looks bad.
  • Change 30 to 300 points in all three places. Predict whether the gap between train and test error for degree 12 grows or shrinks, and why more data helps.

Pause and think: Degree 12 has the lowest training error in the table. Why is it still the worst choice here?

Because its test error (about 0.30) is roughly ten times that of degree 2 (about 0.03). It used its extra flexibility to chase the noise in the 20 training points. Low training error only shows the model can fit what it saw; the test column shows what it will do on new data.

Pause and think: Degree 1 and degree 12 have similar test errors. Are they failing for the same reason? How can we tell from the table?

No. Degree 1 is bad on both columns (about 0.16 train, 0.27 test): it is too simple for a curved rule, which is underfitting. Degree 12 is excellent on train and bad on test: a big gap between the columns is the sign of overfitting.

Key takeaways

  • Machine learning builds models from examples instead of hand-written rules.
  • Features are the inputs, labels are the answers, parameters are the numbers the model learns.
  • Training is a loop: predict, measure loss, compute the gradient, update parameters, repeat.
  • The goal is generalisation: always evaluate on data the model has never seen.
  • Underfitting means too simple; overfitting means memorising noise.
  • Use plain code when the rules are known; use ML when examples are easier to get than rules.

Key terms

  • Machine learning: Building programs that improve at a task by learning patterns from data rather than following hand-written rules.
  • Feature: An input variable that describes an example, such as a house's size.
  • Label: The correct answer for an example that a supervised model learns to predict.
  • Parameters: The adjustable numbers inside a model that training sets, such as weights and biases.
  • Loss function: A formula that turns the model's errors into a single number to minimise.
  • Generalisation: How well a trained model performs on new data it has not seen.
  • Overfitting: When a model learns the training data, including its noise, too closely and does worse on new data.

← 1.1 Six Concepts Every AI Engineer Must Know · 2.2 Labeled vs Unlabeled: Two Ways Machines Learn →