Lesson 3.1 · 20 min
Neural Network Bias: What It Is and Why It Matters
If a neuron can already multiply every input by a learned weight, why does it also need one extra number that ignores the input completely?
In short: A bias is a learned constant added to a neuron's weighted sum: z = w·x + b. It shifts the neuron's output up or down, so the neuron can fit lines that do not pass through the origin and can choose where its activation switches on. Without bias, many simple patterns become impossible or much harder to learn.
The problem: a line stuck at the origin
Let us start with a tiny running example. We want to predict a house price from its size. In our toy town every house costs a fixed $300k for the land plus $200k per 100 m² of floor space. So a 100 m² house costs $500k, a 200 m² house costs $700k, and so on. Written as a rule: price = 2 · size + 3 (price in $100k, size in 100s of m²).
The simplest artificial neuron computes a weighted sum of its inputs: it multiplies each input by a number called a weight and adds the results. With one input, that is just z = w · x. Here is the catch: when x = 0, the output is always 0, no matter what w is. The neuron can only draw lines that pass through the point (0, 0), the origin.
Our house rule says a house of size zero still costs $300k (the land). A line forced through the origin can never say that. It can tilt (change w), but it cannot lift itself up. That missing ability to lift and lower is exactly what the bias gives us.
Think of it like a taxi meter A taxi fare is a per-kilometre rate times the distance (the weight times the input) plus a fixed flag-down fee you pay the moment you sit in the car (the bias). A meter with no flag-down fee can only charge for distance. It cannot model the fixed part of the price, however carefully it tunes the per-kilometre rate.
What a bias is, in plain words
A bias is one extra learnable number per neuron that is added after the weighted sum and before the activation function. It does not multiply any input. It is simply added.
Two words that are easy to mix up: an activation function is the non-linear function f applied to z (for example ReLU(z) = max(0, z) or the sigmoid σ(z) = 1 / (1 + e⁻ᶻ)). The pre-activation is z itself. The bias lives inside z, so it changes what the activation function sees.
Weights and biases are both parameters: numbers the network learns during training. The weights control the slope or sensitivity to each input. The bias controls the offset: the output the neuron gives when all inputs are zero, and therefore where the neuron's decision boundary sits.
Pause and think: A neuron has weights w = [0.5, −1.0], bias b = 2, and receives inputs x = [0, 0]. What is its pre-activation z?
z = 0.5·0 + (−1.0)·0 + 2 = 2. With all inputs at zero the weights contribute nothing, so the output is decided entirely by the bias. Without a bias it would be stuck at 0.
Two jobs a bias does
Job 1: shift a line or plane away from the origin. With one input, z = w·x + b is the familiar line y = mx + c from school. w is the slope and b is the intercept, the height where the line crosses the y-axis. With many inputs, w·x + b = 0 describes a flat boundary (a hyperplane). Without b, every such boundary must pass through the origin. With b, it can sit anywhere.
Job 2: set the threshold of the activation. Take a ReLU neuron, max(0, w·x + b). It outputs zero until w·x + b becomes positive. With w = 1, the neuron switches on when x > −b. So a bias of −1 means "only fire when x is above 1", and a bias of +1 means "fire as soon as x is above −1". The bias moves the switching point left and right.
The same idea holds for a sigmoid neuron. σ(w·x + b) crosses 0.5 exactly where w·x + b = 0, that is at x = −b / w. The weight decides how steep the S-curve is; the bias decides where its middle sits. A classifier that says "spam if probability > 0.5" therefore has its decision point set by the bias.
Why the name "bias"? The word here means "a built-in lean" towards a higher or lower output, before the evidence (the inputs) is considered. It is unrelated to social bias in AI fairness, and it is also different from the bias in the statistical "bias–variance trade-off". Same word, three different ideas.
How a bias is learned, step by step
A bias is trained exactly like a weight: by gradient descent (covered in the next lesson). We measure how wrong the network is with a loss function, compute how the loss changes when we nudge each parameter (its gradient), and move each parameter a little in the direction that lowers the loss.
One training step for a single neuron with bias
- Start with guesses: Initialise
wwith a small random number. The bias is very often initialised to0; this is safe because the random weights already make neurons differ from each other. - Forward pass: Compute the prediction
ŷ = w·x + bfor each training example. - Measure the error: Use a loss such as mean squared error,
L = mean((ŷ − y)²). - Compute gradients: For this loss,
∂L/∂w = 2·mean((ŷ − y)·x)and∂L/∂b = 2·mean(ŷ − y). The bias gradient is just the average error, because the bias behaves like a weight on an input that is always 1. - Update:
w ← w − η·∂L/∂wandb ← b − η·∂L/∂b, whereη(eta) is the learning rate. Repeat until the loss stops falling.
Notice the neat interpretation of ∂L/∂b: if the neuron is, on average, predicting too high, the average error is positive and the bias goes down. If it is predicting too low, the bias goes up. The bias absorbs the overall offset in the data.
Code you can run: with and without bias
We train two single-neuron models on our house-price data with plain gradient descent. One has a bias, one does not. Then we look at how the bias moves a ReLU neuron's switching point.
bias_demo.py
import numpy as np
# Toy data: house size (100s of m²) -> price (in $100k). True rule: price = 2*size + 3
size = np.array([1.0, 2.0, 3.0, 4.0, 5.0])
price = 2 * size + 3
def train(use_bias, lr=0.02, steps=5000):
w, b = 0.0, 0.0
for _ in range(steps):
pred = w * size + (b if use_bias else 0.0)
err = pred - price
w -= lr * 2 * np.mean(err * size) # dL/dw for mean squared error
if use_bias:
b -= lr * 2 * np.mean(err) # dL/db
pred = w * size + (b if use_bias else 0.0)
return w, b, np.mean((pred - price) ** 2)
for use_bias in (False, True):
w, b, mse = train(use_bias)
print(f"bias={use_bias!s:5} w={w:.3f} b={b:.3f} MSE={mse:.4f}")
# A neuron with ReLU: bias decides when it "switches on"
x = np.array([-2.0, -1.0, 0.0, 1.0, 2.0])
for b in (-1.0, 0.0, 1.0):
out = np.maximum(0, 1.0 * x + b)
print(f"b={b:+.0f} ReLU(x + b) = {out}")Output:
bias=False w=2.818 b=0.000 MSE=1.6364 bias=True w=2.000 b=3.000 MSE=0.0000 b=-1 ReLU(x + b) = [0. 0. 0. 0. 1.] b=+0 ReLU(x + b) = [0. 0. 0. 1. 2.] b=+1 ReLU(x + b) = [0. 0. 1. 2. 3.]
The bias-free model tries to compensate by making the slope steeper. It over-predicts large houses and under-predicts small ones, and no amount of extra training fixes that: the best line through the origin still has an error of about 1.64. Adding one number, the bias, removes the error completely.
Pause and think: In the bias-free run, the learned slope is 2.818, not 2. Why would gradient descent choose a slope that is "wrong"?
Because it is the best it can do under the constraint. The model must pass through the origin, so to get close to prices that start at 3 it tilts the line upward. 2.818 is the least-squares slope for a line forced through (0, 0); it trades errors on small houses against errors on big ones.
With bias vs without bias
How much does a bias cost? Very little. A dense layer that maps 512 inputs to 256 neurons has 512 × 256 = 131,072 weights but only 256 biases, about 0.2% extra. That small price buys a lot of flexibility.
Bias in real networks
In frameworks, bias is on by default. A PyTorch nn.Linear(512, 256) layer creates a weight matrix of shape 256 × 512 and a bias vector of length 256; you turn the bias off with bias=False. Keras Dense layers have the matching use_bias argument.
Where people deliberately remove bias A convolution or linear layer followed directly by Batch Normalization usually has bias=False. BatchNorm subtracts the mean of each feature, which cancels any constant the previous layer added, and then adds its own learned shift (called β). Many modern large language models, including the Llama family, also use linear layers without bias inside their Transformer blocks; their designers found the bias terms added little. The exact choice varies by model, so check the architecture you are using.
- Classifiers: the bias of the final layer often ends up reflecting how common each class is. If 90% of emails are not spam, the "not spam" output gets a head start.
- Regression: the output bias often learns something close to the average target value, so the weights only need to explain deviations from that average.
- Hidden layers: each neuron uses its bias to pick its own threshold, so different ReLU neurons turn on at different points. Together they can build bent, piecewise shapes that a single straight line could not.
Common mistakes and when not to use bias
Mistake: thinking weights alone are enough A common belief is that a big enough network can learn any offset from the weights alone. In a network without biases, an all-zero input produces zero in every layer (for activations where f(0) = 0, like ReLU and tanh), and every decision boundary of the first layer passes through the origin. The network can sometimes work around this, but it is wasting capacity on a problem one cheap number solves.
- Do not apply weight decay blindly to biases. Weight decay (L2 regularisation) pulls parameters towards zero to prevent overfitting. Many training recipes exclude biases (and normalisation parameters) from weight decay because pulling the offset to zero rarely helps and can hurt.
- Do not initialise biases to large values. Zero (or a small constant) is the usual start. A large negative bias on a ReLU neuron can switch it off for every input, and a neuron that never fires gets no gradient: a "dead ReLU".
- Remove bias before a normalisation layer that subtracts the mean. It is cancelled anyway, so it only wastes memory and compute.
- Keep it when the data is not centred. If the target has a non-zero average (prices, temperatures, counts), the output layer almost certainly needs a bias.
Rule of thumb: keep the bias unless you can name the component that already provides the shift.
Worked example, step by step
So far we trusted the code to do the updates. Let us do two of them by hand, so we can see the weight and the bias move together. We use only two houses: size 1 costs 5, and size 3 costs 9. Both follow our rule price = 2·size + 3. We start at w = 0, b = 0 and use a learning rate η = 0.05.
One update, by hand
- Predict: With
w = 0andb = 0both predictions are 0. - Errors: Error = prediction − truth:
0 − 5 = −5and0 − 9 = −9. Both are negative, so we predict too low. The loss is the mean of the squares:(25 + 81) / 2 = 53. - Weight gradient:
∂L/∂w = 2·mean(error·size) = 2·((−5·1) + (−9·3)) / 2 = −32. The big house counts three times as much, because its input is 3. - Bias gradient:
∂L/∂b = 2·mean(error) = 2·(−5 − 9) / 2 = −14. No input appears here. Every house counts the same. - Update both:
w = 0 − 0.05·(−32) = 1.6andb = 0 − 0.05·(−14) = 0.7. New predictions:1.6·1 + 0.7 = 2.3and1.6·3 + 0.7 = 5.5. The loss falls from 53 to 9.77.
| Update | w | b | Predictions | Loss |
|---|---|---|---|---|
| Start | 0 | 0 | 0 and 0 | 53 |
| After 1 | 1.6 | 0.7 | 2.3 and 5.5 | 9.77 |
| After 2 | 2.26 | 1.01 | 3.27 and 7.79 | 2.23 |
Look at what happened after the second update: w is already 2.26, which is above the true slope 2, while b is only 1.01, far below the true 3. The weight races ahead because its gradient is multiplied by the inputs. It then has to come back down as the bias slowly rises and takes over the constant part. This is normal. While the bias is still too small, the weight covers for it, just like the bias-free model did with its slope of 2.818.
A quick way to spot a missing bias Plot the errors against the input. If small inputs are all predicted too low and large inputs all too high (or the other way round), the model is tilting a line to make up for an offset it cannot express.
Practice: try it yourself
Now we build a tiny spam detector with one sigmoid neuron. An email is spam when it has more than 4 links. We train the neuron twice, with and without a bias, and look at where each one puts its switching point.
practice_bias_threshold.py
import numpy as np
# Toy task: an email is spam (1) when it has more than 4 links.
links = np.array([0., 1., 2., 3., 5., 6., 7., 8.])
spam = np.array([0., 0., 0., 0., 1., 1., 1., 1.])
def sigmoid(z):
return 1 / (1 + np.exp(-z))
def train(use_bias, lr=0.1, steps=20000):
w, b = 0.0, 0.0
for _ in range(steps):
p = sigmoid(w * links + b) # predicted spam probability
err = p - spam # error signal for each email
w -= lr * np.mean(err * links) # weight: error times input
if use_bias:
b -= lr * np.mean(err) # bias: plain average error
return w, b
for use_bias in (False, True):
w, b = train(use_bias)
pred = (sigmoid(w * links + b) > 0.5).astype(int)
acc = np.mean(pred == spam)
edge = f"{-b / w:.2f}" if use_bias else "0.00 (pinned)"
print(f"bias={use_bias!s:5} w={w:.3f} b={b:.3f} switch point: links={edge}")
print(f" predictions {pred} accuracy={acc:.3f}")Output:
bias=False w=0.264 b=0.000 switch point: links=0.00 (pinned) predictions [0 1 1 1 1 1 1 1] accuracy=0.625 bias=True w=3.566 b=-14.011 switch point: links=3.93 predictions [0 0 0 0 1 1 1 1] accuracy=1.000
Now change it:
- Change the rule so spam starts above 6 links: set
spamto[0, 0, 0, 0, 0, 0, 1, 1]. Before running, predict where the switch point moves and whetherbbecomes more or less negative. - Subtract 4 from every value in
linksso the data is centred near zero. Predict the accuracy of the bias-free neuron now. Does it still need a bias? - Lower
stepsfrom 20000 to 200. Predict whether the neuron with bias already reaches accuracy 1.0, and remember the worked example: which parameter is the slow one?
Pause and think: In the run with bias, w is positive (3.566) and b is strongly negative (−14.011). Why does a spam detector need a negative bias here?
The neuron fires when w·links + b > 0. With a positive weight, any email with links would push z above zero. The negative bias is a hurdle the evidence must clear: it takes about 3.93 links (14.011 / 3.566) before w·links beats it. The bias encodes "assume not spam until there are enough links".
Pause and think: The bias-free neuron predicts class 0 for the email with 0 links. Did it learn that this email is safe?
No. With no bias and an input of 0, the pre-activation is exactly 0 and the sigmoid gives exactly 0.5 for any weight. Our rule "spam if p > 0.5" then says 0 by a tie, not by learning. The neuron has no way to move that output away from 0.5.
Key takeaways
- A bias is a learned constant added to a neuron's weighted sum: z = w·x + b.
- Weights set slope and sensitivity; the bias sets the offset and the point where the activation switches on (x = −b/w).
- Without bias, every line or boundary is pinned to the origin and zero input gives zero output.
- A bias is trained like a weight on an always-1 input; its gradient is the average error signal.
- Biases are on by default; drop them only when a following layer, such as BatchNorm, already provides the shift.
Key terms
- Bias: A learnable constant added to a neuron's weighted sum that shifts its output independently of the inputs.
- Weight: A learnable number that scales one input, controlling how strongly that input influences the neuron.
- Pre-activation: The value z = w·x + b computed before the activation function is applied.
- Activation function: A non-linear function such as ReLU or sigmoid applied to the pre-activation to produce the neuron's output.
- Decision boundary: The set of inputs where w·x + b = 0; the bias moves it away from the origin.
- Dead ReLU: A ReLU neuron whose pre-activation is negative for all inputs, so it always outputs zero and receives no gradient.
← 2.9 Contrastive Learning: Training by Comparison · 3.2 Gradient Descent: Rolling Downhill to the Optimum →