Modern AI Engineering

Lesson 3.7 · 22 min

RMSNorm: Simpler Normalization for Transformers

Most modern open LLMs quietly deleted half of LayerNorm, the mean subtraction and the bias, and trained just as well. Why does the simpler version work?

In short: RMSNorm normalizes a vector by dividing it by its root mean square, x / √(mean(x²) + ε), then multiplies by a learned scale γ. Unlike LayerNorm it does not subtract the mean and has no shift β. It keeps the property that matters most, control over the scale of activations, while being simpler and somewhat cheaper, which is why many modern LLMs such as the Llama family use it.

Why normalization is needed in deep networks

A large language model is a tall stack of identical blocks, often dozens of layers deep. Each block adds its output to a running vector called the residual stream. If the size of those vectors drifts, growing a little in one layer and a little more in the next, activations can explode or shrink towards zero, and gradients follow. Training then becomes unstable: the loss spikes, or the model needs a tiny learning rate to survive.

Normalization layers fix this by rescaling each token's vector to a standard size before it enters the next computation. In the previous lesson we met BatchNorm and LayerNorm. Transformers use per-token normalization because it does not depend on the batch. This lesson is about the leaner per-token normalizer that most recent LLMs chose: RMSNorm.

Think of it like a volume limiter A radio station runs every song through a limiter so nothing is painfully loud or inaudibly quiet. The limiter only adjusts the loudness; it does not shift the music's pitch. LayerNorm is like a limiter that also re-centres the pitch. RMSNorm is the plain limiter: it only fixes the volume, and it turns out that is the part deep networks really need.

A quick recap of Layer Normalization (LayerNorm)

For one token's hidden vector x with d numbers, LayerNorm does two things: re-centring (subtract the mean so the values average to 0) and re-scaling (divide by the standard deviation so they have spread 1). Then it applies a learned per-feature scale γ and shift β.

Computing this needs two passes over the vector (first the mean, then the variance around that mean) and stores 2d learned parameters per layer.

What RMSNorm is and how it works

RMSNorm (Root Mean Square Layer Normalization) was proposed by Biao Zhang and Rico Sennrich in 2019. Their idea: the benefit of LayerNorm comes mainly from re-scaling, not re-centring. So drop the mean, drop β, and scale the vector by its root mean square (RMS): the square root of the average of the squared values. The RMS is a measure of a vector's typical magnitude; it equals the vector's length divided by √d.

RMSNorm on one token vector

  1. Square: Square every value of the vector.
  2. Average: Take the mean of the squares.
  3. Root: Add ε and take the square root. This is RMS(x).
  4. Divide: Divide every value by RMS(x). The vector now has RMS 1, but its direction is unchanged.
  5. Scale: Multiply element-wise by the learned γ so the model can give each feature its preferred size.

Geometrically, RMSNorm keeps the direction of the vector and resets its length to √d (before γ). LayerNorm additionally moves the vector so its components average to zero, which changes the direction too.

The math behind RMSNorm with a concrete numeric example

Take a tiny token vector with d = 4: x = [1, 2, 3, 4], with γ = [1, 1, 1, 1] and ignore ε.

RMSNorm and LayerNorm on x = [1, 2, 3, 4], step by step
StepRMSNormLayerNorm
Statistic 1mean of squares = (1 + 4 + 9 + 16)/4 = 7.5mean μ = (1 + 2 + 3 + 4)/4 = 2.5
Statistic 2(none needed)variance = (2.25 + 0.25 + 0.25 + 2.25)/4 = 1.25
DivisorRMS = √7.5 ≈ 2.739std = √1.25 ≈ 1.118
Centre?NoSubtract 2.5: [−1.5, −0.5, 0.5, 1.5]
Result[0.365, 0.730, 1.095, 1.461][−1.342, −0.447, 0.447, 1.342]

Check the RMSNorm result: squares are 0.133 + 0.533 + 1.2 + 2.133 = 4.0, mean 1.0, so RMS = 1. The values keep their signs and their ratios (1 : 2 : 3 : 4). LayerNorm's output, in contrast, is centred at zero.

Pause and think: What is RMSNorm([3, −4]) with γ = 1 (ignore ε)?

Mean of squares = (9 + 16)/2 = 12.5, RMS = √12.5 ≈ 3.536. Result ≈ [0.849, −1.131]. Note that the signs survive, and the result has RMS 1.

LayerNorm vs RMSNorm: the key differences

Scale invariance means multiplying the input by any positive constant gives the same output: RMSNorm(10x) = RMSNorm(x). This is the property that keeps activations from blowing up layer after layer, and both normalizers have it. Shift invariance means adding a constant to every element gives the same output; only LayerNorm has this, and in practice LLMs do not seem to need it.

Why modern LLMs prefer RMSNorm

  • It works as well. In the original paper and in many later models, replacing LayerNorm with RMSNorm gave comparable quality. The re-scaling part is what stabilises training.
  • It is cheaper. One reduction instead of two, no mean subtraction, no bias add. The original paper reported noticeable speed-ups over LayerNorm, with the size depending on the model and implementation. The normalization is a small part of total compute, but it runs at every layer for every token, and in memory-bound GPU kernels every saved pass over the data helps.
  • Fewer parameters and simpler code. Half the learnable vectors, simpler fused kernels, and less to go wrong in mixed-precision training.
  • Design momentum. Once strong open models adopted it (T5 used an RMS-style norm; the Llama family made RMSNorm with pre-normalization a popular template), later models copied the recipe. Many current open-weight LLMs, including Mistral, Qwen and Gemma models, use RMSNorm; exact details such as ε and how γ is parameterised vary by model.

Honest caveat RMSNorm's advantage over LayerNorm is modest and practical, not dramatic. Some models still use LayerNorm successfully, and the speed difference depends heavily on hardware and kernel implementation. Treat "RMSNorm is the modern default" as a strong trend, not a law.

A code example

Here are both normalizers in numpy, applied to our example vector, plus checks of the invariance properties and the parameter counts for a 4,096-wide model (a typical hidden size for a model of about 7 billion parameters).

rmsnorm.py

import numpy as np
def layer_norm(x, gamma, beta, eps=1e-6):
mu = x.mean(-1, keepdims=True)                 # 1) mean
var = ((x - mu) ** 2).mean(-1, keepdims=True)  # 2) variance around the mean
return gamma * (x - mu) / np.sqrt(var + eps) + beta
def rms_norm(x, gamma, eps=1e-6):
rms = np.sqrt((x ** 2).mean(-1, keepdims=True) + eps)  # only one statistic
return gamma * x / rms                         # no mean subtraction, no beta
np.set_printoptions(precision=3, suppress=True)
x = np.array([1.0, 2.0, 3.0, 4.0])                 # one token's hidden vector, d = 4
d = len(x)
g, b = np.ones(d), np.zeros(d)
print("RMS of x      :", round(float(np.sqrt((x ** 2).mean())), 4))
print("RMSNorm(x)    :", rms_norm(x, g))
print("LayerNorm(x)  :", layer_norm(x, g, b))
# Invariance checks
print("RMSNorm(10x)  :", rms_norm(10 * x, g), "<- same: scale invariant")
print("RMSNorm(x+5)  :", rms_norm(x + 5, g), "<- changes: not shift invariant")
print("LayerNorm(x+5):", layer_norm(x + 5, g, b), "<- same: shift invariant")
# Learnable parameters per normalization layer for a 4096-wide model
d_model = 4096
print(f"params per layer  LayerNorm={2 * d_model}  RMSNorm={d_model}")

Output:

RMS of x      : 2.7386
RMSNorm(x)    : [0.365 0.73  1.095 1.461]
LayerNorm(x)  : [-1.342 -0.447  0.447  1.342]
RMSNorm(10x)  : [0.365 0.73  1.095 1.461] <- same: scale invariant
RMSNorm(x+5)  : [0.791 0.923 1.055 1.187] <- changes: not shift invariant
LayerNorm(x+5): [-1.342 -0.447  0.447  1.342] <- same: shift invariant
params per layer  LayerNorm=8192  RMSNorm=4096

In PyTorch, recent versions ship a built-in torch.nn.RMSNorm layer, and many model codebases also define their own few-line version that computes x torch.rsqrt(x.pow(2).mean(-1, keepdim=True) + eps) weight, often doing the computation in 32-bit floats for stability even when the model runs in 16-bit.

Where RMSNorm fits in a Transformer

In a modern decoder-only LLM, RMSNorm is used in the pre-norm arrangement: the input to each sub-layer is normalized, but the residual stream itself is left untouched. One block looks like this:

Some recent models also apply RMSNorm to the query and key vectors inside attention ("QK-norm") to stop attention scores growing too large; whether a given model does this varies.

Common mistakes Forgetting ε, which turns an all-zero vector into a division by zero. Computing the mean of squares in 16-bit floats, where squaring large values can overflow; implementations usually upcast to 32-bit. Loading weights from a model whose γ is stored as an offset (some models multiply by 1 + γ) into code that multiplies by γ directly, which silently breaks the model. And assuming RMSNorm subtracts the mean: it does not.

Pause and think: A teammate swaps LayerNorm for RMSNorm in an already-trained model without retraining. Will the outputs be the same?

Generally no. The trained weights expect centred outputs plus a learned β. RMSNorm skips both, so the activations differ unless every input already had mean 0 and β was 0. Switching normalizers is an architecture change that needs (re)training.

Common mistakes and how to spot them

RMSNorm is only a few lines of code, but three of those few lines can go wrong without any error message. Let us put real numbers on each failure so we can recognise it.

1. Squaring in 16-bit floats. A 16-bit float cannot hold a number larger than about 65,504. Take the vector x = [300, −200, 100, 50]. The first square is 300² = 90,000, which does not fit, so it becomes infinity. The mean of the squares is then infinity, the RMS is infinity, and every value divided by infinity is 0. The layer quietly outputs [0, 0, 0, 0]. In 32-bit floats the same vector gives the correct [1.589, −1.060, 0.530, 0.265]. This is why implementations compute the norm in 32-bit even when the rest of the model runs in 16-bit.

2. Misjudging ε. With ε = 10⁻⁶, take a very small vector x = [0.001, 0.002]. Its mean of squares is 2.5 × 10⁻⁶, about the same size as ε. Adding ε gives 3.5 × 10⁻⁶, so the divisor is 0.00187 instead of 0.00158, and the output is [0.535, 1.069] instead of [0.632, 1.265]. The output RMS is 0.845, not 1. So ε is not always invisible: for tiny activations it weakens the normalization. Without ε at all, an all-zero vector gives 0 / 0, which is NaN.

3. The wrong γ convention. Some models store the scale as an offset and multiply by 1 + γ, so a stored 0 means "no change". If we load those weights into code that multiplies by γ directly, a stored 0 wipes the vector out.

How each mistake shows up, and how to confirm it
SymptomLikely causeHow to confirm
Outputs of a norm layer are all zeros for some tokensSquares overflowed in 16-bitPrint the largest |x| entering the layer; redo the norm in 32-bit and compare
NaN appears right after a norm layerε missing, or placed outside the square root with a zero vectorFeed an all-zero vector through the layer alone
Output RMS clearly below 1 for small inputsε is large compared with mean(x²)Print mean(x²) next to ε
A loaded model produces nonsense from the first layerγ convention mismatch (γ vs 1 + γ)Look at the stored γ values: near 0 means offset style, near 1 means direct style

One test that catches most of these Take the output of the layer before γ is applied and compute its RMS. For any ordinary input it must be very close to 1. If it is 0, NaN or well below 1, one of the mistakes above is in play.

Practice: try it yourself

The lesson opened with a claim: without normalization, the size of a token vector drifts as it passes through many layers. Now we test it. We push one vector through 12 random layers, with and without RMSNorm in front of each layer, and print its RMS along the way.

practice_rmsnorm_stack.py

import numpy as np
def rms(x):
return float(np.sqrt((x ** 2).mean()))
def rms_norm(x, eps=1e-6):
return x / np.sqrt((x ** 2).mean() + eps)      # gamma = 1 for simplicity
def run_stack(gain, use_norm, d=64, layers=12):
rng = np.random.default_rng(0)                 # same weights for both runs
x = rng.normal(0, 1, d)                        # one token vector, RMS near 1
sizes = []
for _ in range(layers):
# A random layer that multiplies the typical size by about 'gain'
W = rng.normal(0, gain / np.sqrt(d), (d, d))
x = W @ (rms_norm(x) if use_norm else x)
sizes.append(rms(x))
return sizes
for gain in (0.5, 1.5):
for use_norm in (False, True):
s = run_stack(gain, use_norm)
label = "with RMSNorm" if use_norm else "no norm     "
print(f"gain={gain}  {label}  RMS after layer 1, 4, 8, 12: "
f"{s[0]:.3f}  {s[3]:.3f}  {s[7]:.3f}  {s[11]:.4f}")

Output:

gain=0.5  no norm       RMS after layer 1, 4, 8, 12: 0.476  0.049  0.003  0.0002
gain=0.5  with RMSNorm  RMS after layer 1, 4, 8, 12: 0.520  0.497  0.467  0.5339
gain=1.5  no norm       RMS after layer 1, 4, 8, 12: 1.428  3.981  19.930  118.0336
gain=1.5  with RMSNorm  RMS after layer 1, 4, 8, 12: 1.561  1.492  1.401  1.6017

Now change it:

  • Multiply the starting vector by 1000: x = 1000 * rng.normal(0, 1, d). Predict what happens to the two "with RMSNorm" lines and to the two "no norm" lines. Which property from the lesson are we testing?
  • Change layers=12 to layers=48 and print s[47]. Before running, estimate the "no norm" size for a gain of 1.5 (it grows about 1.5× per layer). What would this do to a 16-bit float?
  • Write a layer_norm(x) that subtracts the mean and divides by the standard deviation, and use it in place of rms_norm. Predict whether the stack stays stable, and whether the numbers differ much from the RMSNorm run.

Pause and think: With RMSNorm the printed RMS is about 0.5 or about 1.5, not 1. Did the normalization fail?

No. We measure the vector after the weight matrix, and the norm is applied before it. Each layer receives an input of RMS 1 and multiplies it by about the gain, so the output sits near the gain. What matters is that the next layer normalizes again, so the factor is applied once per layer instead of being multiplied layer after layer.

Pause and think: Could we skip normalization and simply initialise every layer with a gain of exactly 1.0?

It would look fine at the start, and careful initialisation does help. But the weights change at every training step. As soon as the layers drift to an average gain slightly above or below 1, the error compounds across the whole depth: even 1.1 per layer is about 3× after 12 layers and far more in a deep model. Normalization makes each layer's input size independent of what the layers below are doing, so training stays stable as the weights move.

Quick summary

  • Deep Transformers need per-token normalization to keep activations at a stable scale.
  • LayerNorm re-centres and re-scales, with learned γ and β.
  • RMSNorm only re-scales: x / √(mean(x²) + ε) · γ. No mean, no β.
  • Both are scale invariant; only LayerNorm is shift invariant, which LLMs seem not to need.
  • RMSNorm is simpler, has half the parameters and is somewhat cheaper; it is the common choice in modern open LLMs, placed before attention, before the FFN, and once at the end.

Key takeaways

  • RMSNorm(x) = γ · x / √(mean(x²) + ε): only re-scaling, no re-centring, no β.
  • It keeps the scale invariance that stabilises deep networks, and drops shift invariance.
  • It is simpler, has half the parameters and is somewhat faster than LayerNorm, with similar quality.
  • Modern LLMs place RMSNorm before attention and before the FFN in each block (pre-norm), plus one final norm.
  • You cannot swap LayerNorm for RMSNorm in a trained model without retraining.

Key terms

  • RMSNorm: A normalization layer that divides a vector by its root mean square and multiplies by a learned scale.
  • Root mean square (RMS): The square root of the average of squared values; a measure of a vector's typical magnitude.
  • Re-centring: Subtracting the mean so values average to zero; done by LayerNorm, skipped by RMSNorm.
  • Scale invariance: The output stays the same when the input is multiplied by a positive constant.
  • Pre-norm: A Transformer layout that normalizes the input of each sub-layer while leaving the residual stream untouched.
  • Residual stream: The running per-token vector that each Transformer block reads from and adds its output to.

← 3.6 Batch Norm vs Layer Norm: When to Use Each · 3.8 Recurrent Neural Networks: Processing Sequences in Order →