Modern AI Engineering

Lesson 4.14 · 27 min

Rotary Position Encoding: Position Without Fixed Lookup Tables

Attention on its own cannot tell "dog bites man" from "man bites dog". How do modern LLMs know word order, using nothing but a rotation?

In short: RoPE (Rotary Position Embedding) encodes position by rotating each query and key vector by an angle proportional to the token's position, using different rotation speeds for different pairs of dimensions. Because rotating both vectors and then taking a dot product only depends on the difference of their angles, the attention score automatically depends on the relative distance between tokens. It adds no parameters, keeps vector lengths unchanged, and is used by most open LLMs today.

The big picture

RoPE, short for Rotary Position Embedding, was introduced in 2021 in the RoFormer paper by Jianlin Su and colleagues. It is the way most open-weight LLMs (Llama, Mistral, Qwen, Gemma and others) tell their attention layers where each token sits in the sequence.

The idea in one sentence: before computing attention scores, rotate each query and key vector by an angle that grows with its position. A token at position 5 is rotated more than a token at position 2. When two rotated vectors are compared with a dot product, the result depends on how far apart the two angles are, which is exactly the distance between the tokens.

Think of it like clock hands Imagine each token carries a clock whose hand moves forward a fixed step for each position. To compare two tokens, you only look at the angle between their hands. Two tokens 3 positions apart always have the same angle between their hands, whether they are at 2 and 5 or at 1,000 and 1,003. That is RoPE in a picture: absolute rotations, relative comparisons.

Why a Transformer needs position information

Attention compares every query with every key using dot products and then takes a weighted average. Nothing in that recipe knows the order of tokens. If you shuffle the input tokens, each token gets exactly the same set of scores, just in a shuffled order. Mathematically, attention is permutation-equivariant: permute the inputs and the outputs are permuted the same way, but otherwise unchanged.

So without extra help, "dog bites man" and "man bites dog" would look like the same bag of words. Word order carries meaning, so we must inject position somehow. The question is how, and that is where approaches differ.

Pause and think: Pause and think: a causal mask already makes token 3 unable to see token 4. Does that give the model full position information?

Only partly. The mask tells a token which tokens are earlier, but within the visible past it still cannot tell "the word just before me" from "a word 50 tokens back" without some position signal. Models with a causal mask still use explicit position encodings like RoPE.

Older approaches and their problems

How position has been added to Transformers
ApproachHow it worksMain drawback
Sinusoidal absolute (original Transformer, 2017)Add a fixed pattern of sines and cosines for each position to the token embeddingPosition is mixed into the content vector; relative distance is only indirectly available
Learned absolute (BERT, GPT-2)Learn one vector per position index up to a maximum, add it to the embeddingNo vector exists beyond the maximum (e.g. 512 or 1,024), so the model cannot handle longer inputs
Relative position bias (T5 and others)Add a learned bias to each attention score based on the (bucketed) distance between tokensExtra parameters and extra work inside every attention score computation
ALiBi (2021; used in BLOOM, MPT)Subtract a penalty proportional to distance from each attention scoreA fixed linear bias that strongly favours nearby tokens

The wish list that motivated RoPE: (1) attention scores should depend on relative distance, since "two words back" means the same thing anywhere in a document; (2) no extra learned parameters; (3) it should not disturb the size of the content vectors; and (4) it should fit into ordinary, fast attention code.

The core idea behind RoPE

RoPE does not add anything to the token embeddings. Instead, it acts inside attention, right after the queries and keys are computed and before their dot product. It splits each query and key vector into pairs of numbers: (x₁, x₂), (x₃, x₄), and so on. Each pair is treated as a point on a 2D plane and rotated by an angle. The angle is position × frequency, and each pair has its own frequency.

  • Fast pairs rotate a lot per position (for example 1 radian per token). They are sensitive to small distances: is this the next word or the one after?
  • Slow pairs rotate very little per position. They change meaningfully only over hundreds or thousands of tokens, so they can encode long-range distance.
  • Values (V) are not rotated. Only Q and K, because only their dot product decides the attention weights.

The 2D rotation math

Rotating a 2D point (x, y) counter-clockwise by an angle φ is done with a rotation matrix:

Rotation matrices have two properties we need:

  • Rotations preserve length. ‖R(φ)·v‖ = ‖v‖. The content of a vector is not inflated or shrunk by its position.
  • Rotations compose by adding angles, and undoing one is rotating backwards. R(a)ᵀ = R(−a), and R(−a)·R(b) = R(b − a).

Example: rotate (1, 0) by 90°. cos 90° = 0 and sin 90° = 1, so we get (1·0 − 0·1, 1·1 + 0·0) = (0, 1). The point moved from pointing right to pointing up, and its length is still 1.

How RoPE is applied to Q and K

RoPE inside one attention layer

  1. Project: Compute the query q and key k for every token as usual (q = x·W_q, k = x·W_k).
  2. Pick frequencies: For a head of width d there are d/2 pairs. Pair i gets frequency θᵢ = base^(−2i/d), with base = 10,000 in the original paper. Pair 0 has θ = 1; later pairs get geometrically slower.
  3. Rotate by position: For a token at position m, rotate pair i of its query and key by the angle m·θᵢ.
  4. Score: Compute rotated-q · rotated-k / √dₖ as usual. The score now contains relative-position information.
  5. Continue: Softmax and the weighted sum of V are unchanged. V is not rotated.

In code, nobody builds the big rotation matrix. Because it is block-diagonal, rotating is just a few element-wise multiplications with precomputed cos and sin tables. Implementations differ in which dimensions are paired: the original paper pairs neighbours (x₀, x₁), (x₂, x₃), while many popular codebases pair the first half with the second half (x₀ with x_d/2). Both are valid as long as the same convention is used in training and inference.

Why the dot product captures relative position

This is the heart of RoPE. Take one pair, query at position m and key at position n:

The absolute positions m and n have vanished. Only their difference (n − m) remains. The same is true for every pair, so the full score depends on the content of q and k and on their relative distance, never on where in the document they happen to be. Move both tokens 1,000 positions later and the score is identical.

A small numeric example

Use one pair and a rotation of 30° per position. Let q = k = (1, 0), so before rotation they point the same way and their dot product is 1.

  • Query at position 1, key at position 3. q rotates 30° → (cos 30°, sin 30°) = (0.866, 0.5). k rotates 90° → (0, 1). Dot product = 0.866·0 + 0.5·1 = 0.5.
  • Query at position 4, key at position 6. q rotates 120° → (−0.5, 0.866). k rotates 180° → (−1, 0). Dot product = (−0.5)(−1) + 0.866·0 = 0.5.
  • Both pairs are 2 positions apart, so the angle between them is 60° both times, and the score is cos 60° = 0.5 both times.

Now the same check with a real 8-dimensional RoPE (4 pairs, base 10,000) and random vectors:

rope_demo.py

import numpy as np
rng = np.random.default_rng(0)
def rope(x, pos, base=10000.0):
"""Rotate each pair (x[2i], x[2i+1]) by angle pos * theta_i."""
d = x.shape[-1]
theta = base ** (-np.arange(0, d, 2) / d)    # one frequency per pair
ang = pos * theta
cos, sin = np.cos(ang), np.sin(ang)
x1, x2 = x[0::2], x[1::2]
out = np.empty_like(x)
out[0::2] = x1 * cos - x2 * sin
out[1::2] = x1 * sin + x2 * cos
return out
d = 8
q, k = rng.standard_normal(d), rng.standard_normal(d)
print("frequencies:", np.round(10000.0 ** (-np.arange(0, d, 2) / d), 4))
# Same distance (3) at different absolute positions -> same score
for m, n in [(5, 2), (50, 47), (1000, 997)]:
s = rope(q, m) @ rope(k, n)
print(f"query at {m:>4}, key at {n:>4} (distance {m-n}): score = {s:.4f}")
# Different distances -> different scores
for dist in [0, 1, 10, 100]:
print(f"distance {dist:>3}: score = {rope(q, dist) @ rope(k, 0):.4f}")
# Rotation never changes a vector's length
print("norm before/after:", round(np.linalg.norm(q), 4), round(np.linalg.norm(rope(q, 123)), 4))

Output:

frequencies: [1.    0.1   0.01  0.001]
query at    5, key at    2 (distance 3): score = -1.5865
query at   50, key at   47 (distance 3): score = -1.5865
query at 1000, key at  997 (distance 3): score = -1.5865
distance   0: score = -1.4680
distance   1: score = -1.6954
distance  10: score = -1.1246
distance 100: score = -0.3711
norm before/after: 1.8627 1.8627

Pause and think: In the output, the fastest pair rotates 1 radian per position. Roughly how many positions does it take to complete a full turn, and why do we also need the slow pairs?

A full turn is 2π ≈ 6.28 radians, so about 6.3 positions. After that the fast pair repeats, so on its own it cannot tell distance 1 from distance 7.3. Slow pairs change only over long distances and break that ambiguity, like the hour hand complementing the second hand.

Comparison, real-world use and pitfalls

Real-world use: stretching context windows RoPE is used by Llama, Mistral, Qwen, Gemma, GPT-NeoX and PaLM, among others. Because positions are just angles, a model trained on 4K tokens can be adapted to longer contexts by changing how angles are computed: position interpolation squeezes new positions into the trained angle range, while NTK-aware scaling and YaRN adjust frequencies unevenly so fast pairs keep local detail. Raising the base (as Llama 3 did, with 500,000) slows all rotations so long distances stay distinguishable. These methods usually need some fine-tuning on long text to work well.

Common mistakes Rotating V as well as Q and K (RoPE only touches Q and K). Mixing the two pairing conventions between a checkpoint and your inference code, which silently scrambles positions. Assuming a RoPE model works beyond its trained length without scaling: quality usually collapses. And forgetting to apply the correct absolute position to each new token when using a KV cache, since cached keys were rotated with their own positions.

Going one level deeper

There is a second way to write RoPE that makes the proof one line long. Treat each pair (a, b) as a single complex number a + bi. Rotating the pair by an angle φ is then just a multiplication by e^(iφ) = cos φ + i·sin φ. Quick check with the earlier example: (1 + 0i) × (cos 90° + i·sin 90°) = i, which is the point (0, 1).

The dot product has a complex form too. For two pairs written as complex numbers z and w, the dot product is the real part of z × w̄, where w̄ flips the sign of the imaginary part of w.

One pair, with q = (1, 2) and k = (3, −1)

  1. Write them as complex numbers: q = 1 + 2i and k = 3 − i, so k̄ = 3 + i.
  2. Multiply once: q × k̄ = (1 + 2i)(3 + i) = 3 + i + 6i + 2i² = 1 + 7i. The real part, 1, is the ordinary dot product: 1·3 + 2·(−1) = 1.
  3. Add the positions: Query at position m, key at position n: (q·e^(imθ)) × conj(k·e^(inθ)) = q × k̄ × e^(i(m−n)θ). The two absolute angles merge into one difference.
  4. Plug in numbers: With 30° per position, query at 1 and key at 3, the difference is −60°. e^(−i60°) = 0.5 − 0.866i.
  5. Take the real part: (1 + 7i)(0.5 − 0.866i) has real part 0.5 + 7 × 0.866 ≈ 6.56. That is the RoPE score for this pair at this distance.

Look at what happened to the 7. In a plain dot product the imaginary part of q × k̄ is thrown away. With RoPE, the score for one pair is A·cos(Δ) + B·sin(Δ), where A + Bi = q × k̄ and Δ is the angle for the distance. So each pair contributes a wave in distance, and the learned q and k set its height and its shift. That is how a head can learn to prefer “two tokens back” over “the same token”.

Wavelengths for the lesson's 8-dimensional example (base 10,000): 2π / θᵢ positions per full turn.
Pairθᵢ (radians per position)Full turn everyGood at telling apart
01≈ 6.3 positionsNeighbouring tokens
10.1≈ 62.8 positionsPositions within a sentence or two
20.01≈ 628 positionsParagraph-scale distances
30.001≈ 6,283 positionsDocument-scale distances

The slowest pairs matter for long inputs. If a model only ever trains on a few thousand tokens, a pair with a wavelength longer than that never completes a turn during training. Longer inputs then push it to angles it has not met before. This is one common explanation for why RoPE models degrade past their trained length, and why the scaling methods above work by squeezing or slowing the angles.

Practice: try it yourself

We will write RoPE for a single pair using Python's built-in complex numbers, with no numpy and no rotation matrix. We check that the score depends only on the distance, try the shortcut of rotating just one vector, and print the wavelength of each pair.

practice_rope_complex.py

import cmath
import math
def rotate(pair, pos, theta):
# Treat the pair (a, b) as the complex number a + bi.
# Multiplying by e^(i * angle) rotates it by that angle.
return complex(*pair) * cmath.exp(1j * pos * theta)
def dot(z1, z2):
# Dot product of two 2-D vectors written as complex numbers
return (z1 * z2.conjugate()).real
q, k = (1.0, 2.0), (3.0, -1.0)
theta = math.radians(30)          # 30 degrees per position
print("no rotation      :", round(dot(complex(*q), complex(*k)), 4))
for m, n in [(1, 3), (4, 6), (100, 102)]:
score = dot(rotate(q, m, theta), rotate(k, n, theta))
print(f"query at {m:3d}, key at {n:3d}: score = {score:.4f}")
# Shortcut: leave q alone and rotate k by the distance only
print("rotate k by 2 only:", round(dot(complex(*q), rotate(k, 2, theta)), 4))
# Wavelength: positions needed for one pair to turn a full circle
for i, th in enumerate([1.0, 0.1, 0.01, 0.001]):
print(f"pair {i}: theta = {th:<5} full turn every {2 * math.pi / th:7.1f} positions")

Output:

no rotation      : 1.0
query at   1, key at   3: score = 6.5622
query at   4, key at   6: score = 6.5622
query at 100, key at 102: score = 6.5622
rotate k by 2 only: 6.5622
pair 0: theta = 1.0   full turn every     6.3 positions
pair 1: theta = 0.1   full turn every    62.8 positions
pair 2: theta = 0.01  full turn every   628.3 positions
pair 3: theta = 0.001 full turn every  6283.2 positions

Now change it:

  • Swap the placements so the key comes before the query: use [(3, 1), (6, 4), (102, 100)]. Predict first: is the score still 6.5622?
  • Set q, k = (1.0, 0.0), (1.0, 0.0). Predict the score at distance 2 from the lesson's small numeric example.
  • Change the angle to math.radians(90). Using q × k̄ = 1 + 7i, predict the score for a key 2 positions after the query.

Pause and think: Without rotation the score of this pair is 1.0. At a distance of 2 it is 6.56. Rotation never changes a vector's length, so how can the score get bigger?

A dot product depends on lengths and on the angle between the vectors. q = (1, 2) and k = (3, −1) start almost at right angles, which is why the plain score is small. Turning one of them by 60° relative to the other brings them close to parallel, so the same lengths give a much larger score. A model can learn q and k that line up best at a chosen distance, and that is a head that prefers a certain offset.

Pause and think: Pair 3 needs about 6,283 positions for a full turn. A model is trained only on inputs of up to 2,048 tokens. Roughly what share of that pair's circle did training cover, and why does it matter?

About one third (2,048 / 6,283 ≈ 0.33). For that pair, any distance beyond 2,048 lands on an angle the model never saw during training, so its behaviour there is untested. That is why feeding a much longer input tends to hurt quality, and why context-extension methods rescale positions or frequencies to bring the angles back into the familiar range.

Key takeaways

  • Attention alone ignores order, so Transformers need a position signal.
  • RoPE rotates each pair of query/key dimensions by position × frequency, with fast and slow pairs.
  • In the dot product the absolute angles cancel, so scores depend only on relative distance.
  • It adds no parameters, preserves vector length, leaves V alone and fits standard attention kernels.
  • Going beyond the trained context needs scaling methods such as position interpolation, NTK-aware scaling or YaRN.

Key terms

  • RoPE: Rotary Position Embedding: encoding position by rotating query and key vectors.
  • Rotation matrix: A matrix that turns a vector by an angle without changing its length.
  • Frequency θᵢ: How many radians pair i rotates per position; base^(−2i/d).
  • Relative position: The distance between two tokens, as opposed to their index in the sequence.
  • Permutation-equivariant: Shuffling the inputs simply shuffles the outputs the same way, so order is invisible.
  • ALiBi: A position method that subtracts a distance-proportional penalty from attention scores.
  • Position interpolation: Extending context by squeezing new positions into the angle range seen during training.

← 4.13 Cross-Attention: Connecting Encoder Output to the Decoder · 4.15 Feed-Forward Networks: The Transformer's Memory Layer →