Modern AI Engineering

Lesson 4.8 · 29 min

Self-Attention: How Tokens See One Another

In “The animal didn't cross the street because it was too tired”, how does a model work out that “it” means the animal and not the street?

In short: Self-attention lets every token in a sequence look at every other token and build a new, context-aware version of itself. Each token produces a query, a key and a value; queries are compared with keys by dot product, scaled by √dₖ, turned into weights with softmax, and used to average the values. Running several of these in parallel (multi-head attention) lets a Transformer capture many kinds of relationships at once.

What is Self Attention?

Self-attention is an operation that updates each token's vector by mixing in information from the other tokens of the same sequence, weighted by how relevant each one is. “Self” means the sequence attends to itself (in contrast with cross-attention, where one sequence attends to a different one). The output is one new vector per token, of the same size, that now carries context.

Before self-attention, the vector for “it” is the same in every sentence. After self-attention in a trained model, the vector for “it” in our sentence contains a large share of information from “animal”, so later layers can treat it as referring to the animal.

Think of it like a group discussion Each word walks into a room with a question (“who am I referring to?”), a name tag describing itself (“I am a noun, an animal”) and some notes to share. Each word compares its question with everyone's name tag, listens most to the best matches, and updates its own notes with a blend of what they shared. That is self-attention: question = query, name tag = key, notes = value.

Why do we need Self Attention?

Word meaning depends on context. “Bank” in “river bank” and “bank loan” are different. “It” can point to almost anything. A fixed embedding, one vector per word, cannot capture this, so a model needs a way to let words inform each other.

  • Resolve references: link “it”, “they”, “this” to what they mean.
  • Disambiguate: pick the right sense of “bank”, “bat”, “spring”.
  • Long-range links: connect a verb to its subject many words away (“The keys … are”).
  • Parallelism: recurrent networks pass context step by step; self-attention links all pairs at once, which trains fast on GPUs.

Earlier sequence models (RNNs, LSTMs) carried context through a single memory vector updated word by word, so distant information faded and training could not be parallelised across positions. Self-attention gives every pair of tokens a direct connection, with weights that depend on the actual content of the sentence.

Query, Key, and Value vectors

Each token vector x is multiplied by three learned weight matrices to produce three new vectors:

  • Query q = x · W_Q: what this token is looking for.
  • Key k = x · W_K: what this token offers to be matched against.
  • Value v = x · W_V: the information this token passes on if it is attended to.

Why three separate vectors instead of using x directly? Because the role of a word when searching differs from its role when being found and from what it contributes. A pronoun's query might look for “nouns that can be tired”, while a noun's key advertises “I am an animate noun”. The three matrices W_Q, W_K, W_V are learned during training, just like every other weight.

The library lookup A query is the search phrase we type; keys are the catalogue entries; values are the books themselves. Unlike a real library, attention does not return a single best book: it returns a blend of all books, weighted by how well each entry matches the search.

Stacking all tokens as rows, we get matrices Q = X·W_Q, K = X·W_K and V = X·W_V. With n tokens and key size dₖ, Q and K are n × dₖ.

Step-by-step working of Self Attention

The five steps

  1. Project: Multiply each token vector by W_Q, W_K and W_V to get its query, key and value.
  2. Score: For each token's query, take the dot product with every key. A larger dot product means a better match.
  3. Scale: Divide all scores by √dₖ. With dₖ = 64, divide by 8.
  4. Softmax: Apply softmax across each row so each token's weights over all tokens are positive and sum to 1. (In a decoder, future positions are masked to −∞ first.)
  5. Mix: Multiply the weights by the values: each token's output is a weighted average of all value vectors.

A simple example walk-through

Let us run every step with real numbers on three tokens, “cat sat mat”, each a 4-dimensional vector, projected to dₖ = 2. The weight matrices hold small integers (−1, 0, 1) so we can check the arithmetic by hand.

self_attention_by_hand.py

import numpy as np
np.set_printoptions(precision=2, suppress=True)
tokens = ["cat", "sat", "mat"]
X = np.array([[1.0, 0.0, 1.0, 0.0],     # cat
[0.0, 1.0, 0.0, 1.0],     # sat
[1.0, 1.0, 0.0, 0.0]])    # mat   (3 tokens x d_model=4)
rng = np.random.default_rng(0)
W_q = rng.integers(-1, 2, size=(4, 2)).astype(float)   # learned in real models
W_k = rng.integers(-1, 2, size=(4, 2)).astype(float)
W_v = rng.integers(-1, 2, size=(4, 2)).astype(float)
# Step 1: project every token into a query, a key and a value
Q, K, V = X @ W_q, X @ W_k, X @ W_v
for name, M in [("Q", Q), ("K", K), ("V", V)]:
print(name, "=", M.tolist())
# Step 2: score every query against every key (dot products)
scores = Q @ K.T
print("scores = Q @ K.T =", scores.tolist())
# Step 3: scale by sqrt(d_k) so large dimensions do not blow up softmax
d_k = K.shape[1]
scaled = scores / np.sqrt(d_k)
# Step 4: softmax each row into weights that sum to 1
weights = np.exp(scaled - scaled.max(1, keepdims=True))
weights /= weights.sum(1, keepdims=True)
print("weights =\n", weights)
# Step 5: each output is the weighted average of the value vectors
out = weights @ V
for t, o in zip(tokens, out):
print(f"output[{t}] = {o}")

Output:

Q = [[0.0, -1.0], [-1.0, -2.0], [1.0, -1.0]]
K = [[-1.0, 1.0], [1.0, 2.0], [-1.0, 2.0]]
V = [[-1.0, 1.0], [1.0, 0.0], [0.0, 1.0]]
scores = Q @ K.T = [[-1.0, -2.0, -2.0], [-1.0, -5.0, -3.0], [-2.0, -1.0, -3.0]]
weights =
[[0.5  0.25 0.25]
[0.77 0.05 0.19]
[0.28 0.58 0.14]]
output[cat] = [-0.26  0.75]
output[sat] = [-0.72  0.95]
output[mat] = [0.29 0.42]

Check “cat” by hand. Its query is [0, −1]. Dot with the keys: [0,−1]·[−1,1] = −1, [0,−1]·[1,2] = −2, [0,−1]·[−1,2] = −2. Scaled: −0.71, −1.41, −1.41. Softmax: e⁰ = 1 and e^(−0.71) ≈ 0.49 (after subtracting the max) give 1/1.99 ≈ 0.50 and 0.49/1.99 ≈ 0.25, 0.25. Output = 0.50·[−1, 1] + 0.25·[1, 0] + 0.25·[0, 1] ≈ [−0.25, 0.75]; the code prints −0.26 because it uses unrounded weights.

Pause and think: In the output, “sat” gives itself only 5% weight. Is that a bug?

No. Attention weights depend on how well the query matches each key, not on position. Here the query of “sat” ([−1, −2]) matches the key of “cat” ([−1, 1]) best, scoring −1, versus −5 for its own key ([1, 2]). Trained models often attend strongly away from the current token.

Why Self Attention works so well

  • Content-based, dynamic weights: weights are computed from the actual tokens in each input, so the same layer handles “it = animal” in one sentence and “it = street” in another.
  • Direct paths: any token can influence any other in a single step, so long-range relationships are easy to learn.
  • Parallel computation: all scores are one matrix multiplication, ideal for GPUs.
  • Stackable: outputs have the same shape as inputs, so layers can be stacked to build richer and richer representations.

Limits and common mistakes Self-attention is order-blind: without positional encodings, shuffling the tokens only shuffles the outputs. Its cost grows with the square of the sequence length (a 10× longer input means roughly 100× more attention scores). And attention weights are not a reliable explanation of why a model decided something; they show where information flowed in one layer, not the full reasoning.

Multi-Head Self Attention

One attention operation produces one set of weights per token, so it tends to capture one kind of relationship at a time. Multi-head attention runs h attention operations (heads) in parallel, each with its own W_Q, W_K, W_V of a smaller size. One head might follow the previous token, another might link pronouns to nouns, another might focus on punctuation.

Because each head works in a smaller space (64 instead of 512 dimensions), the total compute is about the same as one big head, but the model gains several independent views of the sentence. Many modern LLMs also share keys and values between groups of heads (grouped-query attention) to save memory during generation.

Where Self Attention is used

AreaHow self-attention is usedExamples
Large language modelsCausal (masked) self-attention in every decoder layerGPT series, Llama, most chat assistants
Text understandingBidirectional self-attention in encodersBERT, sentence-embedding models
VisionImages split into patches; patches attend to each otherVision Transformer (ViT, 2020)
SpeechAudio frames attend to each other in the encoderWhisper
Image generationAttention layers inside the denoising network, plus cross-attention to the text promptStable Diffusion and similar diffusion models
ScienceAttention over protein sequences and residue pairsAlphaFold 2

Back to “it was too tired” In a trained model, some heads in middle layers give the token “it” high weight on “animal”. If we change the sentence to “…because it was too wide”, the weight shifts towards “street”. Nothing about the layer changed; only the input did. That is the power of content-based weights.

Common mistakes and how to spot them

Self-attention is only a few lines of code, and that makes its bugs sneaky. The code still runs and still returns numbers of the right shape. Here are the slips we see most often when someone writes it by hand, with a quick test for each.

Typical self-attention bugs and how to catch them.
MistakeWhat we observeQuick test
Softmax over columns instead of rowsColumns sum to 1, rows do notSum each row of the weights; every sum must be 1
Forgot to divide by √dₖWeights are almost one-hot from the first training step; learning is slowPrint the largest weight per row; values near 1.00 everywhere are a warning
Mask applied after softmaxRows sum to less than 1Sum each row again after masking
Q and K swappedThe score matrix is transposed: token i's row holds how others look at itCheck a pair we understand: does the row of “it” point at “animal”?
No positional informationShuffled input gives the same outputs, shuffledFeed a sentence and its reverse; compare one token's output

Two of these are worth doing with numbers. Wrong axis: in our hand-computed weight matrix, each row sums to 1, but the first column sums to 0.50 + 0.77 + 0.28 = 1.55. If a test on the columns passes, the softmax ran in the wrong direction.

Masking too late: suppose one row of weights is [0.5, 0.3, 0.2] and the third token must be hidden. Zeroing it after softmax leaves [0.5, 0.3, 0], which sums to 0.8, so the output is no longer a proper average. Setting its score to −∞ before softmax gives [0.625, 0.375, 0]. The two visible tokens keep their 5 : 3 ratio and the row sums to 1 again.

A three-step sanity check for any attention code

  1. Check the shapes: With n tokens, the scores and weights must be n × n and the output must be n × the value size.
  2. Check the rows: Every row of weights is positive and sums to 1. Masked cells are exactly 0.
  3. Check a case we can predict: Make all scores equal. Every output must then be the plain average of the visible value vectors.

Practice: try it yourself

Random weights never show why attention works. So we will set the weights by hand. We build a three-token example where the query of “it” is designed to find an animate noun, and check that the output of “it” really picks up the features of “animal”.

practice_it_finds_animal.py

import numpy as np
np.set_printoptions(precision=2, suppress=True)
tokens = ["animal", "street", "it"]
# Hand-made features (illustrative): [is_animate, is_place, is_pronoun]
X = np.array([[1.0, 0.0, 0.0],
[0.0, 1.0, 0.0],
[0.0, 0.0, 1.0]])
# Hand-set projections instead of learned ones (d_k = 1 keeps it readable)
W_q = np.array([[0.0], [0.0], [1.5]])    # only the pronoun asks a question
W_k = np.array([[2.0], [-2.0], [0.0]])   # animate says "yes", place says "no"
W_v = np.eye(3)                          # values pass the features on unchanged
Q, K, V = X @ W_q, X @ W_k, X @ W_v
scores = Q @ K.T / np.sqrt(K.shape[1])   # 3 x 3: every query against every key
weights = np.exp(scores - scores.max(1, keepdims=True))
weights /= weights.sum(1, keepdims=True) # softmax per row
out = weights @ V                        # blend the values
print("scores for 'it':", scores[2])
for t, w, o in zip(tokens, weights, out):
print(f"{t:6s} weights={w}  output={o}")

Output:

scores for 'it': [ 3. -3.  0.]
animal weights=[0.33 0.33 0.33]  output=[0.33 0.33 0.33]
street weights=[0.33 0.33 0.33]  output=[0.33 0.33 0.33]
it     weights=[0.95 0.   0.05]  output=[0.95 0.   0.05]

Now change it:

  • Flip the key signs to [[-2.0], [2.0], [0.0]], as if the sentence ended “because it was too wide”. Predict the new weights for “it”.
  • Change 1.5 in W_q to 0.0. Predict the weights and output for “it” when it asks no question at all.
  • Set the pronoun's key to 2.0 (the last row of W_k). Predict how “it” now splits its weight between “animal” and itself.

Pause and think: The score of “it” for “street” is −3, yet its weight is 0.00, not a negative number. And “it” itself, with score 0, still gets 0.05. What does this tell us about how to read attention scores?

Only the differences between scores in a row matter. Softmax always returns positive weights that sum to 1, so a negative score just means “much less than the best match”, never negative attention. A score of 0 is not “no attention” either: it is 3 below the top score, which leaves a small share. Adding the same number to every score in a row would change nothing.

Pause and think: The output for “it” is [0.95, 0, 0.05]. Its own “pronoun” feature fell from 1 to 0.05. Has the model forgotten that “it” is a pronoun?

The attention output alone nearly has. But in a Transformer layer this output is added to the token's original vector through the residual connection, so “it” keeps its own features and gains the animate signal on top. Attention supplies the context; the residual path keeps the identity.

Key takeaways

  • Self-attention lets each token build a context-aware vector by mixing information from all tokens.
  • Each token gets a query (what it seeks), a key (what it offers) and a value (what it shares) from learned matrices.
  • Attention(Q, K, V) = softmax(Q·Kᵀ / √dₖ) · V; scaling by √dₖ keeps softmax trainable.
  • Multi-head attention runs several smaller attentions in parallel to capture different relationships.
  • Self-attention is order-blind (needs positions) and costs grow with the square of sequence length.

Key terms

  • Self-attention: An operation where every token in a sequence attends to every token of the same sequence to update its representation.
  • Query: A vector describing what a token is looking for, compared against keys.
  • Key: A vector describing what a token offers, matched against queries.
  • Value: A vector holding the information a token contributes to others' outputs.
  • Scaled dot-product attention: softmax(Q·Kᵀ / √dₖ) · V, the standard attention formula.
  • Attention head: One independent attention operation with its own Q, K, V projections.
  • Softmax: A function that turns a list of scores into positive weights summing to 1.

← 4.7 Encoder vs Decoder: Two Sides of the Transformer · 4.9 Attention Math: Queries, Keys, and Values Unpacked →