Modern AI Engineering

Lesson 4.6 · 30 min

The Transformer Architecture: Built on Attention

Every major chatbot, image captioner and code assistant is built from the same few repeated parts. What are they, and how does a sentence flow through them?

In short: The Transformer turns tokens into vectors, adds position information, and then refines those vectors through a stack of identical layers. Each layer has two parts: multi-head attention, where tokens exchange information, and a feed-forward network, which processes each token on its own, both wrapped in residual connections and layer normalization. The original design had an encoder that reads the input and a decoder that writes the output; today we also use encoder-only and decoder-only versions.

Why the Transformer was needed

Before 2017, the best translation systems were recurrent encoder–decoder models (LSTMs) with attention added on top. They worked, but recurrence forced them to process words one after another. Training was slow, and information from distant words faded.

The paper “Attention Is All You Need” (Vaswani et al., 2017) asked a bold question: what if we drop recurrence completely and build the whole model from attention plus simple per-token layers? The result, the Transformer, trained much faster on GPUs because every position is processed in parallel, and it set new quality records on English–German and English–French translation. It became the foundation for BERT, GPT and nearly every large model since.

Think of it like a newsroom meeting Each token is a reporter holding one word. In every round (layer), reporters first talk to each other to gather relevant context (attention), then each goes back to their desk and thinks alone to update their notes (feed-forward network). After many rounds, every reporter's notes describe their word in the light of the whole story.

Our running example is translation, the task the Transformer was invented for: English “I love cats” in, French “J'aime les chats” out.

The two halves of the architecture

The original Transformer has two stacks. The encoder reads the whole input sentence and produces one context-rich vector per input token. The decoder generates the output sentence one token at a time, looking both at what it has written so far and at the encoder's vectors.

Hyperparameters of the original base Transformer (2017)
SettingValueMeaning
Layers (N)6 encoder + 6 decoderHow many times the block repeats
d_model512Size of each token vector
Attention heads8Parallel attention patterns per layer (64 dimensions each)
d_ff2,048Hidden size of the feed-forward network
Parametersabout 65 millionTiny by today's standards

Tokenization, Embedding, and Positional Encoding

Tokenization cuts text into tokens and maps them to integer IDs (the original paper used a byte-pair encoding vocabulary shared between the two languages). Embedding looks up each ID in a learned table to get a d_model-sized vector, so “cats” becomes a list of 512 numbers that the model can refine.

Attention by itself treats its input as an unordered set: if we shuffle the tokens, each token gets the same attention result, just in a shuffled position. So we must inject order. Positional encoding adds a position-dependent vector to each embedding. The original paper used fixed sine and cosine waves of different frequencies:

Small numbers. For position 1 and the first pair of dimensions (i = 0): PE = sin(1) ≈ 0.84 and cos(1) ≈ 0.54. For position 0 they are sin(0) = 0 and cos(0) = 1. These small numbers are simply added to the token embedding. Many later models use learned position embeddings (GPT-2, BERT) or rotary position embeddings (RoPE), which rotate query and key vectors by an angle that depends on position (Llama and many others).

The Attention Mechanism and Multi-Head Attention

Attention lets each token gather information from other tokens. Each token vector is multiplied by three learned matrices to create a query (what am I looking for?), a key (what do I contain?) and a value (what will I share?). The query of one token is compared with the keys of all tokens by dot product; softmax turns the scores into weights; the output is the weighted sum of values.

Multi-head attention runs several attention operations in parallel, each with its own smaller Q, K, V projections (8 heads of 64 dimensions in the base model). Different heads can learn different relationships, such as “look at the previous word” or “look at the subject of the verb”. Their outputs are concatenated and mixed by one more matrix.

The original Transformer uses attention in three places: (1) encoder self-attention, where input tokens attend to all input tokens; (2) decoder masked self-attention, where each output token attends only to earlier output tokens; and (3) cross-attention, where decoder queries attend to the encoder's keys and values, letting the translation look at the source sentence.

Feed-Forward Networks, Residual Connections, and Layer Normalization

After attention mixes information between tokens, a position-wise feed-forward network (FFN) processes each token on its own, with the same weights at every position: expand to a bigger size, apply a non-linearity, shrink back.

Two helpers keep deep stacks trainable. A residual connection adds a sub-layer's input to its output: x + Sublayer(x). The sub-layer only has to learn a change to x, and gradients have a direct path back through the addition, which makes very deep networks stable. Layer normalization rescales each token vector to mean 0 and variance 1 (then applies a learned scale and shift), keeping numbers in a healthy range.

Post-norm vs pre-norm The 2017 paper normalised after the residual addition: LayerNorm(x + Sublayer(x)). Most modern LLMs normalise before the sub-layer: x + Sublayer(LayerNorm(x)), which trains more stably in very deep models. Many also replace LayerNorm with the cheaper RMSNorm.

How the Encoder and Decoder work

Translating “I love cats” step by step

  1. Encode once: The three English tokens pass through all encoder layers in parallel. Out come three vectors, each aware of the whole sentence.
  2. Start the decoder: The decoder receives only a start token. Masked self-attention has just that one token to look at.
  3. Cross-attend: The decoder's query asks the encoder vectors “what should come first?” and attends mostly to “I” and “love”.
  4. Predict: After the FFN and the final linear + softmax, the most likely first token is “J'” (or “J'aime”, depending on the tokenizer).
  5. Feed back and repeat: The new token is appended to the decoder input; the decoder runs again, producing “aime”, “les”, “chats” and finally an end token. The encoder output is reused at every step.

Pause and think: During translation, how many times does the encoder run, and how many times does the decoder run, for a 5-token French output (including the end token)?

The encoder runs once over the English sentence. The decoder runs once per generated token, so 5 times (with a KV cache, each run only processes the newest token). That asymmetry is why the encoder's output is computed once and reused.

How data flows through the entire architecture (code)

Below is one complete Transformer layer in numpy with tiny sizes (d_model = 8, 2 heads, d_ff = 32), following the original post-norm design. The weights are random, so the numbers are meaningless, but every shape and step is real.

one_transformer_layer.py

import numpy as np
rng = np.random.default_rng(42)
vocab, d_model, n_heads, d_ff = 10, 8, 2, 32
tokens = np.array([3, 7, 1, 4])            # token ids for a 4-token sentence
n, d_k = len(tokens), d_model // n_heads
# 1. Embedding lookup + sinusoidal positional encoding
E = rng.normal(size=(vocab, d_model)) * 0.5
pos = np.arange(n)[:, None]; i = np.arange(d_model)[None, :]
angle = pos / 10000 ** (2 * (i // 2) / d_model)
PE = np.where(i % 2 == 0, np.sin(angle), np.cos(angle))
x = E[tokens] + PE
print("embeddings + positions:", x.shape)
def layer_norm(v):
return (v - v.mean(-1, keepdims=True)) / (v.std(-1, keepdims=True) + 1e-5)
def softmax(s):
e = np.exp(s - s.max(-1, keepdims=True)); return e / e.sum(-1, keepdims=True)
# 2. Multi-head self-attention, then residual + LayerNorm (post-norm, as in 2017)
Wq, Wk, Wv, Wo = (rng.normal(size=(d_model, d_model)) * 0.3 for _ in range(4))
Q, K, V = (x @ W for W in (Wq, Wk, Wv))
heads = []
for h in range(n_heads):                   # each head uses its own slice
s = slice(h * d_k, (h + 1) * d_k)
w = softmax(Q[:, s] @ K[:, s].T / np.sqrt(d_k))
heads.append(w @ V[:, s])
print(f"head {h} attention of token 0:", np.round(w[0], 2))
attn = np.concatenate(heads, axis=-1) @ Wo
x = layer_norm(x + attn)
print("after attention + residual + norm:", x.shape)
# 3. Position-wise feed-forward network, then residual + LayerNorm
W1 = rng.normal(size=(d_model, d_ff)) * 0.3
W2 = rng.normal(size=(d_ff, d_model)) * 0.3
x = layer_norm(x + np.maximum(0, x @ W1) @ W2)
print("after feed-forward + residual + norm:", x.shape)
# 4. Output head: project to vocabulary scores (untrained, so random)
logits = x @ E.T                           # weight tying with the embedding
print("next-token probs for last position:", np.round(softmax(logits[-1]), 2))

Output:

embeddings + positions: (4, 8)
head 0 attention of token 0: [0.25 0.42 0.17 0.16]
head 1 attention of token 0: [0.24 0.33 0.21 0.22]
after attention + residual + norm: (4, 8)
after feed-forward + residual + norm: (4, 8)
next-token probs for last position: [0.27 0.04 0.03 0.03 0.08 0.03 0.09 0.35 0.   0.07]

The key observation: the shape never changes inside the stack. A layer takes n × d_model and returns n × d_model, which is why we can stack 6, 32 or 100 identical layers. Only the content of each vector gets richer.

The three variants of the Transformer

Why the Transformer is so powerful

  • Parallel training: all positions are processed at once, so huge datasets can be used on GPU clusters.
  • Direct long-range links: any token reaches any other in one attention step.
  • Scales smoothly: adding layers, width and data has kept improving results, which enabled today's large models.
  • General: the same blocks work for text, images (split into patches), audio and video.
  • Simple, repeated design: one layer type stacked many times is easy to optimise in hardware and software.

Limits we must know Self-attention compute and memory grow with the square of the sequence length, so long contexts are expensive. Generation is still one token at a time for decoder models. And a Transformer is only as good as its training data; the architecture alone does not guarantee truthfulness.

Pause and think: If we removed the positional encodings from a Transformer encoder, what would happen to “dog bites man” vs “man bites dog”?

Each token would get the same output vector in both sentences (just in a different order), because self-attention without position information cannot tell order. The model could not tell who bit whom.

Worked example, step by step

We said the base model has about 65 million parameters and that most of them sit in the feed-forward layers. Let us check both claims with plain multiplication. A weight matrix from size a to size b has a × b weights plus b biases. We use d_model = 512 and d_ff = 2,048.

Counting one encoder layer

  1. Attention: Four matrices of 512 × 512 (for queries, keys, values and the output mix), each with 512 biases: 4 × (262,144 + 512) = 1,050,624.
  2. Feed-forward, expand: 512 → 2,048: 512 × 2,048 + 2,048 = 1,050,624. One FFN matrix already matches the whole attention block.
  3. Feed-forward, shrink: 2,048 → 512: 2,048 × 512 + 512 = 1,049,088. FFN total: 2,099,712.
  4. Layer norms: Two norms, each with a scale and a shift of 512 numbers: 2 × 1,024 = 2,048.
  5. Add up: 1,050,624 + 2,099,712 + 2,048 = 3,152,384 per encoder layer. The FFN holds almost exactly two thirds of it.

A decoder layer has one more attention block (cross-attention) and one more norm: 2 × 1,050,624 + 2,099,712 + 3,072 = 4,204,032. Six encoder layers give about 18.9 million and six decoder layers about 25.2 million, so the two stacks hold about 44.1 million.

The rest is the embedding table. The paper used a shared vocabulary of about 37,000 tokens, and 37,000 × 512 is about 18.9 million. That brings our count to roughly 63 million, close to the reported figure of about 65 million. The small gap comes from details such as the exact vocabulary size.

A quick rule Every big matrix in a layer scales with d_model × d_model. So doubling the vector size roughly quadruples the parameters per layer, while adding a layer only adds one more copy.

Practice: try it yourself

We will test the claim that attention is order-blind. We run self-attention on “dog bites man” and on “man bites dog”, first without positional encoding and then with it, and compare the vector that comes out for “dog”.

practice_order_blind.py

import numpy as np
rng = np.random.default_rng(5)
n, d = 3, 4                                 # 3 tokens, 4 numbers each
E = rng.normal(size=(n, d))                 # embeddings for "dog", "bites", "man"
def self_attention(x):
scores = x @ x.T / np.sqrt(d)
w = np.exp(scores - scores.max(1, keepdims=True))
return (w / w.sum(1, keepdims=True)) @ x
def positions(n):
# Sinusoidal positional encoding, same formula as the lesson
pos = np.arange(n)[:, None]; i = np.arange(d)[None, :]
angle = pos / 10000 ** (2 * (i // 2) / d)
return np.where(i % 2 == 0, np.sin(angle), np.cos(angle))
order_a = [0, 1, 2]                         # dog bites man
order_b = [2, 1, 0]                         # man bites dog
# Without positions: compare the vector for "dog" in both sentences
dog_a = self_attention(E[order_a])[0]       # dog is first in sentence A
dog_b = self_attention(E[order_b])[2]       # dog is last in sentence B
print("no positions  : dog differs by", round(float(np.abs(dog_a - dog_b).sum()), 4))
# With positions added before attention
dog_a = self_attention(E[order_a] + positions(n))[0]
dog_b = self_attention(E[order_b] + positions(n))[2]
print("with positions: dog differs by", round(float(np.abs(dog_a - dog_b).sum()), 4))
print("PE for position 0:", np.round(positions(n)[0], 2))
print("PE for position 2:", np.round(positions(n)[2], 2))

Output:

no positions  : dog differs by 0.0
with positions: dog differs by 2.1852
PE for position 0: [0. 1. 0. 1.]
PE for position 2: [ 0.91 -0.42  0.02  1.  ]

Now change it:

  • Change order_b to [1, 0, 2] (“bites dog man”) and read “dog” from index 1 instead of 2. Predict the “no positions” difference before running.
  • In the two “with positions” lines, use 0.01 * positions(n). Predict: does the difference stay near 2.19, drop to exactly 0, or become small but not 0?
  • Change d from 4 to 8 and print positions(n)[2] again. Predict which of the 8 numbers will be closest to their position-0 values.

Pause and think: Position 0 has encoding [0, 1, 0, 1] and position 2 has [0.91, −0.42, 0.02, 1]. The first two numbers changed a lot, the last two hardly at all. Why is that useful and not a flaw?

The first pair is a fast wave and the last pair is a very slow wave. Fast waves tell nearby positions apart but repeat after a few tokens. Slow waves barely move between neighbours but keep changing over hundreds or thousands of positions, so they separate far-apart tokens. Together they give every position a unique pattern at both small and large scales.

Pause and think: Suppose we widen the base model from d_model = 512 to 1,024 and d_ff from 2,048 to 4,096, keeping 6 layers. Roughly how do the parameters in one encoder layer change, and why?

They grow about four times, to roughly 12.6 million. Every large matrix has both sides doubled (512 × 512 becomes 1,024 × 1,024, and 512 × 2,048 becomes 1,024 × 4,096), and doubling both sides multiplies the entries by four. Only the small bias and norm terms grow by two.

Key takeaways

  • The Transformer replaced recurrence with attention, enabling parallel training (Vaswani et al., 2017).
  • Input pipeline: tokenize → embed → add positional encoding (sinusoidal, learned or RoPE).
  • Each layer = multi-head attention (tokens talk) + feed-forward network (each token thinks), with residuals and layer norm.
  • The encoder reads the input once; the decoder writes token by token using masked self-attention and cross-attention.
  • Variants: encoder-only (BERT), decoder-only (GPT and most LLMs), encoder–decoder (T5, original Transformer).

Key terms

  • Encoder: The stack that reads the whole input and produces context-rich vectors for each input token.
  • Decoder: The stack that generates output tokens one at a time, using masked self-attention (and cross-attention in encoder–decoder models).
  • Positional encoding: Vectors added to token embeddings so the model can tell token order.
  • Multi-head attention: Several attention operations in parallel, each learning a different pattern, then combined.
  • Feed-forward network (FFN): A small two-layer network applied to each token independently inside every layer.
  • Residual connection: Adding a sub-layer's input to its output: x + Sublayer(x).
  • Layer normalization: Rescaling each token vector to a standard mean and variance to keep training stable.

← 4.5 RNNs vs Transformers: A Fundamental Architecture Shift · 4.7 Encoder vs Decoder: Two Sides of the Transformer →