Transformer Architecture Explained Step by Step
By Modern AI Engineering · · 8 min read
The Transformer is the neural network design behind modern language models. It takes a sequence of tokens, turns each one into a vector, and passes the vectors through a stack of identical blocks. In each block, every token first gathers information from the other tokens through attention, and then a small network processes each token separately. At the top, the model predicts the next token.
This article follows one piece of text through the whole structure, one stage at a time. You need to know what a vector is. Nothing more.
What is a Transformer?
A Transformer is an architecture, a plan for how the parts of a network are arranged. Research on machine translation introduced it. Earlier models for text, called recurrent neural networks, read one token after another and carried a memory along. That was slow, and information from far back faded.
The Transformer dropped the step-by-step reading. It looks at all tokens at once and lets each one connect directly to any other. This solved the fading problem and allowed training to run in parallel on GPUs, which made much larger models practical.
The Transformer step by step
Here is the full path, from text in to prediction out, for a model that generates text.
- Tokenization. The text is cut into tokens, and each token is replaced by an ID number.
- Embedding. Each ID is looked up in a table and replaced by a vector.
- Position. Information about the order of the tokens is added, because attention alone has no sense of order.
- Self-attention. Each token looks at the other tokens and pulls in information from the relevant ones.
- Feed-forward network. Each token vector is processed separately by a small two-layer network.
- Residual connections and normalization. Around both sublayers, the input is added back to the output and the values are kept in a steady range.
- Repeat. Steps four to six form one block. The model stacks many blocks.
- Output. The final vector at the last position is turned into a score for every token in the vocabulary, and the scores become probabilities.
How does self-attention work?
Attention is the part that makes a Transformer a Transformer. Each token produces three vectors from its own vector: a query, a key and a value.
The query says what this token is looking for. The key says what this token offers. The value is the information it will pass on if chosen.
To update one token, the model compares its query with the key of every token. Each comparison gives a score. A function called softmax turns the scores into weights that add up to one. The token then takes a weighted mix of all the values. Tokens with high weights contribute a lot. Tokens with low weights contribute almost nothing.
Take the sentence: the animal did not cross the road because it was tired. When the model updates the token it, the query for it matches the key for animal strongly. So the vector for it absorbs information about the animal. The model has worked out what the pronoun refers to.
Before the softmax, the scores are divided by a fixed number tied to the vector size. This keeps them from growing too large, which would make training unstable. The whole operation is called scaled dot-product attention.
Why multi-head attention?
One round of attention can track one kind of relationship. Language has many at once: which noun a pronoun refers to, which verb belongs to which subject, which adjective describes which thing.
Multi-head attention runs several attention operations side by side. Each head has its own way of making queries, keys and values, so each can learn to look for something different. Their results are joined and mixed back into one vector per token.
What do the feed-forward layer, residuals and normalization do?
Attention moves information between tokens. The feed-forward network then works on each token alone. It widens the vector, applies an activation function and narrows it again. Much of what the model knows is thought to be stored in these layers.
A residual connection adds the input of a sublayer to its output. This gives the signal a direct path through a very deep stack, so the gradients used in training can reach the early layers.
Normalization keeps the numbers in a steady range from layer to layer. The original design used layer normalization. Many newer models use a lighter version called RMSNorm.
Encoder vs decoder: what is the difference?
The original Transformer had two halves. The encoder reads the whole input and builds a rich representation of it. The decoder writes the output one token at a time. In the decoder, a rule called causal masking stops each token from looking at tokens that come after it. A third kind of attention, cross-attention, lets the decoder look at the output of the encoder.
Later models often keep only one half.
| Type | How it reads | Typical use |
|---|---|---|
| Encoder-only | Every token sees the whole input, in both directions. | Classification, search, embeddings. |
| Decoder-only | Each token sees only the tokens before it. | Text generation. Most chat models are this type. |
| Encoder-decoder | The encoder reads the input, and the decoder writes with cross-attention to it. | Translation and other input-to-output tasks. |
How do Transformers know word order?
Attention treats its input as a set. Without extra information, the dog bit the man and the man bit the dog would look alike. So position has to be supplied.
The original design added a fixed pattern of numbers to each embedding, different for each position. Many current models use rotary position embedding, or RoPE. It rotates the query and key vectors by an angle that depends on position, so the attention score between two tokens reflects how far apart they are.
How to study the Transformer
Learn it in the order the data flows: tokens, embeddings, positions, attention, feed-forward, output. Work through one small attention example by hand, with vectors of two or three numbers. It takes an hour and removes most of the mystery.
Module 4 of the AI Engineering Bootcamp, Transformers and How They Think, has fifteen lessons that follow this path. They include separate lessons on the math of queries, keys and values, scaled dot-product attention, causal masking, multi-head attention, cross-attention and RoPE. The labs let you inspect attention weights yourself.
Frequently asked questions
What is a Transformer in AI?
It is a neural network design that processes all tokens of a sequence together and uses attention to let each token draw on the others. It is the base of modern language models.
What does the phrase attention is all you need mean?
It is the title of the research paper that introduced the Transformer. The point was that a model built on attention, with no recurrent layers, was enough for strong results on sequence tasks.
What is the difference between an encoder and a decoder?
An encoder reads the full input in both directions and builds a representation of it. A decoder generates output one token at a time and can only look at earlier tokens.
Are GPT models Transformers?
Yes. They are decoder-only Transformers. They use causal masking and are trained to predict the next token.
Why did Transformers replace RNNs?
They process tokens in parallel, so they train faster on large data, and attention links distant tokens directly, so long-range information does not fade.