Modern AI Engineering

Lesson 4.7 · 25 min

Encoder vs Decoder: Two Sides of the Transformer

BERT and GPT are both Transformers, yet one is great at understanding and the other at writing. What single design choice makes the difference?

In short: An encoder reads the whole input at once, letting every token look both left and right, and outputs a context-rich vector per token: ideal for understanding. A decoder generates text one token at a time and uses a causal mask so each token sees only earlier tokens: ideal for writing. Transformers come in three types, encoder-only, decoder-only and encoder–decoder, and we choose among them based on whether the task is understanding, open-ended generation or mapping an input to an output.

What is a Transformer?

A Transformer is a neural-network architecture for sequences, introduced in 2017. It turns each token into a vector and refines those vectors through a stack of layers. In each layer, self-attention lets tokens gather information from other tokens, and a small feed-forward network processes each token on its own.

The original Transformer had two stacks: an encoder and a decoder. Later models kept one or both. This lesson is about what each stack does and why the difference matters when we pick a model. Our running example: a support team that wants to (a) detect the topic of each incoming ticket and (b) draft replies.

Think of it like a reader and a writer The encoder is a careful reader: it may read the whole letter, jump back and forth, and only then decide what each sentence means. The decoder is a writer: it writes one word at a time and can only look back at what it has already written, never at words it has not written yet.

A word about tokens

Both encoder and decoder work on tokens: small pieces of text such as words or word parts, each mapped to an integer ID by a tokenizer. “The bank of the river” might be five tokens. Each ID is looked up in an embedding table to get a vector, and positional information is added so the model knows token order.

From then on, everything happens to these per-token vectors. An encoder outputs one refined vector per input token. A decoder outputs, at its last position, a probability for every possible next token.

What is an Encoder?

An encoder takes the full input sequence and produces a contextual representation of every token: a vector that describes the token in the context of the whole input. Its self-attention is bidirectional: token 2 can look at token 1 (left) and token 5 (right) alike.

This matters for meaning. In “The bank of the river”, the word “bank” is only clear once we see “river”, which comes after it. An encoder lets “bank” attend to “river” directly, so its vector leans towards the riverside meaning.

Encoders are usually pre-trained with masked language modelling: hide some tokens and predict them from both sides. BERT (Google, 2018) masked about 15% of tokens, for example “The [MASK] of the river” → predict “bank”. Because the hidden token is replaced, looking right is not cheating.

  • Input: a complete sequence (sentence, document, query).
  • Output: one vector per token, often pooled into one vector for the whole input.
  • Good at: classification, sentiment, named-entity recognition, embeddings for search, reranking.
  • Examples: BERT, RoBERTa, many sentence-embedding models.

What is a Decoder?

A decoder generates a sequence one token at a time. At each step it reads all tokens produced so far and predicts a probability for the next one; the chosen token is appended and the process repeats. This is called autoregressive generation.

Its self-attention is causal (also called masked): token t may attend only to tokens 1…t. A causal mask sets every “look at the future” score to −∞ before softmax, so those weights become 0. Without it, training would let each position peek at the very token it must predict.

Decoders are pre-trained with next-token prediction (causal language modelling): given “The bank of the”, predict “river”. In an encoder–decoder model, each decoder layer has an extra cross-attention sub-layer whose queries come from the decoder and whose keys and values come from the encoder output, so the writer can consult the reader's notes.

  • Input: the tokens generated so far (plus the prompt).
  • Output: a probability distribution over the next token.
  • Good at: open-ended generation: chat, writing, code, reasoning in text.
  • Examples: the GPT series, Llama, Mistral and most chat LLMs.

The one big difference

In one line The encoder lets every token see the whole sequence (bidirectional attention); the decoder lets each token see only itself and earlier tokens (causal attention). Everything else follows from this mask.

Let us see it in numbers. We compute attention weights for “The bank of the river” twice: once with no mask (encoder) and once with a causal mask (decoder). The token vectors are random, so the exact values are illustrative, but the masking pattern is exactly what real models do.

encoder_vs_decoder_mask.py

import numpy as np
tokens = ["The", "bank", "of", "the", "river"]
n = len(tokens)
rng = np.random.default_rng(1)
X = rng.normal(size=(n, 4))                     # toy token vectors
scores = X @ X.T / 2.0                          # raw attention scores (√4 = 2)
def softmax_rows(s):
e = np.exp(s - s.max(1, keepdims=True))
return e / e.sum(1, keepdims=True)
# Encoder: bidirectional, every token sees every token
enc = softmax_rows(scores)
# Decoder: causal mask, token i may only see tokens 0..i
mask = np.triu(np.ones((n, n)), k=1).astype(bool)   # True above the diagonal
dec = softmax_rows(np.where(mask, -np.inf, scores))
for name, w in [("ENCODER", enc), ("DECODER", dec)]:
print(name)
for i, t in enumerate(tokens):
row = " ".join(f"{v:4.2f}" for v in w[i])
print(f"  {t:5s} -> {row}")
# Can 'bank' (position 1) use 'river' (position 4) to pick its meaning?
print("bank->river  encoder:", round(enc[1, 4], 2), " decoder:", round(dec[1, 4], 2))

Output:

ENCODER
The   -> 0.54 0.13 0.12 0.08 0.14
bank  -> 0.13 0.34 0.22 0.14 0.16
of    -> 0.15 0.27 0.24 0.18 0.16
the   -> 0.10 0.17 0.19 0.33 0.21
river -> 0.17 0.20 0.16 0.21 0.27
DECODER
The   -> 1.00 0.00 0.00 0.00 0.00
bank  -> 0.28 0.72 0.00 0.00 0.00
of    -> 0.23 0.41 0.36 0.00 0.00
the   -> 0.12 0.22 0.24 0.42 0.00
river -> 0.17 0.20 0.16 0.21 0.27
bank->river  encoder: 0.16  decoder: 0.0

Pause and think: In the output, why is the row for “river” identical in the encoder and decoder matrices, while the row for “The” differs completely?

“river” is the last token, so the causal mask hides nothing from it: it sees all five tokens either way. “The” is first, so in the decoder it can only see itself (weight 1.00), while in the encoder it spreads attention over all five tokens.

The three types of Transformers

  • Encoder-only (BERT family): understanding tasks; outputs vectors, not free text.
  • Decoder-only (GPT family, Llama, most chat LLMs): generation; also handles understanding by phrasing it as text (“Topic of this ticket: …”).
  • Encoder–decoder (original Transformer, T5, BART, Whisper for speech-to-text): input-to-output tasks like translation and summarisation.

Choosing a type for a new task

  1. Is the output a label, score or vector?: If yes, an encoder-only model is usually the efficient choice.
  2. Is the output free-form text with no fixed source?: Chat, creative writing, code generation: use a decoder-only model.
  3. Is there a clear input that maps to a different output?: Translation, summarisation, speech to text: an encoder–decoder fits well, though large decoder-only models also do these tasks well.
  4. Check constraints: Latency, cost and hardware: a small encoder can classify thousands of tickets per second on a CPU, while a large decoder is slower and pricier per call.

Let's tabulate the difference

When to use which one?

TaskGood choiceWhy
Route tickets to billing, shipping or returnsEncoder-onlyLabel output; fast and cheap at high volume
Embed help articles for semantic searchEncoder-only (embedding model)Bidirectional context gives strong vectors
Draft a reply to a customerDecoder-onlyOpen-ended text generation
Translate help articles into SpanishEncoder–decoder or a large decoder-only LLMClear input→output mapping
Transcribe a support phone callEncoder–decoder (e.g. Whisper)Audio encoder, text decoder
General assistant that does all of the aboveDecoder-only LLMFlexible, but larger and costlier per request

Common misconception “Decoder-only models can't understand text because they only look left.” They understand very well: by the time a decoder reaches the end of the prompt, the last positions have seen everything before them. The real trade-off is efficiency and fit: for pure classification or embeddings at scale, a small encoder is often cheaper and just as accurate.

Pause and think: We need to score 2 million product reviews per day as positive or negative on modest hardware. Which type would we try first and why?

An encoder-only model (fine-tuned for sentiment). It produces a label in a single forward pass with no token-by-token generation, so it is far cheaper and faster than prompting a large decoder for each review.

Worked example, step by step

The mask does more than decide who sees whom. It also decides what a model can practise on. Let us take one sentence, “The bank of the river”, and write out the training examples each type gets from it in a single pass.

Decoder training: next-token prediction. One pass over 5 tokens gives 4 questions.
PositionWhat the model may seeWhat it must predict
1Thebank
2The bankof
3The bank ofthe
4The bank of theriver

All four questions are answered in the same forward pass. The causal mask makes this honest: position 2 cannot see “of”, so it has to guess it.

Encoder training: masked-token prediction. Masking about 15% of 5 tokens hides roughly one token.
Input the model seesHidden tokenContext it may use
The [MASK] of the riverbankBoth sides: “The” on the left, “of the river” on the right

What the two tables tell us

  1. Count the questions: The decoder practises on every position. The encoder practises only on the hidden tokens, here one out of five. Per pass, the decoder gets more questions from the same text.
  2. Compare the clues: The encoder's single question comes with richer clues. It sees “river” while guessing “bank”. The decoder at position 1 must guess “bank” from “The” alone.
  3. Look at what is learned: Guessing the next token is the same skill as writing. Filling a gap is the skill of reading closely. Each training task builds the skill the model is later used for.
  4. Spot the mismatch: An encoder never practised continuing text, so asking it to write a paragraph does not work well. A decoder never saw the right-hand side during training, so its vector for an early token cannot reflect later words.

A practical consequence When we take a single vector from a decoder to represent a whole text, the last token is the natural choice: it is the only position that has seen everything. With an encoder, any position (or the average of all) has seen the full input.

Practice: try it yourself

We will run an experiment the weight tables cannot show. We replace the last word of the sentence and measure how much each token's output vector moves, once with encoder-style attention and once with decoder-style attention.

practice_who_is_affected.py

import numpy as np
rng = np.random.default_rng(3)
tokens = ["The", "bank", "of", "the", "river"]
X = rng.normal(size=(5, 4))             # toy token vectors
def attend(X, causal):
scores = X @ X.T / np.sqrt(X.shape[1])
if causal:                          # hide every position to the right
future = np.triu(np.ones_like(scores), k=1).astype(bool)
scores = np.where(future, -np.inf, scores)
w = np.exp(scores - scores.max(1, keepdims=True))
return (w / w.sum(1, keepdims=True)) @ X
# Swap the last word for a different one: "river" -> some other word
X2 = X.copy()
X2[4] = rng.normal(size=4)
for name, causal in [("encoder", False), ("decoder", True)]:
moved = np.abs(attend(X2, causal) - attend(X, causal)).sum(axis=1)
print(name)
for tok, m in zip(tokens, moved):
print(f"  {tok:5s} output moved by {m:.2f}")

Output:

encoder
The   output moved by 0.05
bank  output moved by 0.09
of    output moved by 0.11
the   output moved by 0.21
river output moved by 5.31
decoder
The   output moved by 0.00
bank  output moved by 0.00
of    output moved by 0.00
the   output moved by 0.00
river output moved by 5.31

Now change it:

  • Replace the first token instead: change X2[4] to X2[0]. Predict first: which decoder rows move now?
  • Replace the middle token: use X2[2]. Predict exactly which decoder rows stay at 0.00.
  • Change k=1 to k=0 in np.triu, so each token is also hidden from itself. Predict what the first decoder row prints, and why.

Pause and think: In the decoder, changing the last token left the first four outputs at exactly 0.00. Why does this property let a decoder reuse earlier work when it appends a new token, and why can an encoder not do the same?

In a decoder, earlier positions never look right, so adding or changing a later token cannot alter their results. They can be computed once and kept. In an encoder every token attends to every other, so a new token changes all outputs (all five rows moved in our run) and everything has to be recomputed.

Pause and think: In the encoder run, “The” moved by only 0.05 while “river” moved by 5.31. Does the small number mean encoders make little use of words to the right?

No. These are random, untrained vectors, so “The” happened to give “river” a small weight. The important fact is that the number is not zero: the path exists. Training can make that path strong where it helps, for example from “bank” to “river”. In the decoder the path is cut, so no amount of training can use it.

Summary

Encoders and decoders share the same building blocks (embeddings, attention, feed-forward layers, residuals, normalization) but differ in one mask. Encoders attend in both directions and output vectors, which makes them great readers for classification and embeddings. Decoders attend only to the past and output next-token probabilities, which makes them writers for generation. Encoder–decoder models combine both through cross-attention for input-to-output tasks. Today, decoder-only models power most chat LLMs, while encoders remain the efficient choice for understanding at scale.

Key takeaways

  • Encoder: bidirectional attention, outputs one context vector per token; great for understanding.
  • Decoder: causal attention, generates one token at a time; great for writing.
  • The causal mask (future scores set to −∞) is the key difference.
  • Three types: encoder-only (BERT), decoder-only (GPT, most chat LLMs), encoder–decoder (T5, Whisper).
  • Choose by task: labels and embeddings → encoder; open-ended text → decoder; input→output mapping → encoder–decoder.

Key terms

  • Encoder: A Transformer stack that reads the whole input bidirectionally and outputs a contextual vector per token.
  • Decoder: A Transformer stack that generates tokens one at a time using causal self-attention.
  • Bidirectional attention: Attention where each token can attend to tokens on both its left and right.
  • Causal mask: A mask that blocks attention to future tokens by setting their scores to −∞ before softmax.
  • Cross-attention: Attention where decoder queries look at encoder keys and values.
  • Masked language modelling: Pre-training by hiding some tokens and predicting them from the surrounding context.

← 4.6 The Transformer Architecture: Built on Attention · 4.8 Self-Attention: How Tokens See One Another →