Modern AI Engineering

Lesson 4.5 · 25 min

RNNs vs Transformers: A Fundamental Architecture Shift

Why did almost the whole field switch from RNNs to Transformers within a few years of 2017, and are RNN ideas really gone?

In short: Both RNNs and Transformers process sequences such as text. An RNN reads one token at a time and squeezes everything it has seen into a single memory vector, which makes it slow to train and forgetful over long distances. A Transformer uses self-attention to let every token look directly at every other token in parallel, which trains much faster on GPUs and handles long-range context well, at the cost of compute that grows with the square of the sequence length.

What both of them are for

Many kinds of data are sequences: ordered lists where position matters. A sentence is a sequence of tokens, speech is a sequence of audio frames, a stock chart is a sequence of prices. “Dog bites man” and “man bites dog” contain the same words but mean different things, so a model must understand both the items and their order.

Recurrent Neural Networks (RNNs) and Transformers are two neural-network designs for sequences. Both turn each input token into a vector that captures its meaning in context, and both can be used to classify a sequence, translate it, or generate a new one. They differ in how information moves between positions. Our running example: the sentence “The keys that the man left on the kitchen table are missing”. To choose “are” (not “is”), a model must connect it to “keys”, eight tokens earlier.

Think of it like two readers The RNN reader reads one word at a time through a narrow slot and keeps a short mental note, rewriting it after each word. By the end of a long paragraph, early details have faded. The Transformer reader sees the whole page at once and can glance from any word to any other word instantly, as often as needed.

What is an RNN?

An RNN keeps a hidden state: a vector that acts as the network's running memory. At each step it combines the current token with the previous hidden state to produce a new hidden state. The same weights are reused at every step, which is what “recurrent” means.

How an RNN reads “The keys … are”

  1. Start empty: h₀ is a vector of zeros: no memory yet.
  2. Read “The”: Combine x₁ with h₀ to get h₁.
  3. Read “keys”: Combine x₂ with h₁ to get h₂. Now the memory should note “plural subject”.
  4. Keep going: Each of the next tokens (that, the, man, left, on, the, kitchen, table) overwrites part of the memory.
  5. Predict at “are”: Only h₁₀ is available. Whether “plural” survived eight updates decides whether the model picks “are” or “is”.

Improved variants were invented to protect the memory. The LSTM (Long Short-Term Memory, Hochreiter and Schmidhuber, 1997) adds gates, small learned switches that decide what to keep, forget and output. The GRU (Gated Recurrent Unit, 2014) is a simpler gated version. Before 2017, LSTMs were the standard for translation, speech recognition and text generation.

The problem with an RNN

  • No parallelism across time: hₜ needs hₜ₋₁, so a 1,000-token sequence needs 1,000 strictly ordered steps. GPUs, which are fast because they do many operations at once, sit mostly idle. Training on huge datasets becomes slow.
  • Vanishing and exploding gradients: to learn, the error signal travels backwards through every step (backpropagation through time). At each step it is multiplied by similar factors. If they are below 1 the signal shrinks towards zero; above 1 it blows up. Long-range lessons are hard to learn.
  • Memory bottleneck: everything seen so far must fit in one fixed-size vector. Early details get overwritten.
  • Long path length: information from token 1 reaches token 100 only after passing through 99 updates.

Pause and think: An LSTM fixes vanishing gradients better than a plain RNN. Does it also fix the parallelism problem?

No. An LSTM still computes step t from the state at step t−1, so it remains sequential in time. Gates help memory and gradient flow, not parallel training.

What is a Transformer?

The Transformer was introduced in 2017 in the paper “Attention Is All You Need” by Vaswani and colleagues at Google. It removed recurrence completely. Its key operation is self-attention: every token builds a query, compares it with a key from every other token, and takes a weighted mix of their values. In one step, “are” can look straight back at “keys” and give it a high weight.

Because attention alone does not know word order, Transformers add positional information to each token (fixed sinusoidal patterns in the original paper; learned or rotary position encodings in many later models). Each layer then runs attention followed by a small feed-forward network, and many layers are stacked. All positions in a layer are computed at the same time with large matrix multiplications, exactly what GPUs are built for.

From recurrence to attention

  1. Simple RNN: Elman's recurrent network popularises the hidden-state loop for sequences.
  2. LSTM: Gated memory cells make longer dependencies learnable.
  3. Seq2seq and attention: Encoder–decoder LSTMs translate sentences; Bahdanau and colleagues add attention so the decoder can look back at all encoder states.
  4. Transformer: Attention without recurrence; faster to train and better at translation.
  5. BERT, GPT and LLMs: Pre-trained Transformers take over almost all language tasks, then vision and audio.
  6. Recurrence returns, differently: State-space models such as Mamba and RNN-style designs such as RWKV revisit linear-time sequence processing, often in hybrids with attention.

The key difference in one line

In one line An RNN passes information along the sequence one step at a time through a single memory; a Transformer lets every token look directly at every other token, all at once.

Everything else follows from that: speed of training (parallel vs sequential), long-range memory (direct link vs many hops), and cost (quadratic attention vs linear recurrence).

RNN vs Transformer side by side (code)

Let us measure the long-range problem directly. We build a tiny untrained RNN and a single attention step, then nudge the first token and check how much the last output changes, for sequences of length 5, 20 and 50.

rnn_vs_attention.py

import numpy as np
rng = np.random.default_rng(0)
d = 4                                   # numbers per token
Wx = rng.normal(size=(d, d)) * 0.5      # input -> hidden weights
Wh = rng.normal(size=(d, d)) * 0.5      # hidden -> hidden weights
def run_rnn(X):
h = np.zeros(d)
for x in X:                         # strictly one token after another
h = np.tanh(x @ Wx + h @ Wh)    # new memory = f(token, old memory)
return h
def attention(X):
scores = X @ X.T / np.sqrt(d)       # every pair of tokens in ONE matmul
w = np.exp(scores - scores.max(1, keepdims=True))
w /= w.sum(1, keepdims=True)        # softmax per row
return w @ X                        # each output mixes all tokens
# How much does the FIRST token still influence the LAST output?
# Nudge token 0 and measure how much the last output moves.
for n in [5, 20, 50]:
X = rng.normal(size=(n, d))
X2 = X.copy(); X2[0] += 1.0
rnn_effect = np.abs(run_rnn(X2) - run_rnn(X)).sum()
att_effect = np.abs(attention(X2)[-1] - attention(X)[-1]).sum()
print(f"n={n:2d}  RNN steps={n:2d}  RNN effect={rnn_effect:.1e}  "
f"attention effect={att_effect:.1e}")
print("RNN: sequential steps grow with n; attention: one parallel step")

Output:

n= 5  RNN steps= 5  RNN effect=4.8e-03  attention effect=3.7e-01
n=20  RNN steps=20  RNN effect=8.3e-07  attention effect=9.4e-02
n=50  RNN steps=50  RNN effect=0.0e+00  attention effect=8.8e-03
RNN: sequential steps grow with n; attention: one parallel step

In the RNN, the first token's influence collapses exponentially: about 5×10⁻³ after 5 steps, under 10⁻⁶ after 20, and below floating-point precision (printed as 0) after 50. In attention, the influence also shrinks, but only because this untrained attention spreads its weight over more tokens. It shrinks slowly, and a trained model can learn to put high weight on a distant token whenever it matters, because the link is direct. The RNN also needed n sequential steps; attention needed one matrix operation.

Let's tabulate the difference

When to use which one?

SituationBetter choiceWhy
Chatbot, summarisation, code generationTransformerBest quality; huge pre-trained models are available
Classifying or embedding textTransformer (encoder)Strong pre-trained encoders exist
Tiny microcontroller reading a sensor streamRNN / GRUVery small, constant memory, processes one reading at a time
Short time series with little dataEither; often a simple RNN or classic modelA big Transformer may overfit or be overkill
Extremely long sequences on a tight memory budgetConsider state-space or hybrid modelsLinear-time processing; an active research area

Common misconception “Transformers have no limits on context.” They have a maximum context length, and attention cost grows with the square of the length. Many efficiency tricks (sliding-window attention, KV-cache compression, hybrids with recurrent layers) exist precisely because of this.

Pause and think: Our team must process 1-million-step sensor logs on a small edge device with very little memory. Why might a recurrent or state-space model beat a standard Transformer here?

A standard Transformer would need attention over a million positions (n² growth) and a growing KV cache, which a small device cannot hold. A recurrent or state-space model carries a fixed-size state, so memory stays constant no matter how long the log is.

Going one level deeper

“Linear versus quadratic” is easy to say and easy to misjudge. Let us count. For a sequence of n tokens, one RNN layer does n updates, one after another. One attention layer computes a score for every pair of tokens, so n × n scores, all in one parallel step. If we stored those scores naively as 4-byte numbers, the table below shows the memory for a single head in a single layer.

Simple counts, not measurements. Real systems avoid storing the full score table, but the amount of work still grows the same way.
Tokens (n)RNN steps in a rowAttention scores (n²)Score table at 4 bytes each
1010100400 bytes
10010010,00040 KB
1,0001,0001,000,0004 MB
10,00010,000100,000,000400 MB
100,000100,00010,000,000,00040 GB

So which one is faster? It depends on which cost hurts more: total arithmetic, or waiting in line.

How to reason about speed

  1. Count total work: At n = 1,000 the attention layer does a million scores while the RNN does a thousand updates. Attention does far more arithmetic.
  2. Count the waiting: The RNN's thousand updates must run in order. Update 1,000 cannot start before update 999 ends. Attention's million scores have no order, so a GPU can do them side by side.
  3. Short and medium inputs: Here parallel hardware wins. The Transformer finishes a training pass sooner even though it does more arithmetic.
  4. Very long inputs: Multiply n by 10 and attention work grows by 100. At some length, memory runs out before patience does. That is where linear-time designs become attractive again.

Depth is different too In an RNN, the number of steps between the first and last token grows with n. In a Transformer, it is fixed by the number of layers, whatever the length. A short chain of steps is easier to train through, and that is a second reason gradients behave better.

Practice: try it yourself

We will build the smallest RNN possible, with one number as its memory, and watch a single early signal fade. We also track the gradient: how much the last memory would change if we nudged the first one.

practice_tiny_rnn.py

import math
def run_rnn(inputs, w_x, w_h):
# A one-number RNN: new memory = tanh(w_x * input + w_h * old memory)
h, trace, grad = 0.0, [], 1.0
for t, x in enumerate(inputs):
h = math.tanh(w_x * x + w_h * h)
trace.append(h)
if t > 0:
grad *= w_h * (1 - h * h)   # d(h_t)/d(h_t-1), chained step by step
return trace, grad
# The only signal is at step 1. Every later input is 0.
inputs = [1.0] + [0.0] * 9
for w_h in [0.5, 0.9, 1.5]:
trace, grad = run_rnn(inputs, w_x=1.0, w_h=w_h)
shown = " ".join(f"{h:.3f}" for h in trace[:4])
print(f"w_h={w_h}: memory {shown} ... last={trace[-1]:.4f}  grad={grad:.1e}")
# Attention has no chain: the last token reads the first one directly.
# With equal scores, 10 tokens share the weight evenly, whatever the distance.
n = len(inputs)
print(f"attention: token {n} reads token 1 in one hop, weight {1 / n:.2f}")

Output:

w_h=0.5: memory 0.762 0.363 0.180 0.090 ... last=0.0014  grad=1.6e-03
w_h=0.9: memory 0.762 0.595 0.490 0.414 ... last=0.1892  grad=1.0e-01
w_h=1.5: memory 0.762 0.815 0.840 0.851 ... last=0.8585  grad=3.5e-04
attention: token 10 reads token 1 in one hop, weight 0.10

Now change it:

  • Change [0.0] 9 to [0.0] 49. Predict first: what happens to last and grad for w_h=0.9?
  • Change the first input from 1.0 to -1.0. Predict the sign of last for w_h=1.5, and say what that tells us the memory is storing.
  • Add 1.0 to the list of w_h values. Predict: does the memory fade faster or slower than with 0.9, and does it ever reach exactly 0?

Pause and think: With w_h=1.5 the memory does not fade: it ends at 0.8585. Yet the gradient is 3.5e-04, smaller than for 0.5. How can the memory be strong while the learning signal is weak?

tanh is saturated. Near 0.86 the curve is flat, so a small nudge to an earlier memory barely changes the later ones: each step's factor 1.5 × (1 − h²) is well below 1. The RNN holds on to a value, but training cannot easily adjust how it got there. Remembering and being trainable are two different things.

Pause and think: The attention line prints a weight of 0.10 for token 1. That is small. Why is this still a better position than the RNN with w_h=0.5, whose memory ended at 0.0014?

The 0.10 is only the untrained starting point: equal scores spread weight evenly over 10 tokens. Training can raise that one score directly and move the weight close to 1, because the link is a single step. The RNN's 0.0014 is the result of nine shrinking steps in a row, and the same chain shrinks the gradient that would be needed to fix it.

Summary

RNNs and Transformers both model sequences. RNNs read in order and keep one memory vector, so they are sequential, struggle with long-range dependencies, and are slow to train at scale. Transformers replace recurrence with self-attention, giving every token a direct, parallel view of every other token, which made today's large language models possible, at the cost of quadratic attention. RNN-style ideas live on in efficient modern architectures for very long or streaming inputs.

Key takeaways

  • Both RNNs and Transformers model sequences; they differ in how tokens share information.
  • RNNs read step by step with one memory vector: sequential, forgetful over long ranges, slow to train at scale.
  • LSTMs and GRUs ease forgetting with gates but stay sequential.
  • Transformers use self-attention for direct, parallel token-to-token links, enabling large-scale training.
  • Transformers pay quadratic attention cost; recurrent and state-space ideas return for very long or streaming inputs.

Key terms

  • Sequence: An ordered list of items, such as tokens, where position carries meaning.
  • RNN: A network that processes a sequence one step at a time, updating a hidden state with shared weights.
  • Hidden state: The RNN's running memory vector, updated after each token.
  • LSTM: An RNN variant with gates that control what to keep, forget and output.
  • Vanishing gradient: When the learning signal shrinks towards zero as it travels back through many steps.
  • Self-attention: An operation where each token computes weights over all tokens and mixes their information.

← 4.4 Embeddings: Encoding Meaning as Vectors · 4.6 The Transformer Architecture: Built on Attention →