Modern AI Engineering

Lesson 6.5 · 24 min

Attention Sinks: The Hidden Cost of Extended Context

Why does a chatbot suddenly start writing gibberish the moment we delete the first few, seemingly useless, tokens of a long conversation?

In short: Trained LLMs dump a large share of their attention onto the first few tokens, even when those tokens mean nothing; these tokens are called attention sinks. If a streaming system evicts them to save memory, the attention distribution shifts and the model breaks down. Keeping a handful of sink tokens plus a sliding window of recent tokens (the StreamingLLM recipe), or giving the model a dedicated learned sink, lets it run stably over very long streams.

What is a Large Language Model?

A Large Language Model (LLM) is a neural network that reads a sequence of tokens (pieces of words) and predicts the next one. It generates text by repeating that prediction: write a token, add it to the input, predict again. Nearly all modern LLMs are decoder-only Transformers, stacks of layers built around attention.

Our running example: a voice assistant that stays on all day in a car, listening and replying. Its conversation never really ends, so the token stream grows to hundreds of thousands of tokens. How do we keep it running without running out of memory or losing its mind?

What is attention?

In attention, the current token builds a query and compares it with the key of each earlier token. The scores go through softmax, which turns them into weights that are positive and must add up to exactly 1. The output is the weighted average of the earlier tokens' values.

That "must add up to 1" rule is the seed of this whole lesson. A head cannot say "nothing here is relevant, give me zero". It has to put its weight somewhere.

The problem of streaming with long conversations

During generation the model keeps keys and values for every past token in the KV cache. For an endless stream this cache grows without limit: memory runs out, and every new token gets slower because it must read more cache. Also, models are trained on a maximum length (say 4K or 8K tokens), and quality often collapses beyond it.

Think of it like a boat's ballast A sailing boat carries heavy ballast low in the hull. It does nothing useful by itself, but it keeps the boat balanced. Throw it overboard to save weight and the boat capsizes. Attention sinks are the model's ballast: they hold weight that has nowhere better to go.

The naive fix and why it fails

The obvious fix is window attention: keep only the most recent W tokens in the cache and evict the oldest ones. Memory becomes fixed and each step stays fast.

The StreamingLLM paper (Xiao et al., 2023, Efficient Streaming Language Models with Attention Sinks) tested exactly this on models such as Llama 2. As long as the text fit in the window, things were fine. But as soon as the very first tokens were evicted, perplexity (a measure of how surprised the model is by text; lower is better) shot up and the output fell apart. Losing a handful of tokens out of thousands should not matter, yet it was catastrophic.

The other baseline, recomputing the window from scratch for each new token, kept quality but was far too slow, because it recomputes the keys and values of the whole window every step.

What is an attention sink?

When researchers inspected attention maps, they saw that in most layers and heads (beyond the first couple of layers), a large share of attention went to the first token of the sequence, often far more than to any meaningful word. This happened even when the first token was just a start-of-text marker or a newline.

An attention sink is a token that soaks up attention weight without contributing much information. The head uses it as a "no-op": when nothing in the context is relevant, it parks weight on the sink, whose value vector tends to be small, so the output changes little.

Why the first tokens become a sink

  • Softmax forces a decision. Weights must sum to 1, so a head with nothing to look at needs a safe place to put them.
  • The first token is always visible. Under the causal mask, every later token can see token 0, in every training example. No other position is guaranteed to be present for all queries. So the model learns to use it as a universal dumping ground.
  • It is about position, not meaning. The StreamingLLM authors replaced the first four tokens with plain newline tokens and the effect largely remained. What matters is being at the start.
  • Several sink tokens help. Models trained without a special start token spread the sink over the first few tokens, which is why StreamingLLM keeps about 4 initial tokens.

Pause and think: Why can a token in the middle of the text not serve as a reliable sink during training?

Because under the causal mask, tokens before it cannot see it. Only the very first tokens are visible to every query in every training sequence, so they are the only consistent place to park attention.

A step-by-step numeric walkthrough

Take one query that scores 10 cached tokens. Token 0 is the sink: score 4.0, value near 0. The rest have scores between 0.1 and 0.6. Watch what happens to the softmax when we evict the sink.

Evicting the sink, by the numbers

  1. Full cache: exp(4.0) ≈ 54.6 dominates; the other nine exps add to about 13. The sink gets about 0.81 of the weight, each other token at most 0.03. Output ≈ +0.09.
  2. Evict to a window of 4: Only tokens 6–9 remain, with scores 0.2–0.6. Their weights must still sum to 1, so each jumps to roughly 0.20–0.30, about ten times larger than before.
  3. See the damage: The output becomes ≈ +0.55, completely different from +0.09. Every later layer now receives inputs unlike anything seen in training, and errors compound.
  4. Keep the sink: With token 0 plus tokens 6–9, the sink again absorbs most of the weight (~0.90), recent tokens stay small, and the output ≈ +0.06 is close to the original.

attention_sink_demo.py

import numpy as np
def softmax(x):
e = np.exp(x - x.max())
return e / e.sum()
# Scores of the current query against 10 cached tokens (illustrative values).
# Token 0 is the sink: a large score, but its value vector is near zero.
scores = np.array([4.0, 0.5, 0.2, 0.1, 0.3, 0.4, 0.2, 0.6, 0.3, 0.5])
values = np.array([0.0, 1.0, -1.0, 0.5, 2.0, -0.5, 1.5, 0.8, -1.2, 1.0])
def show(name, keep):
w = softmax(scores[keep])
out = w @ values[keep]
print(f"{name:22s} sink weight={w[0] if keep[0] == 0 else 0:.2f}  "
f"max other={w[keep != 0].max():.2f}  output={out:+.3f}")
show("full cache", np.arange(10))
show("window only (last 4)", np.arange(6, 10))          # sink evicted
show("sink + last 4", np.array([0, 6, 7, 8, 9]))        # StreamingLLM-style
def streaming_cache(positions, n_sink=4, window=4):
# keep the first n_sink tokens forever plus the most recent `window` tokens
if len(positions) <= n_sink + window:
return positions
return positions[:n_sink] + positions[-window:]
print("cache after 20 tokens:", streaming_cache(list(range(20))))

Output:

full cache             sink weight=0.81  max other=0.03  output=+0.093
window only (last 4)   sink weight=0.00  max other=0.30  output=+0.549
sink + last 4          sink weight=0.90  max other=0.03  output=+0.055
cache after 20 tokens: [0, 1, 2, 3, 16, 17, 18, 19]

The fix: StreamingLLM

StreamingLLM keeps a small fixed cache made of two parts: the first few tokens (about 4) as sinks, and a rolling window of the most recent tokens. Everything in between is evicted. No retraining is needed; it works with existing models.

The paper reported stable language modelling over streams of up to about 4 million tokens, and large speedups versus recomputing a sliding window. It also showed that adding one dedicated, learnable sink token at the start of every training sample lets a model rely on that single token instead of several.

Streaming is not long-term memory StreamingLLM keeps the model fluent forever; it does not make it remember forever. Anything evicted from the middle is gone. If our car assistant must recall the address the driver said two hours ago, we need retrieval or a summary memory on top, not just sinks.

StreamingLLM and modern attention sinks

Later models build the sink directly into the architecture. A common approach is a learned sink logit: each attention head gets an extra learnable score that joins the softmax denominator but has no value attached. The head can send weight to this "nobody" slot, so it no longer needs to hijack the first real token. OpenAI's gpt-oss models (2025) use learned per-head sinks of this kind. A closely related idea, sometimes called "softmax plus one", adds a constant 1 to the softmax denominator for the same reason.

Importance of attention sinks

  • Streaming and long chats: keeping sinks is a cheap, essential trick whenever a KV cache is truncated.
  • KV cache compression and quantization: methods that evict or compress cache entries usually protect the first tokens, because damaging them hurts quality far more than their count suggests.
  • Interpretability: sinks explain why attention maps often show a bright first column; it is a "no-op" signal, not proof that the first word matters.
  • Architecture design: learned sinks and related gating ideas are now part of how some new models are built.

Pause and think: Our car assistant uses window-only eviction and starts producing nonsense after about 8,000 tokens, which equals the window size. What is the most likely cause and the cheapest fix?

At 8,000 tokens the first tokens are being evicted, so the attention sink disappears and the softmax distribution shifts. The cheapest fix is StreamingLLM-style caching: always keep the first ~4 tokens plus the rolling window.

Going one level deeper

We said a learned sink is an extra score that joins the softmax but has no value attached. Let us see with small numbers what that does, and then look at the second detail that keeps streaming stable: position re-indexing.

One head, four tokens, nothing relevant

  1. Scores: The four real tokens score 0.1, 0.2, 0.0 and 0.1 (illustrative). Their exponentials are about 1.11, 1.22, 1.00 and 1.11, which add up to 4.43.
  2. Without a sink: The weights must sum to 1, so each token gets roughly a quarter. The head is forced to average four irrelevant values.
  3. With a sink logit of 3.0: exp(3.0) ≈ 20.09 joins the denominator: 4.43 + 20.09 = 24.52. The real tokens together get 4.43 / 24.52 ≈ 0.18. The other 0.82 goes to the sink and contributes nothing.
  4. When something matters: If one token scores 5.0, exp(5.0) ≈ 148 dwarfs the sink's 20, so most of the weight goes to that token. The sink only wins when nothing else does.
Position re-indexing in a StreamingLLM-style cache after 20 tokens (4 sinks + window of 4)
Kept tokensOriginal position in the textPosition the model is given
Sink tokens0, 1, 2, 30, 1, 2, 3
Window tokens16, 17, 18, 194, 5, 6, 7

Without re-indexing, the window tokens would carry positions that keep growing, eventually past anything seen in training. A common bug in home-made streaming caches is to keep the sinks but forget this step: output stays fine for a while, then drifts once positions exceed the training length.

Practice: try it yourself

We will write a tiny attention head with an optional learned sink logit and compare two situations: a query with nothing relevant in the context, and a query with one clear match.

practice_learned_sink.py

import math
def attend(scores, values, sink_logit=None):
# softmax over the scores; an optional sink logit joins the denominator
# but has no value attached, so weight sent there adds nothing to the output
exps = [math.exp(s) for s in scores]
denom = sum(exps) + (math.exp(sink_logit) if sink_logit is not None else 0.0)
w = [e / denom for e in exps]
out = sum(wi * vi for wi, vi in zip(w, values))
return w, out
values = [1.0, -1.0, 2.0, 0.5]                    # illustrative value numbers
cases = {"nothing relevant": [0.1, 0.2, 0.0, 0.1],
"one clear match": [0.1, 0.2, 5.0, 0.1]}
for name, scores in cases.items():
for label, sink in [("no sink", None), ("sink logit 3.0", 3.0)]:
w, out = attend(scores, values, sink)
print(f"{name:17s} {label:15s} weight on real tokens={sum(w):.2f} "
f"largest={max(w):.2f} output={out:+.2f}")

Output:

nothing relevant  no sink         weight on real tokens=1.00 largest=0.28 output=+0.55
nothing relevant  sink logit 3.0  weight on real tokens=0.18 largest=0.05 output=+0.10
one clear match   no sink         weight on real tokens=1.00 largest=0.98 output=+1.96
one clear match   sink logit 3.0  weight on real tokens=0.88 largest=0.86 output=+1.73

Now change it:

  • Set the sink logit to 0.0. That is the “softmax plus one” idea, since exp(0) = 1. Predict the weight on real tokens in the “nothing relevant” case.
  • Set the sink logit to 6.0. Predict what happens to the “one clear match” case. Is a stronger sink always better?
  • Change the matching score from 5.0 to 8.0 and keep the sink at 3.0. Predict how much weight the sink still takes.

Pause and think: With nothing relevant, the output is +0.55 without a sink and +0.10 with one. Why is +0.10 closer to what the head should produce?

When no token is relevant, the best contribution is close to nothing. Without a sink, the weights must still sum to 1, so the head averages four unrelated values and injects noise (+0.55). The sink lets most of the weight go to a slot with no value, so the head stays nearly silent.

Pause and think: A learned sink logit is a parameter of the head, not a token in the cache. Why does that make cache eviction safer than in a model that uses its first token as the sink?

A first-token sink lives in the KV cache, so a window policy can evict it and shift every weight. A sink logit is part of the model's weights: it is present in every softmax no matter which tokens are cached, so trimming the cache cannot remove it.

Key takeaways

  • Softmax weights must sum to 1, so heads need somewhere harmless to put unneeded attention.
  • The first tokens become attention sinks because every later token can always see them.
  • Evicting sinks shifts the attention distribution and breaks window-only streaming.
  • StreamingLLM keeps ~4 sink tokens plus a rolling window: fixed memory, stable output.
  • Newer models build in learned sinks; none of this gives long-term memory of evicted tokens.

Key terms

  • Attention sink: A token, usually the first, that absorbs a large share of attention without adding much information.
  • KV cache: Stored keys and values of past tokens used during generation.
  • Window attention: Keeping only the most recent W tokens in the cache.
  • StreamingLLM: A method that keeps a few initial sink tokens plus a rolling window to stream text indefinitely.
  • Perplexity: A measure of how surprised a model is by text; lower is better.
  • Learned sink: An extra learnable attention slot or logit built into each head to absorb unneeded weight.

← 6.4 Sliding Window Attention: Taming Very Long Contexts · 6.6 Flash Attention: Memory-Efficient Attention at Scale →