Modern AI Engineering

Lesson 5.4 · 23 min

Lost in the Middle: Why LLMs Miss Central Context

You give a model 20 documents and the answer is right there in document 10. Why might it do worse than if you had given it no documents at all?

In short: LLMs tend to use information at the beginning and end of a long context much better than information in the middle. Plotting accuracy against the position of the key fact gives a U-shaped curve. This "lost in the middle" effect, documented in a 2023 study, matters for RAG, long chats and document analysis. We can test for it by moving a known fact through the context, and reduce it by sending fewer, better-ranked chunks, placing the most relevant ones at the edges, and putting the question after the documents.

What is a context window?

The context window is the maximum number of tokens an LLM can take into account at once: the system prompt, the conversation so far, any retrieved documents, the question, and the answer being written. Anything outside the window is invisible to the model. Windows have grown from a few thousand tokens in early chat models to hundreds of thousands or more in many 2025–2026 models.

It is tempting to read a large window as "the model will read everything carefully". The window only sets what the model can see. How well it uses each part is a separate question, and the answer turns out to depend on where the information sits.

Think of it like a long meeting After a three-hour meeting, most people remember the opening agenda and the final decisions, but the details from the hour-two discussion are fuzzy. Psychologists call these the primacy and recency effects. LLMs show a surprisingly similar pattern with long contexts.

What is the lost-in-the-middle problem?

The lost-in-the-middle problem is the tendency of LLMs to answer more accurately when the relevant information is near the start or the end of the context, and less accurately when it is buried in the middle, even though all of it fits comfortably in the window.

The name comes from the 2023 paper "Lost in the Middle: How Language Models Use Long Contexts" by Nelson F. Liu and colleagues (Stanford and others, later published in TACL). They gave models a question plus 10, 20 or 30 retrieved documents, exactly one of which contained the answer, and moved that document from first to last. They also used a synthetic task: find the value for a given key in a long list of random key-value pairs. Across the models they tested, accuracy was highest at the edges and dropped in the middle. In some settings, accuracy with the answer in the middle fell below the model's accuracy with no documents at all.

Let us understand it with an example

Our support bot uses retrieval-augmented generation (RAG): for each question it retrieves 10 policy snippets and pastes them into the prompt. A customer asks "How many days do I have to request a refund?" One snippet says "Refunds are allowed within 45 days of delivery." The other nine are about gift cards, shipping and loyalty points.

  • Refund snippet placed first: the bot answers "45 days" reliably.
  • Refund snippet placed last, right before the question: also reliable.
  • Refund snippet placed fifth: the bot is more likely to say "30 days" (a common default it has seen in training), mix in a gift-card rule, or say the policy does not mention refunds.

Nothing about the information changed. Only its position did. That is what makes this problem sneaky: your retrieval can be perfect and the answer can still be wrong.

Pause and think: Pause and think: why can middle-position accuracy drop below "no documents" accuracy?

Without documents the model answers from what it learned in training (closed-book). With documents, it tries to answer from the context, but if it fails to use the buried relevant one, it may be distracted by the other, irrelevant documents and do worse than if it had just relied on memory.

The U-shaped curve

Accuracy plotted against position forms a U: high at the start (primacy), high at the end (recency), lower in between. Two further patterns are common: the dip tends to get deeper as the context gets longer, and, in the original study, models marketed with longer context windows were not automatically better at using the context they already had.

It varies by model Newer long-context models are trained specifically to retrieve from anywhere in the window, and many now score near-perfectly on simple "find this sentence" tests. The effect tends to reappear on harder tasks: several relevant facts, facts that must be combined, or near-duplicate distractors. Measure your own model on your own task rather than assuming either way.

Why does this happen?

There is no single proven cause. These are the leading explanations, and several likely combine:

  • Patterns in training data. In real text, the most important information is often at the start (titles, instructions, topic sentences) or immediately before what you are predicting (the most recent words). Models learn to weight those positions heavily.
  • Causal attention favours early tokens. With a causal mask, the first tokens are visible to every later position and tend to collect a lot of attention. Researchers studying streaming models observed that LLMs put unusually large attention on the very first tokens, sometimes called attention sinks.
  • Position encodings favour nearby tokens. With RoPE, attention scores tend to decay with distance, so tokens near the end of the context (close to where the answer is being generated) get a natural boost.
  • Limited long-context training. Much fine-tuning data is short. A model may rarely have been rewarded for finding one fact in the middle of 50,000 tokens.
  • Distractors. The middle is usually filled with plausible but irrelevant text, and attention is a finite budget shared across all of it.

Where this hurts us in real life

Common situations where middle content gets under-used
ScenarioWhat goes wrong
RAG with many retrieved chunksThe best chunk lands in position 6 of 12 and is ignored; the answer comes from a weaker chunk or from memory
Long chat sessionsA user preference stated 40 turns ago (now mid-context) is forgotten
Long document review (contracts, reports)A clause on page 30 of 60 is missed in a summary or risk check
Long system promptsRules buried in the middle of a long instruction block are followed less reliably
Agents with long tool historiesAn important tool result from many steps ago is not used

Real-world pattern Teams often discover this when they raise the number of retrieved chunks from 5 to 20 "to be safe" and see answer quality go down. More context means more middle and more distractors.

How to test for the lost-in-the-middle problem

A position sweep for your own model and task

  1. Pick a known fact: Choose a question whose answer is in exactly one document or sentence (the "needle").
  2. Build filler: Surround it with realistic, irrelevant documents from your own domain (the "haystack"). Realistic distractors make the test honest.
  3. Sweep the position: Insert the fact at several depths: 0%, 25%, 50%, 75%, 100% of the context.
  4. Sweep the length: Repeat at several total lengths, e.g. 4K, 16K, 64K tokens.
  5. Score many runs: Ask the question many times per cell (with several different facts), score correctness, and plot accuracy by depth and length.

position_test.py

# 1) Build a position test: same question, the key fact moved through the context
filler = [f"Doc {i}: Store policy note {i} about gift cards." for i in range(1, 10)]
needle = "Doc X: Refunds are allowed within 45 days of delivery."
question = "Q: How many days do customers have to request a refund?"
def build_prompt(depth):
docs = filler[:depth] + [needle] + filler[depth:]
return "\n".join(docs + [question])
for depth in [0, 4, 9]:
lines = build_prompt(depth).splitlines()
pos = next(i for i, l in enumerate(lines) if l.startswith("Doc X"))
print(f"depth {depth}: fact is line {pos + 1} of {len(lines) - 1} docs")
# 2) Fix: put the most relevant chunks at the START and END, weakest in the middle
def edges_first(chunks_by_relevance):
front, back = [], []
for i, c in enumerate(chunks_by_relevance):     # best first
(front if i % 2 == 0 else back).append(c)
return front + back[::-1]
ranked = ["R1", "R2", "R3", "R4", "R5", "R6", "R7"]  # R1 = most relevant
print("retriever order:", ranked)
print("edges-first    :", edges_first(ranked))

Output:

depth 0: fact is line 1 of 10 docs
depth 4: fact is line 5 of 10 docs
depth 9: fact is line 10 of 10 docs
retriever order: ['R1', 'R2', 'R3', 'R4', 'R5', 'R6', 'R7']
edges-first    : ['R1', 'R3', 'R5', 'R7', 'R6', 'R4', 'R2']

Simple single-needle tests became popular in 2023 and many current models pass them easily, so also test harder variants: several needles, questions that require combining two facts, and benchmarks designed for this, such as RULER (2024), which add multi-hop and aggregation tasks.

How to solve the lost-in-the-middle problem

Two more techniques help: ask the model to first quote the relevant passages and then answer from those quotes (this forces it to search the context explicitly), and keep long chats healthy by periodically summarizing older turns and pinning key facts (like user preferences) near the end of the prompt.

Common mistakes Assuming "it fits in the window" means "the model will use it". Raising top-k retrieval to 20 or 50 chunks without testing. Putting the question at the very top followed by 30,000 tokens of documents. Trusting a vendor's single-needle score for a task that needs combining several facts. And reordering chunks by relevance score from a weak retriever, which just puts the wrong chunks at the edges.

Worked example, step by step

The test recipe said "score many runs". How many is "many"? Let us work through one result to see why this matters. Suppose (illustrative numbers) we ran our refund question 20 times at each depth and got 18 correct with the fact first (90%), 14 correct in the middle (70%) and 17 correct at the end (85%). It looks like a clear dip. Is it?

Is a 70% middle really worse than a 90% edge?

  1. Error of the middle cell: √(0.7 · 0.3 / 20) = √0.0105 ≈ 0.10, so about 10 points.
  2. Plausible range: Two standard errors each way: the true middle accuracy could be anywhere from about 50% to 90%.
  3. Compare: That range reaches the 90% we measured at the edge. With 20 runs we cannot say the middle is worse. The dip may be chance.
  4. Repeat with 100 runs: Suppose we again measure 70%. Now the error is √(0.7 · 0.3 / 100) ≈ 0.046, so the range is about 61% to 79%.
  5. Compare again: The edge at 90% with 100 runs has an error of 3 points, so a range of about 84% to 96%. The two ranges no longer overlap. Now the dip is real.
Standard error of a measured accuracy near 60%, from the formula
Runs per cellStandard errorRough range around 60%
1015.5 points29% to 91%
2011.0 points38% to 82%
506.9 points46% to 74%
1004.9 points50% to 70%
4002.4 points55% to 65%

To halve the error we need four times as many runs. That is expensive, so spend runs where they matter: fewer depths (start, middle, end) with more runs each tells us more than many depths with a handful of runs. And use several different facts and questions, not one question repeated, so that a single lucky or unlucky wording does not decide the result.

The mistake this prevents Teams often change the prompt, re-run a small sweep, see the middle go from 60% to 70% and ship the change. With 20 runs per cell, a 10-point move is well inside the noise. The "improvement" may vanish on the next run.

Practice: try it yourself

We cannot call a real model here, so we build a stand-in: a simulated model whose chance of using the fact follows a made-up U-shape. Because we know its true accuracy at every depth, we can see how well a position sweep recovers it, with 20 runs and with 1,000 runs per depth.

practice_position_sweep.py

import random
# A SIMULATED model, not a real one: its chance of using the fact follows a
# made-up U-shape. 0.95 at the edges of the context, 0.60 in the middle.
def simulated_model(depth, rng):
p_correct = 0.60 + 0.35 * (2 * depth - 1) ** 2
if rng.random() < p_correct:
return "You have 45 days from delivery to request a refund."
return "Refunds are usually possible within 30 days."      # the wrong default
def is_correct(answer):
return "45 days" in answer                # simple string check for the needle
def sweep(runs, seed):
rng = random.Random(seed)
row = []
for depth in (0.0, 0.25, 0.5, 0.75, 1.0):
hits = sum(is_correct(simulated_model(depth, rng)) for _ in range(runs))
row.append(100 * hits / runs)
return row
print("depth of the fact       0%    25%    50%    75%   100%")
print("true chance (made up) " + "".join(
f"{100 * (0.60 + 0.35 * (2 * d - 1) ** 2):7.1f}" for d in (0, 0.25, 0.5, 0.75, 1)))
for runs, seed in [(20, 1), (20, 2), (20, 3), (1000, 4)]:
print(f"{runs:5d} runs per depth  " + "".join(f"{v:7.1f}" for v in sweep(runs, seed)))

Output:

depth of the fact       0%    25%    50%    75%   100%
true chance (made up)    95.0   68.8   60.0   68.8   95.0
20 runs per depth    100.0   90.0   50.0   75.0   90.0
20 runs per depth     90.0   70.0   45.0   65.0  100.0
20 runs per depth     95.0   55.0   45.0   60.0   85.0
1000 runs per depth     94.1   69.0   61.6   69.5   95.5

Now change it:

  • Remove the U-shape: set p_correct = 0.8 for every depth. Run the three 20-run sweeps again. Before running, predict whether any row will still look like it has a dip in the middle.
  • Change the wrong answer to "Refunds are possible within 30 days, not 45 days." Predict the measured accuracy at every depth. What does this say about scoring by substring?
  • Change the 20-run sweeps to 100 runs. Using the table from the previous section, predict how far the measured values will typically be from the true ones.

Pause and think: In the first 20-run sweep, the 25% depth scored 90%. A teammate concludes that this model has no problem at 25% depth. What do we tell them?

That 20 runs cannot support the claim. The true value in our simulation is 68.8%. With 20 runs the standard error is about 10 points, so a result of 90% is an unlucky draw about two standard errors high, and such draws do happen: the other two sweeps gave 70% and 55% for the same depth. We need more runs, or at least several repeated sweeps, before we read anything into one cell.

Pause and think: Our scorer checks whether the answer contains "45 days". Name one answer it would wrongly count as correct and one it would wrongly count as wrong.

Wrongly correct: "The policy is 30 days, not 45 days", which contains the text but gives the wrong answer. Wrongly wrong: "You have forty-five days" or "a 45-day window", which are right but do not contain the exact text. A substring check is quick, but we should read a sample of scored answers by hand, and tighten the check or use a more careful grader when the two disagree.

Key points to remember

  • A big context window tells you what fits, not what gets used well.
  • Accuracy versus position is often U-shaped: strong at the start and end, weaker in the middle, and worse with longer contexts.
  • Causes are believed to include training-data patterns, attention favouring early tokens, distance decay in position encodings and distractors.
  • Measure it on your model and task by sweeping the position and length of a known fact.
  • Mitigate with fewer, reranked chunks, edges-first ordering, question-last prompts, quote-then-answer and map-reduce.

Key takeaways

  • The context window limits what the model can see; position affects how well it uses what it sees.
  • Accuracy vs position of the key fact is often U-shaped, and the middle dip tends to deepen with length.
  • Likely causes: training-data patterns, attention favouring early tokens, distance decay in position encodings and distractors.
  • Test with a depth × length sweep on your own task, including harder multi-fact variants.
  • Fix with fewer reranked chunks, edges-first ordering, question-last prompts, quote-then-answer and map-reduce.

Key terms

  • Context window: The maximum number of tokens a model can take into account at once.
  • Lost in the middle: The tendency to under-use information placed in the middle of a long context.
  • Primacy effect: Better use or recall of information that appears first.
  • Recency effect: Better use or recall of information that appears last.
  • Needle in a haystack test: Hiding a known fact at different depths of filler text and checking if the model finds it.
  • Reranker: A model that re-scores retrieved chunks so only the most relevant few are sent to the LLM.
  • Map-reduce: Processing chunks separately and then combining the partial results.

← 5.3 Token Streaming: Rendering Outputs as They Arrive · 6.1 A Timeline of LLM Architecture Improvements →