Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. "Lost in the Middle: How Language Models Use Long Contexts." Transactions of the Association for Computational Linguistics (TACL), 2024. arXiv:2307.03172 — https://arxiv.org/abs/2307.03172
Feeding a model more context doesn't mean it actually uses more context — that's the uncomfortable finding at the center of "Lost in the Middle," a Stanford/UC Berkeley/Samaya AI study that measured, empirically, where in a long prompt a language model actually pays attention. The answer has a shape: a U. Models are reliably good at using information placed at the very start or the very end of their input, and reliably worse — sometimes dramatically worse — at using information buried in the middle. 📉
If you're building anything retrieval-augmented — a RAG pipeline, a document Q&A tool, a long-context agent — this paper isn't background reading. It's a design constraint. 🧠
🔄 The Core Finding: A U-Shaped Curve
The researchers ran controlled experiments where they kept everything about a prompt fixed except one variable: the position of the single piece of information that actually answered the question. Same documents, same distractors, same total length — only the location of the answer moved, from the front of the context, through the middle, to the very end.
Accuracy traced a consistent U shape across every task and nearly every model tested. Performance peaked when the relevant fact sat at the beginning or end of the input, and dropped sharply once it moved toward the center. In some configurations, models did worse with the answer buried in the middle of a stuffed context than they did with no documents at all — relying purely on what they'd memorized during training. Adding more context, in other words, can actively hurt a model's ability to find and use the one fact that matters.
Self-attention, in theory, gives every token in the context equal computational access to every other token — there's no architectural reason position should matter this much. The paper connects this to the serial-position effect from human memory research: people also recall the first and last items in a list far better than the middle ones. Seeing an analogous pattern emerge in transformer-based models, despite their theoretically uniform attention, is one of the paper's more striking observations. 🧬
🔬 How the Researchers Tested This
Two tasks were used to isolate the effect, moving from realistic to synthetic:
| Task | What It Measures | Setup |
|---|---|---|
| 📚 Multi-Document QA | Realistic RAG-style reasoning — mirrors how a retrieval pipeline feeds passages to a generator | Questions from NaturalQuestions-Open, paired with one Wikipedia passage that contains the answer plus a variable number of retrieved distractor passages that don't. Context sizes tested: 10, 20, and 30 documents. |
| 🔑 Key-Value Retrieval | Pure exact-match retrieval, stripped of language and reasoning, to isolate the "can it even find the token" question | A JSON object of random UUID key-value pairs; the model must return the value for one specified key. Context sizes tested: 75, 140, and 300 pairs. |
Models spanned both open and closed weights, and both standard and extended-context variants of the same base model: MPT-30B-Instruct, LongChat-13B (16K), OpenAI's GPT-3.5-Turbo and its 16K variant, and Anthropic's Claude and Claude (100K), with a smaller-scale GPT-4 evaluation as an additional data point.
📊 Key Results
On multi-document QA, GPT-3.5-Turbo's accuracy dropped more than 20 percentage points depending solely on where the answer-bearing document sat in the context. At its worst point — the answer buried near the middle of a 20- or 30-document context — its accuracy fell below its own closed-book baseline, where it was given no documents at all and had to answer purely from memorized knowledge. Handing the model more retrieved evidence made it perform worse than giving it none. 📉
Three other findings stood out:
- 📏 Longer contexts hurt performance overall, independent of position. Averaged across all positions, accuracy on both tasks declined as more documents or key-value pairs were added — models simply struggle more as there's more to sift through.
- 🧵 Bigger context windows don't fix the problem. GPT-3.5-Turbo and its 16K counterpart, evaluated on inputs both could handle, produced nearly identical accuracy curves. A model's ability to accept a long prompt says nothing about its ability to reason effectively across all of it.
- 🔑 Even trivial exact-match retrieval isn't immune. The key-value task requires no reasoning at all — just finding a matching token — yet most models still showed the same U-shaped drop for information placed mid-context. (Claude was a notable exception here, scoring near-perfect at every position and length tested.)
🧩 Why Does This Happen?
The authors ran several follow-up experiments to probe the mechanism behind the U-shape, rather than just document it:
| Variable Tested | Finding |
|---|---|
| 🏗️ Architecture (decoder-only vs. encoder-decoder) | Encoder-decoder models (Flan-T5-XXL, Flan-UL2) stayed largely flat across positions — as long as the input fit within their original training-time sequence length. Once evaluated on sequences longer than what they were trained on, they developed the same U-shaped curve. Bidirectional encoding seems to help, but only within a familiar length range. |
| ❓ Query-Aware Contextualization (placing the question both before and after the documents) | This nearly eliminated the drop on the key-value task — GPT-3.5-Turbo (16K) reached near-perfect accuracy at 300 pairs, versus its lowest point of roughly 46% without this change. On the harder multi-document QA task, though, the fix barely moved the needle, suggesting position sensitivity there comes from something deeper than prompt formatting. |
| 🎓 Instruction Fine-Tuning | Comparing MPT-30B against its instruction-tuned counterpart MPT-30B-Instruct showed both exhibit the same U-shaped curve. Fine-tuning raised overall accuracy but didn't change the shape — this behavior appears to originate in pretraining, not instruction-tuning. |
📥 Is More Context Always Better? A Retrieval Case Study
The paper closes with a practical experiment that speaks directly to RAG system design: a retriever-reader pipeline on open-domain question answering, where a retriever returns an increasing number of documents and a language model reads them to produce an answer.
Retriever recall kept climbing as more documents were added — unsurprisingly, casting a wider net finds more of the right passages. But reader accuracy plateaued long before recall did. Past roughly 20 retrieved documents, adding more barely moved the needle — about a 1.5-point improvement for GPT-3.5-Turbo and roughly a 1-point improvement for Claude — while context length, latency, and cost kept rising. The retrieval system was doing its job; the reader simply couldn't capitalize on the extra evidence.
Past a certain point, retrieving more documents is not a reasoning upgrade — it's mostly cost and latency with a shrinking return. The paper's authors point to two practical levers instead: reranking (surfacing the most relevant passage toward the start of the context, where models actually use it) and ranked-list truncation (returning fewer, better-targeted documents rather than padding the context to some fixed count). 🎯
🛠️ Practical Implications for RAG Engineers
- Rerank before you truncate, not after. If your pipeline can only reliably use information near the start or end of the context, put your highest-confidence passage there — don't just concatenate retrieval results in raw similarity-score order and hope the model reads all of it evenly.
- Stop equating "bigger context window" with "better retrieval quality." A model that accepts 100K tokens is not necessarily better at reasoning across all 100K of them than a model with an 8K window reasoning across a well-curated 8K. Context capacity and context utilization are different properties, and only one of them shows up in a spec sheet.
- Measure accuracy against document count, not just retrieval recall. The open-domain QA case study is a warning: recall metrics can look great while end-to-end answer accuracy quietly plateaus. Track both, and find your pipeline's actual saturation point instead of assuming "more retrieved documents" scales linearly with "better answers."
- Treat position as a first-class variable in evaluation. A RAG system that scores well on average across a benchmark can still fail specifically when the answer lands in the middle of a stuffed context — average accuracy can hide a real, exploitable weak spot.
- Consider query-aware contextualization for retrieval-style subtasks. Placing the query both before and after the retrieved content had a large effect on exact-match retrieval performance in this study. It's a cheap prompt-engineering change worth testing on tasks that resemble key-value lookup more than open-ended reasoning.
❓ Frequently Asked Questions
It measures how a language model's accuracy changes purely as a function of where the relevant piece of information sits within a long input context, using multi-document question answering and a synthetic key-value retrieval task, while holding context length and content otherwise constant.
It describes the pattern where model accuracy is highest when relevant information appears at the very start or very end of the input context, and drops significantly when that information is located near the middle — visually resembling the letter U when plotted against position.
Not on its own. The study found that a model and its extended-context counterpart produced nearly identical accuracy curves on inputs both could process, indicating that the ability to accept more tokens doesn't translate into better reasoning across all of them.
No. The paper's open-domain QA case study found that reader accuracy plateaus well before retriever recall does, meaning documents added past a certain point mostly add latency and cost rather than improving the final answer, and reranking or truncating the retrieved list is often more effective than simply retrieving more.
The core experiments covered MPT-30B-Instruct, LongChat-13B (16K), OpenAI's GPT-3.5-Turbo and its extended 16K variant, and Anthropic's Claude and Claude (100K), with additional experiments on Flan-T5-XXL, Flan-UL2, base MPT-30B, and a smaller-scale evaluation of GPT-4.
📝 Summary
- The core finding → language models use context unevenly: best at the start and end, worst in the middle — a U-shaped accuracy curve that held across nearly every model and task tested.
- Longer contexts compound the problem → accuracy drops as context grows, regardless of where the answer sits.
- A bigger context window isn't a fix → extended-context variants of the same model showed the same curve as their shorter-context counterparts.
- More retrieved documents isn't automatically better → reader accuracy plateaus long before retriever recall does, favoring reranking and truncation over brute-force retrieval volume.
- For RAG builders → treat position, not just recall or context length, as a first-class design and evaluation variable.
Happy building! ✨
Comments
Post a Comment