Context Windows and Their Limits
What the context window really holds, why bigger is not free, how models degrade on long inputs, and strategies for managing context in chat and agent systems.
What is in the window
An LLM is stateless between calls. Everything it knows about the current task must be in the prompt of that call. The context window is the maximum number of tokens the prompt plus the generated output can hold.
Window sizes have grown from 4K tokens (2022) to 200K–1M+ tokens in current frontier models. That makes it tempting to put everything in the prompt. It is rarely the right design.
Three costs of long context
1. Money. Input tokens are billed on every call. A 200K-token prompt at $2 per million is $0.40 per request. A 20-turn chat that resends a growing history pays for the early turns 20 times. Prompt caching softens this, cutting cached tokens to about 10% of the price, but does not remove it.
2. Latency. Prefill time grows with prompt length. Attention work grows faster than linearly at very long lengths. A 100K-token prompt can take several seconds before the first token appears, even on fast hardware.
3. Memory and capacity. For self-hosted models, every token in the window holds KV cache in GPU memory. That is ≈ 320 KB per token for a 70B model, so one 128K-token sequence takes ≈ 43 GB. Long contexts shrink how many users a GPU can serve. See GPU memory and the KV cache.
Quality degrades too
A model that accepts 1M tokens doesn't use them all equally well:
- Lost in the middle: research and practitioner tests repeatedly find that models recall facts at the start and end of a long prompt better than facts buried in the middle.
- Distraction: irrelevant context lowers accuracy. More retrieved chunks is not always better, and precision matters.
- Multi-hop reasoning over long inputs is much harder than finding a single "needle".
Advertised limits are measured on needle-in-a-haystack tests. Test on your own task before relying on the full window.
Managing context in chat
As conversations grow, choose a policy:
| Strategy | How | Trade-off |
|---|---|---|
| Sliding window | Keep the last N turns | Simple. Forgets early facts such as the user's name or goal |
| Summarise and truncate | Replace old turns with a running summary | Keeps the gist. Summaries lose detail and cost an extra LLM call |
| Retrieve from history | Embed past turns and fetch relevant ones | Scales to long histories. Adds retrieval complexity |
| Structured memory | Extract facts (preferences, entities) into a store and inject them | Durable across sessions. Needs extraction quality and privacy controls |
Most production assistants combine them: recent turns verbatim, a running summary, and a small set of retrieved facts about the user.
When long context is the right tool
- One-off deep analysis of a single large document, contract or codebase, where retrieval would miss cross-references.
- Low volume, high value tasks where cost per call doesn't matter.
- Stable prefixes reused across many requests, such as a product manual or policy. Prompt caching makes these cheap after the first call.
Managing context in agents
Agents append every tool call and result to their context, so the context grows with each step. A 30-step agent can reach hundreds of thousands of tokens, and cost grows faster than linearly because each step resends everything before it. Techniques:
- Trim tool outputs. Return summaries or the relevant lines, not a whole web page or log file.
- Compaction. Periodically summarise the history and restart with the summary.
- Sub-agents. Delegate a subtask to a fresh context, and return only the result.
- Files as memory. Write notes and intermediate results to a scratchpad the agent can read back on demand.
Key takeaways
- The context window holds everything the model sees in one call, including system prompt, tools, history, retrieved documents and its own output.
- Long context costs money (linear in tokens), TTFT (prefill time), and GPU memory (KV cache).
- Quality often drops for information buried in the middle of very long inputs. Retrieval of the right few thousand tokens usually beats stuffing.
- Chat and agent systems need an explicit context-management policy, such as truncation, summarisation, retrieval or offloading to memory.