GenAI System Design
1. LLM Fundamentals for System Designers

Context Windows and Their Limits

What the context window really holds, why bigger is not free, how models degrade on long inputs, and strategies for managing context in chat and agent systems.

Lesson 2 of 6 10 min

What is in the window

An LLM is stateless between calls. Everything it knows about the current task must be in the prompt of that call. The context window is the maximum number of tokens the prompt plus the generated output can hold.

System prompt
role, rules, format
Tool definitions
JSON schemas
Retrieved context
documents, search results
Conversation history
all previous turns
User message
Output
counts too
A typical chat request. History and retrieved context usually dominate, and both grow.

Window sizes have grown from 4K tokens (2022) to 200K–1M+ tokens in current frontier models. That makes it tempting to put everything in the prompt. It is rarely the right design.

Three costs of long context

1. Money. Input tokens are billed on every call. A 200K-token prompt at $2 per million is $0.40 per request. A 20-turn chat that resends a growing history pays for the early turns 20 times. Prompt caching softens this, cutting cached tokens to about 10% of the price, but does not remove it.

2. Latency. Prefill time grows with prompt length. Attention work grows faster than linearly at very long lengths. A 100K-token prompt can take several seconds before the first token appears, even on fast hardware.

3. Memory and capacity. For self-hosted models, every token in the window holds KV cache in GPU memory. That is ≈ 320 KB per token for a 70B model, so one 128K-token sequence takes ≈ 43 GB. Long contexts shrink how many users a GPU can serve. See GPU memory and the KV cache.

Quality degrades too

A model that accepts 1M tokens doesn't use them all equally well:

  • Lost in the middle: research and practitioner tests repeatedly find that models recall facts at the start and end of a long prompt better than facts buried in the middle.
  • Distraction: irrelevant context lowers accuracy. More retrieved chunks is not always better, and precision matters.
  • Multi-hop reasoning over long inputs is much harder than finding a single "needle".

Advertised limits are measured on needle-in-a-haystack tests. Test on your own task before relying on the full window.

Managing context in chat

As conversations grow, choose a policy:

StrategyHowTrade-off
Sliding windowKeep the last N turnsSimple. Forgets early facts such as the user's name or goal
Summarise and truncateReplace old turns with a running summaryKeeps the gist. Summaries lose detail and cost an extra LLM call
Retrieve from historyEmbed past turns and fetch relevant onesScales to long histories. Adds retrieval complexity
Structured memoryExtract facts (preferences, entities) into a store and inject themDurable across sessions. Needs extraction quality and privacy controls

Most production assistants combine them: recent turns verbatim, a running summary, and a small set of retrieved facts about the user.

When long context is the right tool

  • One-off deep analysis of a single large document, contract or codebase, where retrieval would miss cross-references.
  • Low volume, high value tasks where cost per call doesn't matter.
  • Stable prefixes reused across many requests, such as a product manual or policy. Prompt caching makes these cheap after the first call.

Managing context in agents

Agents append every tool call and result to their context, so the context grows with each step. A 30-step agent can reach hundreds of thousands of tokens, and cost grows faster than linearly because each step resends everything before it. Techniques:

  • Trim tool outputs. Return summaries or the relevant lines, not a whole web page or log file.
  • Compaction. Periodically summarise the history and restart with the summary.
  • Sub-agents. Delegate a subtask to a fresh context, and return only the result.
  • Files as memory. Write notes and intermediate results to a scratchpad the agent can read back on demand.

Key takeaways

  • The context window holds everything the model sees in one call, including system prompt, tools, history, retrieved documents and its own output.
  • Long context costs money (linear in tokens), TTFT (prefill time), and GPU memory (KV cache).
  • Quality often drops for information buried in the middle of very long inputs. Retrieval of the right few thousand tokens usually beats stuffing.
  • Chat and agent systems need an explicit context-management policy, such as truncation, summarisation, retrieval or offloading to memory.

Go deeper

Finished reading? Mark it done to track your progress.