GenAI System Design
6. Integrating LLMs into Existing Systems

Prompt Caching vs Semantic Caching

The three kinds of LLM caching (exact response caching, provider prompt/prefix caching, and semantic caching), what each saves, how to structure prompts for cache hits, and the risks of each.

Lesson 6 of 6 10 min

Three different caches

Exact response cachePrompt / prefix cacheSemantic cache
KeyHash of full requestToken prefix of the promptEmbedding of the question
ReturnsThe stored responseNothing directly. The model skips re-processing the prefixA stored response for a similar question
Saves100% of the callMost of the input cost and prefill time for the cached part100% of the call
Output correctnessSame as beforeIdentical to uncachedMay be wrong for the new question
WhereYour gateway (Redis)Provider or inference engineYour gateway (vector index)
Hit rateLow for chat, high for fixed inputsHigh for long shared prefixesMedium for FAQ-like traffic

Prompt (prefix) caching

The model's work on a prompt prefix (its KV cache) can be stored and reused when a later request starts with exactly the same tokens. Providers bill cached input tokens at a steep discount, often around 10% of the normal rate, and TTFT drops because prefill is skipped for the cached part. Self-hosted engines do this automatically with prefix caching (see inference engines).

How to get hits:

Tool definitions
static
System prompt
static
Reference docs
static per product
Conversation history
grows, prefix stable
Latest user message
changes
Order content from most stable to least stable. The cache matches from the start of the prompt.
  • Put static content first and variable content last.
  • Don't put timestamps, request IDs or user names at the top of the system prompt, because any change invalidates everything after it.
  • Keep tool definitions in a stable order.
  • Know your provider's rules: minimum cacheable length, cache lifetime (often minutes, sometimes extendable), whether caching is automatic or needs explicit breakpoints, and whether cache writes cost extra.
  • For self-hosted clusters, route conversations to the same replica so the prefix is warm there.

Multi-turn chat and agents benefit most: each turn's prompt is the previous prompt plus a little more, so almost everything is a cache hit.

Exact response caching

Hash the normalised request (model, parameters, messages) and store the response with a TTL.

  • Good for: deterministic pipelines with repeated inputs (classifying the same product titles, re-running evals, identical API calls from retries).
  • Poor for: free-form chat, where exact repeats are rare.
  • Only cache low-temperature, non-personalised responses, and include the model version and prompt version in the key.

Semantic caching

Embed the incoming question. If a previous question is within a similarity threshold, return its stored answer without calling the LLM.

The risk: similar isn't the same.

  • "How do I cancel my subscription?" and "How do I cancel my order?" may be close in embedding space but need different answers.
  • "Can I get a refund after 30 days?" vs "before 30 days?" can differ by one word.
  • Personalised or time-sensitive answers ("my balance", "today's schedule") must never be served from a shared cache.

Making it safe(r):

  • Restrict it to low-risk, non-personalised intents such as FAQs and product info.
  • Use a high similarity threshold, and tune it on labelled pairs of should-match and shouldn't-match questions.
  • Include tenant, locale, user tier and content version in the cache scope, and invalidate when the underlying docs change.
  • Log hits and sample them for quality review.
  • Optionally verify with a cheap model ("does this cached answer address this question?") before returning it.

Which to use

  1. Always: structure prompts for prefix caching. It is free quality-wise and cuts cost substantially for chat, RAG with stable instructions, and agents.
  2. Where inputs repeat exactly: add response caching.
  3. Only for high-volume, low-risk, FAQ-like traffic: consider semantic caching, carefully scoped and monitored.

Key takeaways

  • Exact response caching returns a stored answer for an identical request. It is safe, but hit rates are low for free-form input.
  • Prompt (prefix) caching reuses the model's computed state for a repeated prompt prefix, cutting input cost by up to about 90% and TTFT, with identical outputs.
  • Semantic caching returns a stored answer for a similar question. It has high savings and real correctness risk, so scope it to low-risk intents.
  • Structure prompts static-first and variable-last to maximise prefix-cache hits.

Go deeper

Finished reading? Mark it done to track your progress.