Prompt Caching vs Semantic Caching
The three kinds of LLM caching (exact response caching, provider prompt/prefix caching, and semantic caching), what each saves, how to structure prompts for cache hits, and the risks of each.
Three different caches
| Exact response cache | Prompt / prefix cache | Semantic cache | |
|---|---|---|---|
| Key | Hash of full request | Token prefix of the prompt | Embedding of the question |
| Returns | The stored response | Nothing directly. The model skips re-processing the prefix | A stored response for a similar question |
| Saves | 100% of the call | Most of the input cost and prefill time for the cached part | 100% of the call |
| Output correctness | Same as before | Identical to uncached | May be wrong for the new question |
| Where | Your gateway (Redis) | Provider or inference engine | Your gateway (vector index) |
| Hit rate | Low for chat, high for fixed inputs | High for long shared prefixes | Medium for FAQ-like traffic |
Prompt (prefix) caching
The model's work on a prompt prefix (its KV cache) can be stored and reused when a later request starts with exactly the same tokens. Providers bill cached input tokens at a steep discount, often around 10% of the normal rate, and TTFT drops because prefill is skipped for the cached part. Self-hosted engines do this automatically with prefix caching (see inference engines).
How to get hits:
- Put static content first and variable content last.
- Don't put timestamps, request IDs or user names at the top of the system prompt, because any change invalidates everything after it.
- Keep tool definitions in a stable order.
- Know your provider's rules: minimum cacheable length, cache lifetime (often minutes, sometimes extendable), whether caching is automatic or needs explicit breakpoints, and whether cache writes cost extra.
- For self-hosted clusters, route conversations to the same replica so the prefix is warm there.
Multi-turn chat and agents benefit most: each turn's prompt is the previous prompt plus a little more, so almost everything is a cache hit.
Exact response caching
Hash the normalised request (model, parameters, messages) and store the response with a TTL.
- Good for: deterministic pipelines with repeated inputs (classifying the same product titles, re-running evals, identical API calls from retries).
- Poor for: free-form chat, where exact repeats are rare.
- Only cache low-temperature, non-personalised responses, and include the model version and prompt version in the key.
Semantic caching
Embed the incoming question. If a previous question is within a similarity threshold, return its stored answer without calling the LLM.
The risk: similar isn't the same.
- "How do I cancel my subscription?" and "How do I cancel my order?" may be close in embedding space but need different answers.
- "Can I get a refund after 30 days?" vs "before 30 days?" can differ by one word.
- Personalised or time-sensitive answers ("my balance", "today's schedule") must never be served from a shared cache.
Making it safe(r):
- Restrict it to low-risk, non-personalised intents such as FAQs and product info.
- Use a high similarity threshold, and tune it on labelled pairs of should-match and shouldn't-match questions.
- Include tenant, locale, user tier and content version in the cache scope, and invalidate when the underlying docs change.
- Log hits and sample them for quality review.
- Optionally verify with a cheap model ("does this cached answer address this question?") before returning it.
Which to use
- Always: structure prompts for prefix caching. It is free quality-wise and cuts cost substantially for chat, RAG with stable instructions, and agents.
- Where inputs repeat exactly: add response caching.
- Only for high-volume, low-risk, FAQ-like traffic: consider semantic caching, carefully scoped and monitored.
Key takeaways
- Exact response caching returns a stored answer for an identical request. It is safe, but hit rates are low for free-form input.
- Prompt (prefix) caching reuses the model's computed state for a repeated prompt prefix, cutting input cost by up to about 90% and TTFT, with identical outputs.
- Semantic caching returns a stored answer for a similar question. It has high savings and real correctness risk, so scope it to low-risk intents.
- Structure prompts static-first and variable-last to maximise prefix-cache hits.