Academy
September 25, 2026
LLM Caching in Production: What to Cache
Learn where LLM caching saves time and spend, when it creates risk, and how to measure a cache without hiding poor answers.
- llm caching
- performance
- architecture

LLM caching is not a switch labelled “make it cheaper”. It is a decision to reuse a previous result, a reusable prompt prefix, or a piece of retrieval work instead of paying to repeat it. Done carefully, it reduces latency and provider spend. Done casually, it returns a stale policy answer, loses a provider-side prompt-cache hit, or makes a team believe a slow workflow is healthy because the dashboard only shows cache hits.
This guide separates the cache layers, shows what each can safely store, and gives a practical way to measure the trade-off. The proof is deliberately operational: inspect a real request receipt, compare the cached and uncached paths, and keep a route to bypass the cache when a user needs a fresh answer.
The five caches people mean by LLM caching
The word is overloaded. These layers solve different problems and should not share an expiry policy.
| Layer | Typical key | Good candidate | Main failure mode |
|---|---|---|---|
| HTTP response | URL and request body | Public, deterministic lookups | Serving one user's data to another |
| Application response | Normalised task plus version | Repeated classification | Stale result after a policy change |
| Retrieval | Query, filters and corpus version | Stable knowledge-base search | Answer based on an old document set |
| Semantic cache | Meaningfully similar request | Narrow, repeatable questions | Similar wording hides different intent |
| Provider prompt cache | Exact reusable prefix | Long system prompt or document prefix | A route change misses the prefix hit |
Start with the layer closest to a deterministic result. A product catalogue lookup with a versioned response can tolerate a normal response cache. A free-form support reply is a poor candidate unless the cached object is a reviewed answer to a tightly defined question. Semantic caching is useful, but it is never a substitute for permissions, retrieval filters or an output check.
Cache a decision, not an unbounded conversation
The safer cache unit is a small, named job: classify-refund-reason:v4, extract-invoice-fields:v2, or answer-handbook-question:2026-09. Put the task version, tenant or access scope, model family, input normalisation version and relevant source version in the key. That may feel fussy on day one. It is far cheaper than explaining why a revised employment policy still produced last week's answer.
A useful pre-flight question is: could two users with this key legitimately receive different answers? If the answer is yes because of account permissions, location, date, feature flag or document access, those dimensions belong in the key or the result should not be shared. Hashing a prompt does not solve authorisation; it just makes the unsafe key harder to read.
For a customer-support assistant, one team might cache the retrieved public help-centre passages for five minutes, but never cache the final account-specific response. For invoice extraction, the document checksum plus extractor version can be a durable key because the desired fields do not change once the file is fixed. The distinction matters more than the cache product you choose.
Prompt caching needs a stable prefix
Provider prompt caching is different from saving a completed answer. It reuses processing for an exact or near-exact prefix supplied again to the provider. It is most valuable when a request starts with a long, stable system instruction, tool schema, or document bundle and ends with a short user-specific turn.
Put stable material first and variable material last. Do not inject a timestamp, random trace ID or changing account summary into the opening block if you expect the prefix to repeat. Our guide to prompt caching and routing explains the awkward consequence for a router: changing providers after a warm prefix can cost more than the model-price comparison suggests.
The practical check is simple. Log the provider's reported cached-input tokens separately from input tokens, then compare two otherwise identical calls. If the count never moves, inspect the literal request prefix before negotiating a larger caching budget. The OpenAI API reference also makes clear that API responses and model behaviour can change, which is one reason to pin your task and evaluation versions rather than cache an assumption indefinitely.
A sensible expiry is a product decision
Time-to-live should follow the thing that makes the answer invalid, not a generic “one hour” setting. A cache for a published release note can expire on a content revision. A cache for a support entitlement should expire when the entitlement changes. A cache for a question about “today's price” should normally be bypassed.
Use explicit invalidation where an owning system can emit it. When a handbook page is published, increment the corpus version. When a feature flag changes, bump the task version. When neither is possible, choose a conservative TTL and surface the cache age in internal logs. Expiry is not a compliance control; it is merely a fallback for systems that cannot say precisely what changed.
Measure the whole request, not just hit rate
High hit rate can be bad news if the cache is attached to the wrong work. Measure the following on cached and uncached requests:
| Measure | Why it matters | Healthy question |
|---|---|---|
| Hit rate | Shows reuse, not quality | Which repeated task is it reducing? |
| P50 and P95 latency | Captures the user-visible path | Does the miss path remain acceptable? |
| Input and output tokens | Separates cache savings from model changes | Did prompt-cache tokens actually rise? |
| Error and correction rate | Detects stale or wrong reuse | Are cache hits later reopened? |
| Bypass rate | Shows whether staff trust the cache | Why are people asking for a fresh run? |
Suppose a help centre answer is served in 180 ms from cache but is reopened by an agent twice as often as a fresh answer. The latency chart looks lovely while the operation gets worse. Pair cache metrics with the same groundedness or verification signal you use on fresh traffic. Caching an unverified output merely makes an error cheaper to repeat.
A small rollout that will not surprise the team
Begin with one task where the expected result is bounded and where a human can inspect a sample. Add a shadow lookup that records whether a key would have hit, but does not serve the cached answer. Review false matches and the source-version logic. Then enable the cache for a small traffic slice with a visible bypass path. Keep the old path measurable for at least a week.
This is the same staged discipline used in LLM output evaluation: define the rubric before you optimise the metric. For user-facing text, add a check that the cached response still matches the current source material. For transactional workflows, consider caching only the preparatory work and regenerating the decision itself.
Common mistakes
The first is treating a semantic similarity score as permission to reuse an answer. “Can I change my address?” and “Can I change the delivery address after dispatch?” are close in wording but may have different policy outcomes. The second is omitting tenant and document versions from the key. The third is silently serving stale results without an age signal or retry route.
The fourth is routing a cacheable request to a cheaper provider and accidentally losing a large warm prefix. The fifth is using caching to mask a retrieval system that needs its own index and relevance work. If you have not measured the uncached baseline, you cannot say whether the cache saved anything meaningful.
Next step
Choose one bounded task, write down its key dimensions and invalidation event, then run a week of shadow measurements. When you are ready to compare saved spend with the cost of quality checks, AI cost optimisation and the verification overview provide the wider decision frame.
Frequently asked questions
Does LLM caching make answers less accurate?
Not inherently. It can make answers less appropriate when the cached result is stale, authorised for a different user, or semantically similar but not equivalent. Versioned keys, source-aware invalidation and quality monitoring are the controls that make the difference.
What is the best TTL for an LLM cache?
There is no universal TTL. Tie expiry to the relevant change event where possible, such as a document revision or entitlement update. Use a short conservative TTL only when the system cannot report that event.
Is semantic caching safe for support assistants?
Only for deliberately narrow, well-tested questions, and only with access scope and source versioning in the key. It is safer to cache retrieved public passages than an account-specific final answer.
How do I know whether provider prompt caching is working?
Inspect the provider receipt for cached-input or equivalent token fields, compare repeated calls with an identical prefix, and keep stable instructions and documents before variable user data. A generic hit-rate metric cannot prove a provider prefix was reused.