llm11
← Blog

Academy

September 25, 2026

LLM Caching in Production: What to Cache

Learn where LLM caching saves time and spend, when it creates risk, and how to measure a cache without hiding poor answers.

A person working at a laptop, illustrating LLM caching decisions in production
Photo by Pexels on Pexels

LLM caching is not a switch labelled “make it cheaper”. It is a decision to reuse a previous result, a reusable prompt prefix, or a piece of retrieval work instead of paying to repeat it. Done carefully, it reduces latency and provider spend. Done casually, it returns a stale policy answer, loses a provider-side prompt-cache hit, or makes a team believe a slow workflow is healthy because the dashboard only shows cache hits.

This guide separates the cache layers, shows what each can safely store, and gives a practical way to measure the trade-off. The proof is deliberately operational: inspect a real request receipt, compare the cached and uncached paths, and keep a route to bypass the cache when a user needs a fresh answer.

The five caches people mean by LLM caching

The word is overloaded. These layers solve different problems and should not share an expiry policy.

LayerTypical keyGood candidateMain failure mode
HTTP responseURL and request bodyPublic, deterministic lookupsServing one user's data to another
Application responseNormalised task plus versionRepeated classificationStale result after a policy change
RetrievalQuery, filters and corpus versionStable knowledge-base searchAnswer based on an old document set
Semantic cacheMeaningfully similar requestNarrow, repeatable questionsSimilar wording hides different intent
Provider prompt cacheExact reusable prefixLong system prompt or document prefixA route change misses the prefix hit

Start with the layer closest to a deterministic result. A product catalogue lookup with a versioned response can tolerate a normal response cache. A free-form support reply is a poor candidate unless the cached object is a reviewed answer to a tightly defined question. Semantic caching is useful, but it is never a substitute for permissions, retrieval filters or an output check.

Cache a decision, not an unbounded conversation

The safer cache unit is a small, named job: classify-refund-reason:v4, extract-invoice-fields:v2, or answer-handbook-question:2026-09. Put the task version, tenant or access scope, model family, input normalisation version and relevant source version in the key. That may feel fussy on day one. It is far cheaper than explaining why a revised employment policy still produced last week's answer.

A useful pre-flight question is: could two users with this key legitimately receive different answers? If the answer is yes because of account permissions, location, date, feature flag or document access, those dimensions belong in the key or the result should not be shared. Hashing a prompt does not solve authorisation; it just makes the unsafe key harder to read.

For a customer-support assistant, one team might cache the retrieved public help-centre passages for five minutes, but never cache the final account-specific response. For invoice extraction, the document checksum plus extractor version can be a durable key because the desired fields do not change once the file is fixed. The distinction matters more than the cache product you choose.

Prompt caching needs a stable prefix

Provider prompt caching is different from saving a completed answer. It reuses processing for an exact or near-exact prefix supplied again to the provider. It is most valuable when a request starts with a long, stable system instruction, tool schema, or document bundle and ends with a short user-specific turn.

Put stable material first and variable material last. Do not inject a timestamp, random trace ID or changing account summary into the opening block if you expect the prefix to repeat. Our guide to prompt caching and routing explains the awkward consequence for a router: changing providers after a warm prefix can cost more than the model-price comparison suggests.

The practical check is simple. Log the provider's reported cached-input tokens separately from input tokens, then compare two otherwise identical calls. If the count never moves, inspect the literal request prefix before negotiating a larger caching budget. The OpenAI API reference also makes clear that API responses and model behaviour can change, which is one reason to pin your task and evaluation versions rather than cache an assumption indefinitely.

A sensible expiry is a product decision

Time-to-live should follow the thing that makes the answer invalid, not a generic “one hour” setting. A cache for a published release note can expire on a content revision. A cache for a support entitlement should expire when the entitlement changes. A cache for a question about “today's price” should normally be bypassed.

Use explicit invalidation where an owning system can emit it. When a handbook page is published, increment the corpus version. When a feature flag changes, bump the task version. When neither is possible, choose a conservative TTL and surface the cache age in internal logs. Expiry is not a compliance control; it is merely a fallback for systems that cannot say precisely what changed.

Measure the whole request, not just hit rate

High hit rate can be bad news if the cache is attached to the wrong work. Measure the following on cached and uncached requests:

MeasureWhy it mattersHealthy question
Hit rateShows reuse, not qualityWhich repeated task is it reducing?
P50 and P95 latencyCaptures the user-visible pathDoes the miss path remain acceptable?
Input and output tokensSeparates cache savings from model changesDid prompt-cache tokens actually rise?
Error and correction rateDetects stale or wrong reuseAre cache hits later reopened?
Bypass rateShows whether staff trust the cacheWhy are people asking for a fresh run?

Suppose a help centre answer is served in 180 ms from cache but is reopened by an agent twice as often as a fresh answer. The latency chart looks lovely while the operation gets worse. Pair cache metrics with the same groundedness or verification signal you use on fresh traffic. Caching an unverified output merely makes an error cheaper to repeat.

A small rollout that will not surprise the team

Begin with one task where the expected result is bounded and where a human can inspect a sample. Add a shadow lookup that records whether a key would have hit, but does not serve the cached answer. Review false matches and the source-version logic. Then enable the cache for a small traffic slice with a visible bypass path. Keep the old path measurable for at least a week.

This is the same staged discipline used in LLM output evaluation: define the rubric before you optimise the metric. For user-facing text, add a check that the cached response still matches the current source material. For transactional workflows, consider caching only the preparatory work and regenerating the decision itself.

Common mistakes

The first is treating a semantic similarity score as permission to reuse an answer. “Can I change my address?” and “Can I change the delivery address after dispatch?” are close in wording but may have different policy outcomes. The second is omitting tenant and document versions from the key. The third is silently serving stale results without an age signal or retry route.

The fourth is routing a cacheable request to a cheaper provider and accidentally losing a large warm prefix. The fifth is using caching to mask a retrieval system that needs its own index and relevance work. If you have not measured the uncached baseline, you cannot say whether the cache saved anything meaningful.

Next step

Choose one bounded task, write down its key dimensions and invalidation event, then run a week of shadow measurements. When you are ready to compare saved spend with the cost of quality checks, AI cost optimisation and the verification overview provide the wider decision frame.

Frequently asked questions

Does LLM caching make answers less accurate?

Not inherently. It can make answers less appropriate when the cached result is stale, authorised for a different user, or semantically similar but not equivalent. Versioned keys, source-aware invalidation and quality monitoring are the controls that make the difference.

What is the best TTL for an LLM cache?

There is no universal TTL. Tie expiry to the relevant change event where possible, such as a document revision or entitlement update. Use a short conservative TTL only when the system cannot report that event.

Is semantic caching safe for support assistants?

Only for deliberately narrow, well-tested questions, and only with access scope and source versioning in the key. It is safer to cache retrieved public passages than an account-specific final answer.

How do I know whether provider prompt caching is working?

Inspect the provider receipt for cached-input or equivalent token fields, compare repeated calls with an identical prefix, and keep stable instructions and documents before variable user data. A generic hit-rate metric cannot prove a provider prefix was reused.