llm11
← Blog

Academy

September 25, 2026

Semantic Caching for LLMs: A Careful Guide

Use semantic caching without serving the wrong answer: keys, thresholds, tenant isolation, evaluation samples and a safe rollout plan.

A notebook and laptop representing a careful semantic caching evaluation
Photo by Pexels on Pexels

Semantic caching for LLMs tries to recognise that two requests mean roughly the same thing even when their wording differs. That can cut repeated model calls dramatically for a narrow FAQ or classification job. It can also produce the most convincing kind of wrong answer: a fluent response to a question the user did not quite ask.

The right mindset is not “how low can our similarity threshold go?” It is “which requests are interchangeable enough that reusing an answer is honest?” This guide lays out a controlled answer, with a cache key that respects the caller and source material, a threshold chosen from real examples, and a bypass route when freshness matters.

What semantic caching changes

An ordinary response cache uses an exact key. Change one word and it misses. A semantic cache embeds or otherwise represents the new request, looks for a sufficiently similar stored request, and returns the prior response if it finds one. That makes it useful for repetitions such as “Where can I download my invoice?” and “How do I get a VAT receipt?”.

It does not make the meaning test perfect. A user asking “Can I cancel my plan?” may need a different answer from “Can I cancel after renewal?”, despite sharing most of the same words. Intent, timeframe, account state and access rights can change the outcome. The cache therefore needs stronger boundaries than a vector similarity score.

Good first useWhy it worksAvoid at first
Public FAQ answersContent is reviewed and sharedPersonal account support
Fixed-label classificationOutput space is smallLegal or medical advice
Stable onboarding snippetsVersioned source can invalidateLive prices and availability
Repeated internal definitionsAudience is knownPermissions-sensitive search

Design the key before choosing a threshold

The candidate search may be semantic, but the filter must be exact. Put tenant, audience or role, locale, product version, content-corpus version and task version beside the semantic representation. A support answer from one organisation must never become a “close match” for another. A result created before a policy rewrite must stop qualifying when the policy version changes.

Treat access control as a database filter, not a prompt instruction. This is consistent with the NIST AI Risk Management Framework, which frames trustworthy AI as a lifecycle practice rather than a one-off model setting. The cache is part of that system. It needs the same ownership, logging and review as retrieval or generation.

One practical key might look like: tenant=acme | role=member | locale=en-GB | corpus=2026-09-25.3 | task=help-answer.v2. Similarity is only computed within that precise bucket. It costs a few extra bytes and saves a potentially serious data boundary mistake.

Pick a threshold with a labelled sample

Do not begin at the vector store's default threshold. Collect fifty to one hundred real, anonymised queries from the chosen task. For each candidate pair, ask a domain reviewer one plain question: “Would returning the first answer to the second question be correct and useful?” Mark yes, no or uncertain.

Then test several thresholds against that set. Count false reuse separately from missed opportunities. A 65% hit rate is not a success if one in twenty hits returns a policy answer with the wrong exception. In early deployments, favour precision: a miss simply causes a fresh model call; a false hit changes the answer without the user knowing.

OutcomeWhat it tells youAction
Approved reuseThe task is genuinely repeatableKeep as a candidate hit
False reuseSimilar wording changed the answerAdd a filter or raise threshold
Missed reuseSafe duplicate was generated againConsider a modest threshold change
UncertainReviewer needs more contextDo not cache that slice yet

The same review set becomes an evaluation asset. Re-run it when you change the embedding model, normalisation rules, source content or product policy. How to evaluate LLM output has a useful rubric-first approach for making that routine instead of anecdotal.

Keep the cached object small and inspectable

Cache the reviewed answer, citations and source version, rather than a huge hidden conversation. Store when it was generated, which task template produced it, and a reference to the source material that supported it. On a hit, log the matched question and similarity band internally. You need enough evidence to explain the answer to an operator without retaining more user text than necessary.

For retrieval-backed answers, it is often safer to cache the retrieval result or a document excerpt than the completed prose. The model can still adapt the final wording to the exact request while avoiding a repeated search. That trade-off gives away some latency savings in exchange for a lower chance of presenting the previous person's conclusion as this person's answer.

Freshness and invalidation are part of meaning

A semantic match is only meaningful against the same current facts. Bump a corpus version when a help article changes, invalidate a product-specific bucket when a feature ships, and use a short expiry for time-sensitive content. A cache with perfect query matching still fails if the underlying policy is yesterday's policy.

This is why a generic “24-hour semantic cache” can be dangerous. A stable definition may be fine for months; an incident-status explanation may be wrong within minutes. Make the source owner decide the freshness rule and expose cache age to support staff. If a user asks for the latest information, bypass by default.

A humane rollout

Run the matching logic in shadow mode first. It should record would-hit events without serving them. Review the first hundred candidate pairs with the domain owner, paying special attention to no answers, exceptions, dates and numbers. Next, serve only reviewed public answers, with a clear cache-bypass mechanism for staff.

Measure P95 latency, model tokens, correction or reopen rate, and false-hit rate from the labelled sample. Do not use hit rate as the north-star metric. A cache that correctly answers twenty per cent of repeated requests is valuable; a cache that answers seventy per cent while quietly being wrong is not.

Next step

Choose one stable public question set and make a labelled pair file before adding any live cache. Pair that work with LLM caching in production for expiry and provider-prefix decisions, and use the verification approach for high-stakes responses that should remain fresh.

Frequently asked questions

Is semantic caching the same as RAG?

No. RAG retrieves current source material to inform a new answer. Semantic caching reuses a prior result or retrieval outcome because a new query appears equivalent. They can be combined, but they have different failure modes and invalidation needs.

What similarity score is safe for an LLM cache?

There is no portable safe number. Scores depend on the embedding model, task and source language. Choose a threshold using labelled examples from your own requests, then monitor false reuse after release.

Should I cache user-specific answers?

Usually not as a first use case. If you do, isolate the cache by tenant and access scope, include the relevant account state in the exact filter, and make freshness and deletion behaviour explicit.

Can semantic caching reduce hallucinations?

It can reuse a well-grounded, reviewed answer, but it can also repeat an old or mismatched one. It is a performance technique, not a factuality control; retain grounding and verification checks for answers that need them.