Glossary · term

Semantic Cache

A semantic cache stores a prior query representation together with its response and can reuse that response for a new query judged sufficiently similar in meaning. A typical LLM implementation embeds the new query, searches cached vectors, and returns a stored answer above a configured threshold; otherwise it calls the model and may cache the new result. Because the match is approximate, cache policy is part of answer correctness.

LLMOps2023-12Wave 2 · 2024Maturity: 3/5

Origin and context

The 2023 GPTCache paper described an open-source architecture with embeddings, similarity evaluation, vector storage, and response reuse. AWS later documented the same request-response pattern for managed vector search. Recent Apple research separates curated static answers from dynamically populated cache entries and studies verification near the similarity threshold, showing that production work is moving from simple nearest-neighbor lookup toward explicit quality policies.

Sources: s1, s2, s3

Why it matters

Repeated questions can trigger expensive model inference even when an acceptable answer already exists. A semantic cache can reduce calls and latency across paraphrases, especially for stable support or knowledge tasks. Unlike exact caching, it can also return the wrong answer when two requests look similar but differ by negation or another answer-changing detail. That makes hit precision and invalidation first-class evaluation targets.

Sources: s1, s2, s3

Example

An internal IT assistant stores a VPN-installation question, its embedding, and the approved answer. A paraphrased request can reuse that answer when its similarity score clears the chosen threshold; otherwise the application calls the model. The team tests the cache on paraphrases and answer-changing near-matches, measures incorrect reuse separately from miss rate, and refreshes entries when the underlying instructions change.

Sources: s1, s2, s3

How it differs

Prompt Caching

Prompt caching reuses model-side computation for a matching input prefix and still generates an answer for the current request. A semantic cache normally reuses a completed prior answer for a meaningfully similar request. The latter therefore introduces approximate answer-substitution risk.

Cache-Augmented Generation (CAG)

Cache-augmented generation preloads a bounded knowledge collection into model context and reuses its runtime state to answer new questions. It does not primarily retrieve a previous query's final answer. Semantic response caching and CAG can coexist, but they cache different artifacts and require different invalidation rules.

Maturity and evidence

Maturity is rated 3. The pattern has a peer-reviewed open implementation, managed-infrastructure documentation, and independent research on adaptive verification. Evaluation remains application-specific, and there is no standard for similarity thresholds, verification policies, invalidation, or acceptable mismatch cost.

Sources: s1, s2, s3

Limits and open questions

Embedding similarity is not equivalence. A cache can suppress a needed fresh model call or preserve an answer after its underlying information has changed. Very conservative thresholds may erase the economic benefit, while aggressive thresholds increase incorrect reuse. Operators should choose thresholds against task-specific quality targets, verify borderline matches, measure hit quality alongside latency and cost, refresh stale entries, and keep a miss or review path for uncertain or consequential requests.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-09-03

In the Skills Atlas

This term is also covered in the Skills Atlas as semantic caching skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as prompt caching skill.