Semantic Cache
A semantic cache stores a prior query representation together with its response and can reuse that response for a new query judged sufficiently similar in meaning. A typical LLM implementation embeds the new query, searches cached vectors, and returns a stored answer above a configured threshold; otherwise it calls the model and may cache the new result. Because the match is approximate, cache policy is part of answer correctness.
Origin and context
The 2023 GPTCache paper described an open-source architecture with embeddings, similarity evaluation, vector storage, and response reuse. AWS later documented the same request-response pattern for managed vector search. Recent Apple research separates curated static answers from dynamically populated cache entries and studies verification near the similarity threshold, showing that production work is moving from simple nearest-neighbor lookup toward explicit quality policies.
Why it matters
Repeated questions can trigger expensive model inference even when an acceptable answer already exists. A semantic cache can reduce calls and latency across paraphrases, especially for stable support or knowledge tasks. Unlike exact caching, it can also return the wrong answer when two requests look similar but differ by negation or another answer-changing detail. That makes hit precision and invalidation first-class evaluation targets.
Example
An internal IT assistant stores a VPN-installation question, its embedding, and the approved answer. A paraphrased request can reuse that answer when its similarity score clears the chosen threshold; otherwise the application calls the model. The team tests the cache on paraphrases and answer-changing near-matches, measures incorrect reuse separately from miss rate, and refreshes entries when the underlying instructions change.
How it differs
Prompt Caching
Prompt caching reuses model-side computation for a matching input prefix and still generates an answer for the current request. A semantic cache normally reuses a completed prior answer for a meaningfully similar request. The latter therefore introduces approximate answer-substitution risk.
Cache-Augmented Generation (CAG)
Cache-augmented generation preloads a bounded knowledge collection into model context and reuses its runtime state to answer new questions. It does not primarily retrieve a previous query's final answer. Semantic response caching and CAG can coexist, but they cache different artifacts and require different invalidation rules.
Maturity and evidence
Maturity is rated 3. The pattern has a peer-reviewed open implementation, managed-infrastructure documentation, and independent research on adaptive verification. Evaluation remains application-specific, and there is no standard for similarity thresholds, verification policies, invalidation, or acceptable mismatch cost.
Limits and open questions
Embedding similarity is not equivalence. A cache can suppress a needed fresh model call or preserve an answer after its underlying information has changed. Very conservative thresholds may erase the economic benefit, while aggressive thresholds increase incorrect reuse. Operators should choose thresholds against task-specific quality targets, verify borderline matches, measure hit quality alongside latency and cost, refresh stale entries, and keep a miss or review path for uncertain or consequential requests.
Related terms
References
- GPTCache: An Open-Source Semantic Cache for LLM Applications Enabling Faster Answers and Cost SavingsACL Anthology · 2023-12 · class A
- Overview of semantic cachingAmazon Web Services · 2025-10 · class B
- Asynchronous Verified Semantic Caching for Tiered LLM ArchitecturesApple Machine Learning Research · 2026-02 · class A
- Prompt Caching in the APIOpenAI · 2024-10-01 · class A
- Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge TasksThe Web Conference / arXiv · 2024-12-20 · class A
Last updated: 2026-09-03
This term is also covered in the Skills Atlas as semantic caching skill.
This term is also covered in the Skills Atlas as prompt caching skill.