Atlas · GenAI 2026
Semantic Caching
Serving cached LLM responses for semantically-similar queries via embedding lookup to cut latency and cost (GPTCache).
conceptPeak: 2024Cost & FinOpsAI consensus: 0/3
Prerequisites
- hardEmbedding Models
Similarity is computed over embeddings.
- mediumVector Databases
Cached queries are indexed for nearest-neighbor lookup.
Recommended reference
GPTCache — GitHub (zilliztech/GPTCache)