Glossary · term

Prompt Caching

Prompt caching is an inference optimization that reuses processing already performed for an identical or reusable prefix of model input. Stable material such as system instructions, tool definitions, examples, or long reference documents is placed before changing user content. When a later request matches the cached prefix under a provider's rules, the service can reduce repeated computation, latency, and input cost without changing the visible prompt content.

Agents2024-05-14Wave 1 · 2023Maturity: 4/5

Origin and context

Google announced context caching for Gemini 1.5 Pro in May 2024, saying the feature would let developers send large prompt components once. Anthropic announced prompt caching for Claude in August and described explicit cache breakpoints; OpenAI announced automatic prefix caching in October. Together these releases established a cross-provider product pattern, not a shared cache protocol: eligibility, pricing, retention, observability, and configuration remain provider-specific.

Sources: s1, s2, s3

Why it matters

Agent and retrieval applications often resend large, mostly stable prefixes on every turn. Avoiding redundant processing can make long instructions, many tool schemas, or repeated document context economically practical and more responsive. Prompt caching also changes prompt architecture: stable content should be grouped before request-specific data, and teams need telemetry that separates cached from uncached tokens. It is an efficiency feature, not extra memory or an accuracy technique by itself.

Sources: s1, s2, s3

Example

A contract assistant can place its system policy, output schema, and a reviewed agreement at the beginning of the prompt, followed by each new analyst question. Repeated queries against the same prefix may receive a cache hit. The application should monitor actual cache usage and invalidate assumptions when the agreement, tools, model, or provider configuration changes rather than treating yesterday's hit rate as guaranteed.

Sources: s1, s2, s3

How it differs

Semantic Cache

Prompt caching reuses model-side processing for a matching prompt prefix. A semantic cache typically reuses a prior answer or application result for a meaningfully similar request. Semantic reuse can change which response is returned; prompt caching still runs generation for the current request.

Context Engineering

Context engineering decides what information the model receives and how it is maintained. Prompt caching optimizes repeated processing of that context. Cache-friendly ordering can be one context-engineering tactic, but relevance and correctness take priority over cache hits.

Maturity and evidence

Maturity is rated 4 because multiple major providers documented prompt- or context-caching mechanisms across 2024, while Anthropic and OpenAI documented production API behavior and usage reporting. The rating applies to the optimization pattern, not to stable cross-vendor behavior: exact savings, thresholds, lifetimes, and controls can change with model and API versions.

Sources: s1, s2, s3

Limits and open questions

A cache hit requires provider-specific matching and eligibility conditions, so small prefix changes or low request reuse can erase the benefit. Caching does not expand the context window, improve weak evidence, or guarantee deterministic output. Sensitive content still needs the same data-governance review as any model input. Applications should not hard-code marketing-era discounts or retention assumptions; they should read current provider terms, instrument cache metrics, and benchmark end-to-end latency and cost.

Sources: s1, s2

Related terms

References

Last updated: 2026-08-27

In the Skills Atlas

This term is also covered in the Skills Atlas as prompt caching skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as context engineering skill.