KV cache compression
KV cache compression is a family of inference-time techniques that reduce memory used by the key and value states retained for autoregressive attention. Methods may evict selected token states, cluster or pool them, or store them at lower precision. The goal is to support longer sequences or larger batches with an acceptable quality and latency trade-off; model weights are not being compressed by this operation.
Origin and context
H2O, published at NeurIPS 2023, framed KV-cache eviction around retaining recent tokens and attention heavy hitters. SnapKV, submitted in April 2024, selected clustered positions using attention patterns observed near the end of a prompt. Quantized caches became available in the Hugging Face Transformers interface as another branch of the same operational problem. These are separate methods, not releases of one standard or a feature originated by vLLM.
Why it matters
During token-by-token generation, cached keys and values avoid recomputing attention states for the entire prefix. That speed benefit consumes memory that grows with retained sequence length and active requests, so the cache can limit batch size or long-context serving before model weights do. Compression can exchange some precision, coverage, or extra processing for a smaller memory footprint. This makes it an LLM inference and operations concern, not a training category.
Example
An inference team serving long documents profiles an uncompressed baseline, then compares a low-precision cache with a token-eviction policy. It measures task quality, time to first token, inter-token latency, throughput, and peak memory at realistic concurrency. If quantization saves memory but adds conversion overhead on short requests, the service can enable it only for memory-bound long-context traffic instead of declaring one cache mode globally best.
How it differs
FlashAttention
FlashAttention is an IO-aware exact-attention algorithm that uses tiling to reduce memory reads and writes; its paper also describes a block-sparse extension. KV cache compression instead changes which inference states are retained or how precisely they are stored. A serving stack may combine them, but an efficient attention algorithm does not by itself compress every retained key and value.
Speculative decoding
Speculative decoding reduces serial target-model decoding work by drafting and verifying tokens. KV cache compression targets memory occupied by attention state. Either technique can affect latency and memory, but their mechanisms and failure modes are different.
Maturity and evidence
Maturity is rated 3. The field has multiple peer-reviewed or public research methods and a maintained library implementation, establishing more than a one-paper idea. It remains below 4 because methods cover different operations, hardware and model support varies, and quality, memory, and latency trade-offs require workload-specific evaluation.
Limits and open questions
Eviction can discard states that later become important; quantization can introduce error and may worsen latency when memory is not the bottleneck. Reported speedups depend on sequence length, batch size, model architecture, kernels, and hardware. Offloading a cache to CPU changes placement rather than necessarily compressing it. Teams should name the exact method and budget, and test generation quality as well as memory savings.
Related terms
References
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsarXiv; later NeurIPS · 2023-06-24 · class A
- SnapKV: LLM Knows What You are Looking for Before GenerationUniversity of Illinois Urbana-Champaign / arXiv · 2024-04-22 · class A
- Unlocking Longer Generation with Key-Value Cache QuantizationHugging Face · 2024-05-16 · class A
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessStanford University / arXiv · 2022-05-27 · class A
- Fast Inference from Transformers via Speculative DecodingGoogle Research / arXiv · 2022-11-30 · class A
Last updated: 2026-09-04
This term is also covered in the Skills Atlas as inference optimization skill.
This term is also covered in the Skills Atlas as llm inference serving skill.
This term is also covered in the Skills Atlas as long context modeling skill.