Speculative decoding
Speculative decoding is an inference technique in which a cheaper draft process proposes several future tokens and a target language model verifies them in parallel. An acceptance-and-resampling rule can preserve the target model's output distribution while reducing the number of serial target-model calls. The draft process may be a smaller model, an n-gram method, or another supported proposer; it is an accelerator, not a new training objective.
Origin and context
Leviathan, Kalman, and Matias posted Fast Inference from Transformers via Speculative Decoding in November 2022 and later presented it at ICML 2023. Chen and colleagues independently posted speculative sampling in February 2023. Both works exploit the fact that scoring a short proposed continuation in parallel can cost roughly as much as producing one target-model token serially.
Why it matters
Autoregressive decoding is latency-bound because each token normally depends on the preceding one. When a fast proposer predicts tokens the target model often accepts, one verification pass can advance the sequence by several positions. This can improve inter-token latency without changing target weights. The practical gain depends on model pairing, hardware utilization, batch shape, acceptance rate, and the implementation's overhead.
Example
A small draft model proposes four tokens for the target model's next continuation. The target scores those positions together, accepts the longest valid prefix, and samples a correction at the first rejection under the exact algorithm. If all four are accepted, one expensive verification advances multiple tokens. If acceptance is low, drafting and verification may add work instead of reducing latency.
Maturity and evidence
Maturity is rated 3. The technique has two independent foundational formulations and is implemented in a major open inference engine with multiple proposer options. It remains below 4 because deployment benefits are not uniform, the ecosystem contains materially different variants, and current vLLM documentation still describes compatibility and performance constraints that require workload-specific benchmarking.
Limits and open questions
Distribution-preservation claims apply to exact acceptance schemes, not automatically to every approximate variant. A poorly matched draft model can lower acceptance, consume extra memory, or reduce throughput. Batching, quantization, pipeline parallelism, and sampling settings can change the result; vLLM documents unsupported combinations and cases without latency gains. Teams should benchmark end-to-end service metrics rather than repeat headline speedups from different hardware.
Related terms
References
- Fast Inference from Transformers via Speculative DecodingGoogle Research / arXiv · 2022-11-30 · class A
- Accelerating Large Language Model Decoding with Speculative SamplingDeepMind / arXiv · 2023-02-02 · class A
- Speculative DecodingvLLM · 2026 · class A
Last updated: 2026-09-05
This term is also covered in the Skills Atlas as speculative decoding skill.
This term is also covered in the Skills Atlas as llm decoding strategies skill.
This term is also covered in the Skills Atlas as inference optimization skill.