Atlas · GenAI 2026
vLLM
High-throughput LLM inference engine with PagedAttention
toolPeak: 2024Serving RuntimesAI consensus: 0/3
Prerequisites
No prerequisites.
Recommended reference
docs.vllm.ai — official docs; Kwon et al. (2023) PagedAttention paper
Notes from AI deep research
Anthropic Opus
TW Radar: Adopt! De facto standard self-hosted LLM serving. Continuous batching, OpenAI-compatible API
Related skills
- → is an instance of: LLM Inference Serving(3/3)