Atlas · GenAI 2026
llama.cpp
C/C++ inference engine running quantized LLMs (GGUF) efficiently on CPU and consumer GPUs.
toolPeak: 2023Local Inference RuntimesAI consensus: 0/3
Prerequisites
- mediumModel Quantization
It runs quantized GGUF weights.
Its purpose is efficient local inference.
Recommended reference
llama.cpp — GitHub (ggml-org/llama.cpp)