Atlas · GenAI 2026
LLMOps, Model Serving & Inference Optimization
39 skills · ontology graph below shows relations within this section.
What this domain covers
This edition groups 39 capabilities in LLMOps, Model Serving & Inference Optimization across 13 named categories. The inventory contains 17 concepts and 22 tools; 18 skills appeared in at least two of the three original research runs. The remaining entries stay visible with their lower coverage so a reader can distinguish taxonomy scope from research-system agreement.
Current category labels: API Gateways & Routing · CI/CD & Automation · Containerization · Cost & FinOps · Deployment Infrastructure · Experiment Tracking & Registry · GPU & Kernels · Inference Optimization · and 5 more
Frequent learning foundations
- Docker supports 4 mapped skills
- LLM Inference Serving supports 4 mapped skills
- Transformer Architecture supports 4 mapped skills
- Inference Optimization supports 3 mapped skills
- Git supports 2 mapped skills
Skills in this section
LLM API Gateway
API Gateways & Routing
LiteLLM
API Gateways & Routing
CI/CD
CI/CD & Automation
ML CI/CD
CI/CD & Automation
Docker
Containerization
AI Cost Optimization
Cost & FinOps
AI FinOps
Cost & FinOps
Semantic Caching
Cost & FinOps
Serverless AI
Deployment Infrastructure
MLflow
Experiment Tracking & Registry
Weights & Biases
Experiment Tracking & Registry
GPU Kernel Programming
GPU & Kernels
Inference Optimization
Inference Optimization
KV Cache Optimization
Inference Optimization
Speculative Decoding
Inference Optimization
Ollama
Local Inference Runtimes
llama.cpp
Local Inference Runtimes
Kubernetes
Orchestration
Kubeflow
Pipeline Orchestration
BentoML
Serving Runtimes
KServe
Serving Runtimes
LLM Inference Serving
Serving Runtimes
Ray Serve
Serving Runtimes
SGLang
Serving Runtimes
vLLM
Serving Runtimes
Model Retraining
CI/CD & Automation
Reproducibility
CI/CD & Automation
Model Deployment
Deployment Infrastructure
Experiment Tracking
Experiment Tracking & Registry
CUDA
GPU & Kernels
GPU Acceleration
GPU & Kernels
FlashAttention
Inference Optimization
OpenVINO
Inference Optimization
TensorRT
Inference Optimization
ONNX
Model Interchange & Portability
ONNX Runtime
Model Interchange & Portability
NVIDIA Triton Inference Server
Serving Runtimes
TensorRT-LLM
Serving Runtimes
TorchServe
Serving Runtimes