Atlas · GenAI 2026
AI Evaluation & Observability
24 skills · ontology graph below shows relations within this section.
What this domain covers
This edition groups 24 capabilities in AI Evaluation & Observability across 10 named categories. The inventory contains 17 concepts and 7 tools; 7 skills appeared in at least two of the three original research runs. The remaining entries stay visible with their lower coverage so a reader can distinguish taxonomy scope from research-system agreement.
Current category labels: Benchmarking · Debugging & Diagnostics · Evaluation Design · Evaluation Frameworks · LLM Testing · Monitoring & Drift · Observability & Tracing · Output Quality & Review · and 2 more
Frequent learning foundations
- Model Evaluation supports 5 mapped skills
- Retrieval-Augmented Generation supports 4 mapped skills
- LLM Evaluation Frameworks supports 3 mapped skills
- LLM Observability supports 3 mapped skills
- Statistical Inference supports 3 mapped skills
Skills in this section
Benchmark Analysis
Benchmarking
LLM Benchmarking
Benchmarking
Stochastic System Debugging
Debugging & Diagnostics
Agent Evaluation
Evaluation Design
LLM Evaluation Design
Evaluation Design
LLM-as-Judge
Evaluation Design
DeepEval
Evaluation Frameworks
LLM Evaluation Frameworks
Evaluation Frameworks
LLM Testing
LLM Testing
Data Drift
Monitoring & Drift
ML Monitoring
Monitoring & Drift
LLM Observability
Observability & Tracing
Langfuse
Observability & Tracing
AI Output Verification
Output Quality & Review
RAG Evaluation
RAG Evaluation
Ragas
RAG Evaluation
Hallucination Detection
Reliability & Hallucination
BERTScore
Benchmarking
BLEU
Benchmarking
ROUGE
Benchmarking
TruLens
Evaluation Frameworks
Evidently
Monitoring & Drift
LangSmith
Observability & Tracing
OpenTelemetry
Observability & Tracing