LLM evaluations (evals)
LLM evaluations, commonly shortened to evals, are structured tests that measure how a model or AI system behaves on defined tasks, scenarios, and risk criteria. An eval specifies inputs, expected evidence or scoring rules, execution conditions, and analysis. It may use deterministic checks, human judgment, model-based graders, or several methods together. A benchmark score is one evaluation result, not the whole evaluation program.
Origin and context
Evaluation is older than modern language models, so no single organization invented evals. HELM, introduced in 2022 and published through TMLR in 2023, proposed transparent evaluation across many scenarios and metrics. OpenAI's public Evals repository, released in 2023, supplied a framework and registry for writing and running model tests. A broad survey submitted in July 2023 organized LLM evaluation around what, where, and how to evaluate.
Why it matters
Evals turn claims such as helpful, robust, or safe into testable criteria and make regressions visible before and after deployment. They can compare models, prompts, retrieval systems, tools, and policy changes on representative cases. Their value depends on coverage and governance: contaminated benchmarks, unrepresentative samples, weak graders, or silently changed test conditions can create false confidence. High-impact systems need multiple metrics, documented thresholds, error analysis, and human escalation.
Example
Before changing a support agent's model, a team can freeze a set of routine questions, adversarial requests, tool-call traces, and cases requiring refusal or escalation. It records exact model and prompt versions, runs automatic factual and format checks, asks trained reviewers to inspect ambiguous cases, and compares failure rates by scenario. A single public leaderboard number would not replace this deployment-specific suite because it does not test the team's tools, policies, or users.
How it differs
LLM-as-a-judge
LLM-as-a-judge is one possible grading method inside an eval. Evals are the broader practice: they define cases, metrics, protocols, baselines, and decisions. A suite can use judges alongside exact-match checks and human review, or use no model-based judge at all. The terms should therefore remain separate but strongly linked.
Maturity and evidence
Maturity is rated 4. LLM evaluation has public frameworks, independent taxonomies, transparent benchmark systems, and wide operational use. Core practices are established. It remains below 5 because suites often lack reproducibility across providers, metrics can conflict, benchmark contamination is difficult to detect, and there is no universal evaluation that predicts behavior across every deployment context.
Limits and open questions
An eval measures only the sampled behaviors and assumptions encoded in it. Test leakage, prompt sensitivity, judge bias, small samples, and repeated tuning against a fixed suite can inflate results. Offline tests may miss live interaction effects and rare harms. Teams should version datasets and graders, preserve raw outputs, report uncertainty, review failures qualitatively, and refresh suites when users, tools, models, or policies change.
Related terms
References
- OpenAI EvalsOpenAI · 2023 · class A
- Holistic Evaluation of Language ModelsStanford Center for Research on Foundation Models / arXiv · 2022-11-16 · class A
- A Survey on Evaluation of Large Language ModelsIndependent academic collaboration / arXiv · 2023-07-06 · class A
Last updated: 2026-09-03
This term is also covered in the Skills Atlas as model evaluation skill.
This term is also covered in the Skills Atlas as llm evaluation design skill.
This term is also covered in the Skills Atlas as agent evaluation skill.
This term is also covered in the Skills Atlas as llm evaluation frameworks skill.