Glossary · term

LLM-as-a-judge

LLM-as-a-judge is an evaluation method in which a language model scores, ranks, or critiques other model outputs against instructions or a rubric. A judge may compare two answers, assign a numeric score, or explain defects. It can scale evaluation of open-ended responses that lack a single exact answer, but its verdict is a model output rather than objective ground truth.

Safety2023Wave 1 · 2023Maturity: 3/5

Origin and context

The 2023 MT-Bench and Chatbot Arena work established the current LLM-as-a-judge framing for evaluating chat assistants. The authors compared strong LLM judges with human preferences and documented position, verbosity, and self-enhancement biases. Later independent work asked whether fine-tuned open judges could generalize, and found that high in-domain accuracy can conceal weaknesses in fairness, adaptability, and out-of-domain evaluation.

Sources: s1, s2

Why it matters

A model judge supplies a repeatable scoring interface for open-ended answers where exact-match scoring is insufficient. It can support pairwise comparisons and rubric-based development checks. Scaling the number of judgments also scales any systematic preferences of the judge. The cited studies therefore make judge-human agreement and transfer beyond the evaluation setting important parts of interpreting a score, not optional evidence of universal reliability.

Sources: s1, s2

Example

In an illustrative development check, a team can compare paired answers from two prompts, swap their presentation order, and compare the judge's preferences with human judgments on the same examples. Disagreement is evidence to inspect, not something to hide by averaging scores. The example applies the papers' evaluation concerns; it does not establish that this workflow or judge will be reliable in another domain.

Sources: s1, s2

How it differs

LLM evaluations (evals)

LLM-as-a-judge is one grading method within an evaluation. The surrounding evaluation also determines which tasks and answers are sampled and what a score is intended to measure. A convincing model-generated judgment does not establish that the test represents the intended use.

Maturity and evidence

Maturity is rated 3. There is a defined evaluation framework and independent research testing judge generalization. These sources support an established research method, but do not by themselves establish broad operational adoption across organizations. The rating does not imply that a judge is an objective assessor or that agreement measured on one dataset transfers to another.

Sources: s1, s2

Limits and open questions

The originating study identifies position, verbosity and self-enhancement biases. The independent study finds that fine-tuned judges can perform well in-domain without matching a stronger judge's generalizability, fairness or adaptability. Those findings concern the evaluated models and tasks, not every possible judge. Judge-human agreement is useful evidence within its setting, but cannot turn model outputs into ground truth.

Sources: s1, s2

Related terms

References

Last updated: 2026-09-05

In the Skills Atlas

This term is also covered in the Skills Atlas as llm as judge skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as llm evaluation design skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as model evaluation skill.