Atlas · GenAI 2026

LLM-as-Judge

LLM-as-judge methodology (calibration, bias control)

conceptPeak: 2025Evaluation DesignAI consensus: 1/3

Prerequisites

  • LLM-as-judge is one METHOD within automated evaluation — you need the broader evaluation context first

  • Calibrating a judge model and measuring inter-annotator agreement (vs. human judges) requires statistical skills

Recommended reference

Zheng et al. (2023) 'Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena' — NeurIPS; foundational paper establishing methodology and limitations

Notes from AI deep research

Anthropic Opus

Zheng (2023) MT-Bench. TW Radar: Assess. Wymaga kalibracji. Nie zastepuje human eval w krytycznych

OpenAI Deep Research

Kalibracja i kontrola biasu; nie zastępuje human eval [OA#75]

Related skills