Atlas · GenAI 2026
LLM-as-Judge
LLM-as-judge methodology (calibration, bias control)
conceptPeak: 2025Evaluation DesignAI consensus: 1/3
Prerequisites
LLM-as-judge is one METHOD within automated evaluation — you need the broader evaluation context first
- mediumStatistical Inference
Calibrating a judge model and measuring inter-annotator agreement (vs. human judges) requires statistical skills
Recommended reference
Zheng et al. (2023) 'Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena' — NeurIPS; foundational paper establishing methodology and limitations
Notes from AI deep research
Anthropic Opus
Zheng (2023) MT-Bench. TW Radar: Assess. Wymaga kalibracji. Nie zastepuje human eval w krytycznych
OpenAI Deep Research
Kalibracja i kontrola biasu; nie zastępuje human eval [OA#75]
Related skills
- → is subcategory of: LLM Evaluation Frameworks(3/3)
- → is subcategory of: LLM Evaluation Design(2/3)
- → is part of: LLM Evaluation Frameworks(2/3)