Glossary · term
Eval Drift
Eval drift is the gradual loss of credibility of automated model evaluations. When the evaluating model (LLM-as-a-Judge) becomes too similar to the one being evaluated, their shared errors stop being caught, and the benchmark score overstates the actual quality. This forces periodic calibration and human review of the tests.
Safety2025/26Wave 2 · 2024Maturity: 1/5
Maturity rationale
single R2 source, 2025-26 neologism
References
Author: Weights & Biases