Glossary · term
Generative Test-Set Contamination
A paper (arXiv:2601.04301, January 2026, 11 authors including Stella Biderman) showing that even a single copy of a generative benchmark in the pretraining data allows a model to achieve a loss lower than the "irreducible error" of training on an uncontaminated corpus.
Safety2026Wave 3 · 2025–26Maturity: 2/5
Maturity rationale
single source, early stage
References
Author: Społeczność / Anonimowi