Glossary · term

Generative Test-Set Contamination

A paper (arXiv:2601.04301, January 2026, 11 authors including Stella Biderman) showing that even a single copy of a generative benchmark in the pretraining data allows a model to achieve a loss lower than the "irreducible error" of training on an uncontaminated corpus.

Safety2026Wave 3 · 2025–26Maturity: 2/5

Maturity rationale

single source, early stage

References

Author: Społeczność / Anonimowi