Glossary · term
General Scales for AI Evaluation
A proposal to move from benchmarks (task-specific, saturating, and prone to contamination) to universal scales that measure cognitive demand profiles and ability profiles. A set of roughly 18 rubrics covering a broad range of cognitive requirements allows AI performance on novel tasks to be predicted.
Debate2026Wave 3 · 2025–26Maturity: 2/5
Maturity rationale
single source, early stage
References
Author: Społeczność / Anonimowi