Glossary · term

General Scales for AI Evaluation

A proposal to move from benchmarks (task-specific, saturating, and prone to contamination) to universal scales that measure cognitive demand profiles and ability profiles. A set of roughly 18 rubrics covering a broad range of cognitive requirements allows AI performance on novel tasks to be predicted.

Debate2026Wave 3 · 2025–26Maturity: 2/5

Maturity rationale

single source, early stage

References

Author: Społeczność / Anonimowi