Glossary · term

Superalignment

Superalignment is the research problem of making AI systems substantially more capable than their human supervisors reliably follow human intent and remain within acceptable constraints. It focuses on a capability-gap regime in which people may be unable to evaluate outputs or provide trustworthy direct supervision. It is a specialization of broader AI alignment, not a solved technique, a safety certification, or a synonym for OpenAI's former Superalignment team.

Safety2023-07-05Wave 1 · 2023Maturity: 3/5

Origin and context

OpenAI introduced the label publicly in July 2023 while creating a team co-led by Ilya Sutskever and Jan Leike. Its four-year target and promise to dedicate 20% of secured compute were commitments of that program, not part of the concept's definition or proof of a solution. Axios reported that the separate team disbanded in May 2024 and its work was integrated elsewhere. The term outlived that unit: an ACL 2025 paper called capability-gap alignment the central problem of superalignment, and Anthropic researchers used the phrase again in 2026.

Sources: s1, s3, s4, s5

Why it matters

Many current alignment methods rely on people choosing better outputs, supplying labels, or judging whether a system behaved correctly. If a model exceeds its evaluator in a relevant domain, incorrect or strategically misleading work may look convincing. Superalignment organizes research questions about producing scalable supervision, validating generalization, interpreting internal processes, stress-testing models, and automating parts of alignment research. It is a research agenda and threat model, not evidence that superintelligence exists or is imminent.

Sources: s1, s5, s6

Example

A researcher trains a stronger model using labels produced by a weaker model, then measures how much of the stronger model's latent performance is recovered. This is a weak-to-strong generalization experiment: a tractable proxy for one supervision difficulty, not a full solution to superalignment. Scalable oversight is the broader subproblem and method family concerned with obtaining reliable supervision when evaluators are weaker; Anthropic was studying it before OpenAI's 2023 Superalignment branding. The terms should therefore be linked, not treated as synonyms.

Sources: s2, s5, s6

How it differs

Superintelligence

Superintelligence names a hypothetical capability level or system that broadly exceeds human cognitive performance. Superalignment names the safety and alignment problem posed when a system exceeds the people supervising it; discussing the problem does not establish that such a system currently exists.

Maturity and evidence

Maturity is rated 3. The exact term began as OpenAI's problem-and-program label, but it continued in independent peer-reviewed ACL work and Anthropic alignment research after the original team dissolved. It remains below 4 because definitions still function as an umbrella research agenda, the target capability regime is hypothetical, and existing benchmarks study simplified supervision gaps rather than demonstrating a robust general solution.

Sources: s1, s3, s4, s5

Limits and open questions

The label can blur three different things: a research problem, a portfolio of proposed methods, and a former OpenAI organizational unit. Reports should state which meaning they use. Success on weak-to-strong tasks does not by itself establish honesty, value alignment, out-of-distribution robustness, or supervision of arbitrarily more capable systems. Human intent and acceptable constraints are also contested and underspecified, so machine-learning techniques do not eliminate the governance, institutional, or sociotechnical choices embedded in the objective.

Sources: s1, s2, s3, s4, s5, s6

Related terms

References

Last updated: 2026-09-05