Superalignment
Superalignment is the research problem of making AI systems substantially more capable than their human supervisors reliably follow human intent and remain within acceptable constraints. It focuses on a capability-gap regime in which people may be unable to evaluate outputs or provide trustworthy direct supervision. It is a specialization of broader AI alignment, not a solved technique, a safety certification, or a synonym for OpenAI's former Superalignment team.
Origin and context
OpenAI introduced the label publicly in July 2023 while creating a team co-led by Ilya Sutskever and Jan Leike. Its four-year target and promise to dedicate 20% of secured compute were commitments of that program, not part of the concept's definition or proof of a solution. Axios reported that the separate team disbanded in May 2024 and its work was integrated elsewhere. The term outlived that unit: an ACL 2025 paper called capability-gap alignment the central problem of superalignment, and Anthropic researchers used the phrase again in 2026.
Why it matters
Many current alignment methods rely on people choosing better outputs, supplying labels, or judging whether a system behaved correctly. If a model exceeds its evaluator in a relevant domain, incorrect or strategically misleading work may look convincing. Superalignment organizes research questions about producing scalable supervision, validating generalization, interpreting internal processes, stress-testing models, and automating parts of alignment research. It is a research agenda and threat model, not evidence that superintelligence exists or is imminent.
Example
A researcher trains a stronger model using labels produced by a weaker model, then measures how much of the stronger model's latent performance is recovered. This is a weak-to-strong generalization experiment: a tractable proxy for one supervision difficulty, not a full solution to superalignment. Scalable oversight is the broader subproblem and method family concerned with obtaining reliable supervision when evaluators are weaker; Anthropic was studying it before OpenAI's 2023 Superalignment branding. The terms should therefore be linked, not treated as synonyms.
How it differs
Superintelligence
Superintelligence names a hypothetical capability level or system that broadly exceeds human cognitive performance. Superalignment names the safety and alignment problem posed when a system exceeds the people supervising it; discussing the problem does not establish that such a system currently exists.
Maturity and evidence
Maturity is rated 3. The exact term began as OpenAI's problem-and-program label, but it continued in independent peer-reviewed ACL work and Anthropic alignment research after the original team dissolved. It remains below 4 because definitions still function as an umbrella research agenda, the target capability regime is hypothetical, and existing benchmarks study simplified supervision gaps rather than demonstrating a robust general solution.
Limits and open questions
The label can blur three different things: a research problem, a portfolio of proposed methods, and a former OpenAI organizational unit. Reports should state which meaning they use. Success on weak-to-strong tasks does not by itself establish honesty, value alignment, out-of-distribution robustness, or supervision of arbitrarily more capable systems. Human intent and acceptable constraints are also contested and underspecified, so machine-learning techniques do not eliminate the governance, institutional, or sociotechnical choices embedded in the objective.
Related terms
References
- Introducing SuperalignmentOpenAI · 2023-07-05 · class A
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionICML 2024 / PMLR · 2024-07 · class A
- OpenAI's long-term safety team disbandsAxios · 2024-05-17 · class B
- How to Mitigate Overfitting in Weak-to-strong Generalization?Association for Computational Linguistics · 2025-07 · class A
- Automated Weak-to-Strong ResearcherAnthropic Alignment Science · 2026-04-14 · class A
- Measuring Progress on Scalable Oversight for Large Language ModelsAnthropic · 2022-11-04 · class A
Last updated: 2026-09-05