Alignment tax
An alignment tax is an unwanted loss of capability, task performance, helpfulness, or efficiency associated with an intervention intended to make an AI system better follow human preferences or safety objectives. In empirical model research, the term usually refers to a measured regression relative to an appropriate base or pre-alignment model. In wider safety discourse it can also mean the competitive or resource cost of choosing a safer system, but that broader sense should be stated explicitly rather than assumed.
Origin and context
An Effective Altruism transcript published in April 2020 records Paul Christiano using alignment tax for the cost of insisting on an aligned rather than merely competent system. He said the abstraction or language might come from Eliezer Yudkowsky but was unsure, so the transcript does not establish coinage. A 2021 arXiv-only preprint from Askell and colleagues then used the term for possible model-performance losses from alignment interventions. NeurIPS papers in 2022 and 2023 applied it to measured regressions or drift accompanying alignment optimization and examined mitigations.
Why it matters
The term forces an evaluation to measure both the intended safety or preference gain and possible losses outside the optimized objective. Without that comparison, a high reward-model score or lower harmful-output rate can hide worse translation, reasoning, calibration, usefulness, or behavior on another distribution. Alignment tax is therefore a trade-off diagnosis, not an argument against alignment. It can motivate changes to data, objectives, regularization, evaluation coverage, or training procedure. It also cautions buyers and policymakers against assuming that safety and capability are always opposed: some interventions show little regression or improve both on the tested measures.
Example
A team compares a base model, a supervised model, and an RLHF model on safety evaluations, human preference judgments, and a fixed suite of unrelated capability tasks. The RLHF model improves preference and truthfulness but loses accuracy on several held-out benchmarks. The team can call the measured regression an alignment tax, report it by task and confidence interval, and test a mitigation. It should not publish one universal tax value: the result depends on the baseline, intervention, model scale, data, evaluator, and chosen tasks.
How it differs
Reward hacking
Reward hacking describes optimizing a proxy in a way that defeats its intended goal. Alignment tax describes collateral degradation associated with an alignment intervention. A training run may exhibit both, but capability loss does not by itself prove that the model exploited its reward.
Maturity and evidence
Maturity is rated 3 because the term has a documented 2020 public-use anchor, a 2021 language-model research anchor, and multiple independent, peer-reviewed applications at NeurIPS 2022 and 2023. It remains an informal umbrella rather than a standardized metric: papers operationalize the tax through different tasks, baselines, and forms of drift, and some uses extend beyond capability benchmarks into economic or organizational costs.
Limits and open questions
An observed regression may reflect evaluation noise, data mismatch, training instability, or a deliberate trade-off rather than an unavoidable property of alignment. Public benchmark scores can also miss safety benefits and real deployment costs. Comparisons should control model, compute, data, and decoding conditions where possible and should report which alignment target improved. The phrase must not turn a local result into a general law that safer systems are less capable; the reviewed studies include cases where the measured tax was small or mitigated.
Related terms
References
- A General Language Assistant as a Laboratory for AlignmentarXiv · 2021-12-01 · class A
- Training language models to follow instructions with human feedbackNeurIPS · 2022-12 · class A
- Language Model Alignment with Elastic ResetNeurIPS · 2023-12 · class A
- Paul Christiano: Current Work in AI AlignmentEffective Altruism · 2020-04-03 · class B
Last updated: 2026-09-05
This term is also covered in the Skills Atlas as RLHF skill.
This term is also covered in the Skills Atlas as reward modeling skill.
This term is also covered in the Skills Atlas as model evaluation skill.