Glossary · term

Mesa-optimization

Mesa-optimization occurs when a trained model itself implements an optimization process. The training procedure is the base optimizer and its training target is the base objective; the learned optimizer is the mesa-optimizer and the criterion it searches for is its mesa-objective. A model can perform sophisticated computation without meeting this definition: the claim requires evidence that it is conducting an internal search or optimization process.

Safety2019-06-05Wave 2 · 2024Maturity: 3/5

Origin and context

The 2019 arXiv-only preprint introduced mesa-optimization while analyzing risks from learned optimizers and the possibility that a learned mesa-objective might differ from the base objective. A 2023 arXiv-only preprint interpreted in-context learning as a learned optimization algorithm. The 2024 paper, accepted at NeurIPS 2024, reported transformers that internally estimate and apply task parameters in synthetic autoregressive tasks. These results study identifiable mechanisms, not persistent goals in deployed assistants.

Sources: s1, s2, s3

Why it matters

Training selects models by performance on a loss, but good performance does not uniquely determine the internal algorithm that produces it. If training yields a model that optimizes an internal objective, that objective may generalize differently from the loss outside the training distribution. The concept therefore separates two questions: whether a model contains an optimizer and whether its mesa-objective is aligned with the base objective. Evidence for the first does not automatically establish dangerous misalignment in the second.

Sources: s1, s2, s3

Example

Imagine a training process rewarding a model for succeeding in many simulated tasks. One learned solution could be a direct policy mapping observations to actions. Another could infer a task-specific objective, search over candidate actions, and choose the action that scores best under that inferred objective. Only the second is a mesa-optimizer. Reviewers would still need to identify what is optimized and test whether that criterion changes across environments before making an inner-alignment claim.

Sources: s1, s2, s3

How it differs

Reward hacking

Reward hacking is behavior that exploits a misspecified or imperfect reward signal. Mesa-optimization concerns an internal optimization process learned by the model. Either can occur without the other: a direct policy can exploit reward, and a mesa-optimizer can pursue a mesa-objective that remains aligned in the tested setting.

Mechanistic Interpretability

Mechanistic interpretability is a family of methods for studying internal computation. It may supply evidence about a proposed mesa-optimization algorithm, but it is not itself learned optimization. Behavioral success alone may also underdetermine the internal mechanism.

Maturity and evidence

Maturity is rated 3. The terminology has a stable primary definition in an arXiv-only preprint; a second arXiv-only preprint and a paper accepted at NeurIPS 2024 examine optimization-like transformer mechanisms in controlled tasks. The safety-relevant scope remains unsettled: definitions of search differ, empirical examples are narrow, and the evidence does not establish that deployed frontier models contain persistent mesa-objectives. The concept is established research vocabulary rather than an operationally measured prevalence claim.

Sources: s1, s2, s3

Limits and open questions

Calling every instance of in-context learning or planning mesa-optimization makes the term too broad to test. Researchers should specify the candidate search space, update rule, objective, and causal evidence for the mechanism. They should also separate an optimizer's existence from claims about deceptive alignment, scheming, or goal persistence. Current controlled demonstrations do not justify attributing hidden intentions to ordinary model outputs.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-09-04

In the Skills Atlas

This term is also covered in the Skills Atlas as mechanistic interpretability skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as mathematical optimization skill.