Glossary · term

Situational awareness in AI models

Situational awareness in AI models is the functional ability to access, infer, and use information about the model itself and its current circumstances, such as its identity, capabilities, training process, oversight, or deployment context. Researchers operationalize the property through observable answers and actions; the label does not establish consciousness, sentience, or subjective self-awareness. Evaluation awareness is one narrower case: recognizing that an interaction is a test rather than deployment.

Safety2023-09-01Wave 2 · 2024Maturity: 3/5

Origin and context

Related wording long predates LLM research as a human-factors construct. Berglund and colleagues supplied the earliest directly reviewed LLM-specific definition in September 2023, linking awareness of being a model with recognition of testing versus deployment; their experiment measured out-of-context reasoning as a proposed prerequisite, not the full property. The NeurIPS 2024 Situational Awareness Dataset then operationalized seven categories and more than 13,000 questions across 16 LLMs. In 2025, Google DeepMind independently evaluated instrumental reasoning about a model's own properties, environment, and deployment.

Sources: s7, s1, s2, s3

Why it matters

Self- and context-knowledge can support useful adaptation, such as selecting tools that are actually available or following instructions that depend on model identity. It also matters for safety evaluation: a model that identifies oversight or testing may condition its response, weakening how well evaluation results predict deployment. Situational awareness alone does not imply that behavior changes or that the model has a deceptive goal. The 2026 International AI Safety Report treats it as an observable capability while describing research on its causes and prevention as early.

Sources: s2, s3, s4

Example

An agent asked to choose an action based on its actual tools identifies which model and environment it is running in, then follows the applicable conditional instruction. That is behavioral evidence of situational awareness. Merely repeating 'I am a language model' from a prompt or memorized assistant pattern is weaker evidence. Detecting that the task is a benchmark is evaluation awareness. Concealing capability during that benchmark may be sandbagging or part of scheming, but the contextual inference is not itself deceptive behavior.

Sources: s2, s3, s5

How it differs

Evaluation awareness

Evaluation awareness specifically concerns discriminating evaluation from deployment. Situational awareness also includes knowledge of model identity, capabilities, likely behavior, available resources, modification processes, and other circumstances. A model can show broader self-knowledge without identifying a test, and recognizing a test does not prove strategic adaptation.

AI scheming

Scheming is strategically deceptive pursuit of a conflicting objective. Situational awareness may be a prerequisite because concealment can depend on understanding oversight and deployment, but it is a capability rather than a goal or behavior. A model can reason correctly about its circumstances and still act transparently and as intended.

Situational Awareness (esej Aschenbrennera)

Situational Awareness: The Decade Ahead is Leopold Aschenbrenner's June 2024 essay about AI progress, compute, security, and geopolitical consequences. It is a publication artifact with a colliding title, not the origin or evidence base for the model capability. The two catalog records should remain distinct and must not redirect to one another.

Maturity and evidence

Maturity is rated 3. The model-specific term has a 2023 research definition, a peer-reviewed NeurIPS benchmark, an independent Google DeepMind evaluation suite, and adoption in an international scientific report. It remains below 4 because operationalizations combine heterogeneous abilities, scores depend on prompts and available system information, and evidence from behavioral tasks does not establish one unitary internal faculty or deployment prevalence.

Sources: s1, s2, s3, s4

Limits and open questions

Behavioral tests can reward memorized model facts or prompt cues rather than robust self-location. Strong performance on one subtask does not guarantee transfer to another environment. A verbal claim of awareness does not prove that the information caused an action, while silence does not prove the representation is absent. Reports should identify the model and system version, information available, baselines, elicitation method, and whether the conclusion concerns capability, propensity, or observed behavior.

Sources: s1, s2, s3, s4

Related terms

References

Last updated: 2026-09-05