OpenAI Model Spec
The OpenAI Model Spec is OpenAI's versioned public document describing intended assistant behavior, including objectives, instruction authority, safety boundaries, defaults, and ways to handle conflicts. It is both a behavior-design artifact and a possible audit target. It is not a model card, a complete list of product policies, a contractual guarantee, or proof that a deployed model will produce the specified response in every context.
Origin and context
OpenAI released the first public Model Spec on 8 May 2024 as an evolving account of intended model behavior. Dated snapshots from 11 April and 18 December 2025 record later states of the document; the December snapshot is the exact edition evaluated by an independent 2026 arXiv preprint. The official changelog identifies 18 August 2026 as the latest release reviewed for this page. Because the document's structure and rules change between releases, claims about its content or conformance must identify the snapshot rather than rely on the mutable root page.
Why it matters
A public behavioral specification makes normative product choices easier to inspect than scattered examples or refusal anecdotes. Developers can see how instruction levels and defaults are supposed to interact; researchers can turn individual rules into testable claims; and governance teams can compare published intent with observed behavior. The value is version-sensitive: when wording or hierarchy changes, an evaluation against one edition may not answer whether another edition is followed. The specification also separates desired model behavior from usage policies, deployment controls, and broader safety processes, which remain additional governance layers.
Example
An evaluator auditing instruction conflicts can select a dated Model Spec edition, extract the relevant authority rule, construct ordinary and adversarial multi-turn scenarios, and record whether a named deployed model follows that rule. The report should identify the model snapshot, system configuration, spec edition, elicitation method, and scoring procedure. A failure demonstrates a mismatch under those conditions; it does not by itself show that the entire specification is absent from training or that every deployment behaves identically.
How it differs
Constitutional AI
Constitutional AI is a family of training methods that uses written principles for critique, revision, or feedback. The OpenAI Model Spec is a particular vendor's behavioral specification. A specification can inform training or evaluation without being synonymous with the Constitutional AI method.
Deliberative alignment
Deliberative alignment is a method for teaching models to reason over explicit safety specifications. The Model Spec is the content artifact against which behavior may be trained or evaluated, not the post-training method itself.
Maturity and evidence
Maturity is rated 3 because the named document has multiple dated editions, documented use as a behavioral target, and independent audit research. The independent evidence is an arXiv-only preprint rather than a peer-reviewed final publication, and the specification remains explicitly evolving. The rating does not imply a cross-vendor standard or complete conformance by production models.
Limits and open questions
The specification is normative: it states desired behavior rather than directly measuring deployed behavior. Public editions may omit internal detail, and models, product layers, system instructions, tools, and policies can all change the observed outcome. Comparisons must pin both the spec edition and the tested system. The 2026 audit reports edition-relative results but cannot isolate specification-specific training from broader post-training improvements or evaluation awareness. Legal or safety conclusions should therefore use applicable policy and law in addition to the Model Spec.
Related terms
References
- Model Spec (2024/05/08)OpenAI · 2024-05-08 · class A
- Model Spec (2025/04/11)OpenAI · 2025-04-11 · class A
- How Well Do Models Follow Their Constitutions?arXiv · 2026-05-22 · class B
- Model Spec (2025/12/18)OpenAI · 2025-12-18 · class A
- Model Spec (2026/08/18)OpenAI · 2026-08-18 · class A
- Model Spec changelogOpenAI · 2026-08-18 · class A
Last updated: 2026-09-05
This term is also covered in the Skills Atlas as ai guardrails skill.
This term is also covered in the Skills Atlas as model evaluation skill.
This term is also covered in the Skills Atlas as RLHF skill.