Glossary · term

AI safety cases

An AI safety case is a structured, evidence-backed argument that a specified AI system is sufficiently safe for a specified use and operating context. It connects a top-level safety claim to subclaims, assumptions, reasoning, evidence, counterevidence, and residual risk. The artifact is contextual and revisable: it is not a generic checklist, a declaration that a model is harmless, or a safety certificate issued merely by its author.

Safety2024-05-17Wave 2 · 2024Maturity: 3/5

Origin and context

Safety and assurance cases were used in safety-critical engineering before modern AI. Hawkins and colleagues compared assurance arguments with prescriptive software certification in 2013 and described how evidence supports safety claims. In May 2024, Google DeepMind's first Frontier Safety Framework used a safety case to set a robustness target for a deployment-mitigation level. In August, the UK AI Safety Institute defined AI safety cases and announced research sketches; Buhl and colleagues published a broader frontier-AI governance treatment in October.

Sources: s2, s4, s3, s1

Why it matters

Model evaluations are evidence, but a list of scores does not explain why the evidence covers a deployment's hazards, why mitigations should work, or which assumptions could fail. A safety case makes that reasoning inspectable. It can expose missing tests, weak links, changing operating conditions, and disagreements before a deployment decision. It also provides a structure for combining technical model evidence with access controls, monitoring, incident response, organizational processes, and limits on use.

Sources: s1, s2, s3, s4

Example

For an agent allowed to modify production code, a top claim might be that deployment risk is tolerable within a defined repository and permission boundary. Subclaims could cover dangerous capability, review coverage, rollback, and control-protocol resistance to attack. Evidence could include red-team results, audit samples, access-control tests, and incident drills. An unresolved counterexample or a change in tools should reopen the case rather than be hidden behind an old approval date.

Sources: s1, s2, s3, s4

How it differs

AI control

AI control supplies protocols and adversarial evaluations intended to limit an untrusted model. A safety case is the larger argument that may use those results as evidence, together with other claims about the system and organization. A control benchmark is therefore neither necessary nor sufficient evidence for every safety case.

LLM evaluations (evals)

Evals measure selected behaviors under a protocol. A safety case explains how selected measurements support a safety claim in a defined context and where they do not. Collecting eval scores without an argument, assumptions, and coverage analysis is not a complete safety case.

Maturity and evidence

Maturity is rated 3. The underlying assurance-case method is established in other safety-critical domains, and frontier-AI researchers and a government institute have published concrete definitions and work programs. AI-specific practice remains below 4 because evidence standards, review authority, templates, and treatment of rapidly changing models are unsettled, and current sketches are explicitly not high-confidence guarantees.

Sources: s1, s2, s3, s4

Limits and open questions

A polished argument can still rest on incomplete hazards, invalid assumptions, weak evidence, or conflicts of interest. Case structure does not create independent verification, and evidence can become stale when model weights, tools, users, or deployment boundaries change. Certification and regulatory approval are separate processes that may consume an assurance case but are not produced automatically by it. Important cases need independent challenge, countercases, versioning, and explicit decision ownership.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-09-07

In the Skills Atlas

This term is also covered in the Skills Atlas as ai risk management skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as llm evaluation design skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as ai red teaming skill.