Glossary · term

Frontier Models

Frontier models are foundation models at the leading edge of assessed capability whose dangerous capabilities could create severe public-safety risks. The category is contextual: it moves as capabilities, evaluations, safeguards, and the state of the art change. No universal compute, parameter, or benchmark threshold defines every frontier model. The phrase is therefore a governance category used to focus evaluation and oversight, not a fixed technical model class.

Regulation2023-07-06Wave 1 · 2023Maturity: 4/5

Origin and context

A July 2023 paper on frontier AI regulation defined the target as highly capable foundation models that could possess dangerous capabilities sufficient to pose severe risks to public safety. The Bletchley Declaration later brought frontier-AI language into a multinational policy statement. The 2026 International AI Safety Report continues to assess rapidly advancing general-purpose systems through evidence about capabilities, risks, and safeguards rather than presenting one permanent threshold.

Sources: s1, s2, s3

Why it matters

The category helps direct scarce testing, reporting, incident-response, and external-scrutiny capacity toward models that may enable unusually consequential misuse or loss-of-control scenarios. It also prevents every AI system from being treated as equally risky. Because the frontier moves and evidence is incomplete, organizations need documented evaluation criteria and review dates. A label alone neither proves danger nor demonstrates that suitable safeguards exist.

Sources: s1, s3

Example

A developer preparing a new general-purpose model might test cyber, chemical, biological, autonomy, and safeguard-evasion capabilities against a published evaluation framework. If results cross its stated escalation criteria, the developer can trigger stronger access controls, external review, deployment limits, and post-release monitoring. The assessment should name the evidence and threshold used; simply calling the newest model frontier-grade is not a risk assessment.

Sources: s1, s3

How it differs

GPAI / systemic risk

A frontier model is a moving policy and risk concept. A general-purpose AI model with systemic risk is a specific EU AI Act classification with legal tests and consequences. A model can be described as frontier in research or policy debate without automatically satisfying the EU classification, and the labels should not be used as synonyms.

Compute Governance

Compute governance is an umbrella of interventions that use computing infrastructure, measurement, or access as governance levers. Compute thresholds may help identify models for scrutiny, but they are instruments; they do not exhaust the capability- and risk-based meaning of frontier models.

Maturity and evidence

Maturity is rated 4. The term has an attributable definition, international policy adoption, and continued use in a major multi-country safety assessment. Its operational boundary is not standardized: organizations and jurisdictions select different capability tests, thresholds, and update cycles. That variability prevents treating frontier model as a universal legal or technical classification.

Sources: s1, s2, s3

Limits and open questions

Frontier labels can become circular, promotional, or stale. Compute can be measurable while remaining an imperfect proxy for capability; benchmark results may miss novel risks or be affected by elicitation and access conditions. Governance should combine model and system evaluations, deployment context, safeguards, and post-release evidence. Reviewers should also record uncertainty and avoid importing requirements from one legal regime into another by analogy alone.

Sources: s1, s3

Related terms

References

Last updated: 2026-09-07

In the Skills Atlas

This term is also covered in the Skills Atlas as ai risk management skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as model evaluation skill.