Third-party AI evaluations
A third-party AI evaluation is an assessment conducted by an organization or team outside the developer's evaluation function, under a defined scope, method, access arrangement, and reporting process. It can test capabilities, safeguards, security, social impacts, or claims about performance. Third-party describes the evaluator relationship; it does not by itself prove impartiality, methodological quality, or regulatory authority.
Origin and context
The UK government described independent external evaluation as an emerging frontier-AI safety practice in October 2023. The U.S. National Telecommunications and Information Administration then treated independent evaluation, audits, and red teaming as inputs to AI accountability in March 2024. In October 2024, the UK AI Safety Institute published lessons from conducting pre- and post-deployment evaluations, including access, testing-window, information-security, and capability-elicitation constraints.
Why it matters
Developers know their systems well but also select what to test and disclose. An external evaluator can bring different expertise, methods, incentives, and institutional accountability, and can challenge a developer's claims before or after deployment. Independence is multidimensional: funding, governance, test selection, system access, result ownership, and publication rights all matter. Naming a provider as external without disclosing these conditions is weaker evidence than a transparent evaluation mandate and protocol.
Example
Before a frontier model release, a government institute might receive controlled access to a checkpoint, run preregistered cyber and autonomy tasks, discuss elicitation with the developer, and report scoped findings. The report should identify the model version, tools, safeguards, access restrictions, test window, scoring method, and uncertainty. A vendor rerunning the developer's public benchmark without privileged access may still be external research, but it is a materially different evaluation arrangement.
How it differs
LLM evaluations (evals)
Evals are tests or measurement procedures and may be designed or run internally. Third-party evaluation identifies who conducts or governs the assessment. The same eval can be used in both settings, while independence depends on organizational and contractual conditions rather than the benchmark alone.
AI red teaming
Red teaming is an adversarial testing method that can be internal or external. A third-party evaluation may include red teaming alongside benchmarks, qualitative review, audits, or system-level tests. Neither term guarantees certification or comprehensive safety coverage.
Maturity and evidence
Maturity is rated 3. Independent evaluation is supported by policy from multiple governments and has been operationalized by public institutes. Practice remains below 4 because access terms, conflict safeguards, reporting rights, methodology, and decision consequences vary considerably; frontier-model evaluation science is itself developing, and results are often snapshots of a particular setup.
Limits and open questions
External status can coexist with financial dependence, developer-selected tests, limited access, short timelines, or publication restrictions. Evaluators may also miss system-level risks when they receive only an API or one checkpoint. Findings should not be summarized as verified safe. Reviews should disclose conflicts, access, elicitation, exclusions, confidentiality, versioning, and who decides what follows from the result; regulatory inspection and certification remain separate processes.
Related terms
References
- Emerging processes for frontier AI safetyUK Department for Science, Innovation and Technology · 2023-10-27 · class A
- Independent EvaluationsU.S. National Telecommunications and Information Administration · 2024-03-27 · class A
- Early lessons from evaluating frontier AI systemsUK AI Security Institute · 2024-10-24 · class A
Last updated: 2026-09-07
This term is also covered in the Skills Atlas as llm evaluation design skill.
This term is also covered in the Skills Atlas as model evaluation skill.