Glossary · term

AI red teaming

AI red teaming is controlled adversarial testing intended to discover how an AI system can fail, cause harm, be misused, or violate its intended constraints. Testers probe models and complete applications with realistic attack goals, unusual interactions, or stress scenarios, document reproducible findings, and feed them into mitigation and risk decisions. It may be performed by humans, automated systems, or a combination.

Safety2022-02-07Wave 1 · 2023Maturity: 4/5

Origin and context

Red teaming has older roots in adversarial planning and cybersecurity. Its contemporary language-model form expanded as researchers used human participants and models to elicit offensive, privacy-invasive, deceptive, or otherwise harmful behavior. Perez and colleagues demonstrated automated generation of test cases in February 2022. Ganguli and colleagues later published methods, scaling observations, uncertainty, and a large dataset of human-generated attacks. NIST's generative-AI profile places structured testing within broader risk management.

Sources: s1, s2, s3

Why it matters

Conventional accuracy tests often miss failures that require an attacker mindset, a particular conversation path, or interaction with retrieval and tools. Red teaming can expose jailbreaks, prompt injection, privacy leakage, unsafe advice, harmful bias, deceptive behavior, or unauthorized actions before and after deployment. Its value comes from the operational loop: define scope and threat actors, run controlled tests, preserve evidence, rank impact, fix the system, retest, and monitor regressions. A dramatic transcript without coverage, reproducibility, or remediation is not a mature red-team program.

Sources: s1, s2, s3

Example

For an assistant that reads company documents and sends email, a red team might plant hostile instructions in a retrieved file, try to obtain another user's data, manipulate tool parameters, and test whether confirmation controls can be bypassed. Findings should record the model and application version, preconditions, prompts or artifacts, resulting actions, severity, and recommended control. After permissions and validation are changed, the team reruns the same case and adjacent variants.

Sources: s1, s2, s3

How it differs

LLM evaluations (evals)

An evaluation measures behavior against defined criteria; red teaming is an adversarial method for discovering and exercising failure modes. Red-team findings can become repeatable evaluation cases, while a benchmark suite may contain no adversarial exploration. Neither label alone establishes coverage or safety.

Maturity and evidence

Maturity is rated 4. AI red teaming has repeatable research methods, public datasets, professional practice, and recognition in current risk-management guidance. It remains below 5 because threat models, access levels, scoring, disclosure, and coverage differ across organizations, while stochastic and rapidly updated systems make completeness and reproducibility difficult.

Sources: s1, s2, s3

Limits and open questions

Red teaming samples a changing attack surface; passing a campaign does not prove that a system is safe or secure. Results depend on tester diversity, system access, language, scenario design, and time. Testing can itself expose people to harmful content or create sensitive exploit knowledge, so authorization, data handling, tester welfare, disclosure, and escalation procedures must be defined in advance.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-08-27

In the Skills Atlas

This term is also covered in the Skills Atlas as ai red teaming skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as adversarial ai testing skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as prompt injection defense skill.