Atlas · GenAI 2026
Adversarial AI Testing
AI Red Teaming (proactive, automated)
conceptPeak: 2024Red TeamingAI consensus: 3/3
Prerequisites
Red teaming tests for prompt injection and other vulnerabilities — understanding the attacks is prerequisite for testing them
Red teaming uses automated evaluation to detect failures at scale — eval frameworks provide the testing infrastructure
Recommended reference
Perez et al. (2022) 'Red Teaming Language Models with Language Models' — foundational paper on automated red teaming; plus NIST AI 100-2e2025 'Adversarial ML' report
Notes from AI deep research
Anthropic Opus
Systematyczne szukanie failure modes. DeepTeam: 40+ klas podatnosci
OpenAI Deep Research
Procedury, scenariusze, raportowanie [OA#76]
Google Deep Think
Automatyczne pakiety atakujące [G#84]
Related skills
- → is subcategory of: AI Red Teaming(3/3)
- → is part of: AI Risk Management(3/3)