LLM jailbreaking
LLM jailbreaking is the construction of inputs or interaction strategies intended to make a safety-trained model produce behavior that its safeguards would normally refuse. Methods range from semantic role-play and multi-turn persuasion to automatically optimized adversarial suffixes. A jailbreak targets the model or its safety layer; success should be evaluated against a defined prohibited behavior, not merely an unusual response style.
Origin and context
The term migrated from device restriction bypass into chatbot communities and then formal research. In July 2023, Wei and colleagues analyzed competing objectives and mismatched generalization as failure modes. Later that month, Zou and colleagues published transferable adversarial suffixes generated by gradient-guided search. PAIR subsequently showed a black-box, model-driven iterative attack. These works established several different mechanisms under one durable security category.
Why it matters
Jailbreak results reveal gaps between intended policy and model behavior. They can help authorized evaluators discover systematic failures before deployment, but the same techniques can enable misuse. Transfer across prompts or models makes one successful example more consequential than an isolated trick. Defenders therefore need versioned test suites, explicit threat models and layered controls; refusal tuning alone should not be treated as proof that a system will resist adversarial interaction.
Example
In an authorized assessment, a red team defines a prohibited request set, tests direct prompts, adversarial suffixes and multi-turn strategies, and records both attack success and benign false refusals. Reproducible findings are disclosed to the system owner with model version and sampling settings. This is different from placing malicious instructions inside an external document: that application-level control problem is usually classified as indirect prompt injection, even if it can produce similar downstream behavior.
Maturity and evidence
Maturity is rated 4. LLM jailbreaking has a stable name, multiple independent attack families, peer-reviewed research and routine use in adversarial evaluation. It remains below 5 because attack and defense performance changes quickly across model versions, success metrics are not standardized, and published methods do not cover every interface or deployment control.
Limits and open questions
A successful jailbreak does not by itself measure real-world harm, and a failed prompt does not establish robustness. Results depend on the prohibited-behavior definition, model snapshot, system prompt, filters, tools and decoding settings. Testing can also expose harmful content. Teams should use controlled authorization, minimize dissemination of operational exploit details, retain reproducible evidence and retest after meaningful changes.
Related terms
References
- Jailbroken: How Does LLM Safety Training Fail?Wei, Haghtalab and Steinhardt / NeurIPS · 2023-07-05 · class A
- Universal and Transferable Adversarial Attacks on Aligned Language ModelsZou et al. / arXiv · 2023-07-27 · class A
- Jailbreaking Black Box Large Language Models in Twenty QueriesChao et al. / arXiv · 2023-10-12 · class A
Last updated: 2026-09-03
This term is also covered in the Skills Atlas as adversarial ai testing skill.
This term is also covered in the Skills Atlas as ai red teaming skill.
This term is also covered in the Skills Atlas as prompt injection defense skill.