- Type
- AI security and safety phenomenon
- Emerged
- 2022–2023, with the first public LLMs
- Key techniques
- Role-play, encoding tricks, adversarial suffixes, many-shot and multi-turn attacks
- Countermeasures
- Safety training, input and output classifiers, red teaming
- Notable work
- Anthropic's Constitutional Classifiers programme (2025–2026)
- Related
- Prompt Injection, AI Red Teaming, AI Guardrails, AI Safety
- Type
- AI security and safety phenomenon
- Emerged
- 2022–2023, with the first public LLMs
- Key techniques
- Role-play, encoding tricks, adversarial suffixes, many-shot and multi-turn attacks
- Countermeasures
- Safety training, input and output classifiers, red teaming
- Notable work
- Anthropic's Constitutional Classifiers programme (2025–2026)
- Related
- Prompt Injection, AI Red Teaming, AI Guardrails, AI Safety
Background
Modern assistants are tuned with techniques such as reinforcement learning from human feedback to refuse dangerous requests, and providers rely on that refusal behaviour as a safety layer.[6] From the first wave of public chatbots in late 2022, users began probing the layer's edges. The best-known early example was the "DAN" ("Do Anything Now") family of prompts, which ordered the model to adopt a fictional persona unbound by its rules; researchers soon formalised the underlying failure modes, describing jailbreaks as products of competing training objectives and of safety training that generalises less broadly than model capability.[6] The concept has older roots in adversarial examples — inputs subtly altered to fool neural networks — documented as early as 2013.[7]
Techniques
Jailbreak methods range from the manual to the automated. Manual attacks exploit storytelling — asking for a harmful answer inside a screenplay, a game or a fictional "other AI" — or encodings such as ciphers that may slip past output filters.[1] Automated methods are more scalable. Researchers demonstrated in 2023 that universal adversarial suffixes — strings of seemingly random tokens optimised with gradient methods — could reliably unlock aligned models and transfer between them; later analysis showed such suffixes act as attention hijackers in the model's internal processing.[5][8] Many-shot jailbreaking, described by Anthropic in 2024, exploits long context windows by filling them with hundreds of fabricated question-and-answer examples and then following the pattern; multi-turn attacks such as Crescendo escalate gradually across a conversation, with each step innocuous on its own.[3][4]
Jailbreaking is distinct from prompt injection: injection attacks smuggle malicious instructions into data that an AI system reads, hijacking the application, whereas jailbreaking is a direct user-to-model attack aimed at bypassing the model's own conduct rules. The two are often combined in real attacks.[1]
Defences and Open Problems
Defences are layered: stronger post-training and safety datasets, input and output classifiers that screen traffic for attack patterns, adversarial training, and continuous red teaming in which dedicated testers try to break systems before attackers do. The most prominent current approach is Anthropic's Constitutional Classifiers, introduced in February 2025: guard models trained on synthetic data that filter exchanges against a written constitution of allowed behaviour. In automated evaluations they cut the success rate of advanced attacks from 86% to 4.4% while increasing refusals of harmless queries by only 0.38%.[1] A public red-team demonstration attracted 339 participants and more than 300,000 interactions — roughly 3,700 hours of collective effort — before one universal jailbreak was found; Anthropic paid bounties of US$10,000 and US$20,000 for successful universal attacks. An improved version, Constitutional Classifiers++, was published in 2026, reducing costs forty-fold while keeping the refusal rate on production traffic at 0.05%.[1][2] No deployed model is considered fully robust, however; providers treat jailbreak resistance as a moving target that governs decisions about which capabilities can be safely released.[1]
>See Also
For Malaysian organisations deploying AI assistants, jailbreak resistance is an emerging operational and compliance issue. Financial institutions operating under Bank Negara Malaysia's technology risk management framework are expected to security-test customer-facing systems, and breaches that expose personal data can trigger obligations under the Personal Data Protection Act 2010.[9][10] Malaysia's AI Governance Framework and its newly established AI Safety Institute both place emphasis on safe and responsible deployment — themes that matter as local models and Bahasa Melayu assistants multiply, since safety filters can behave differently across languages.
References
- ↑Anthropic. (2025, February 3). Constitutional Classifiers: Defending against universal jailbreaks. https://www.anthropic.com/research/constitutional-classifiers
- ↑Cunningham, H., et al. (2026). Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks. ICLR 2026. https://openreview.net/forum?id=eNvsH5Ye2V
- ↑Anthropic. (2024, April). Many-shot Jailbreaking. https://www-cdn.anthropic.com/af5633c94ed2beb282f6a53c595eb437e8e7b630/Many_Shot_Jailbreaking__2024_04_02_0936.pdf
- ↑Russinovich, M., et al. (2024). The Crescendo Multi-Turn LLM Jailbreak Attack. arXiv. https://arxiv.org/abs/2404.01833
- ↑Zou, A., et al. (2023). Universal and Transferable Adversarial Attacks on Aligned Language Models. arXiv. https://arxiv.org/abs/2307.15043
- ↑Wei, A., Haghtalab, N., & Steinhardt, J. (2023). Jailbroken: How Does LLM Safety Training Fail? arXiv. https://arxiv.org/abs/2307.02483
- ↑Szegedy, C., et al. (2013). Intriguing properties of neural networks. arXiv. https://arxiv.org/abs/1312.6199
- ↑MIT Press. (2025). Universal Jailbreak Suffixes Are Strong Attention Hijackers. Transactions of the ACL. https://direct.mit.edu/tacl/article/doi/10.1162/TACL.a.695/137193/
- ↑Bank Negara Malaysia. (2026). Official website. https://www.bnm.gov.my/
- ↑National Cyber Security Agency Malaysia. (2026). Official website. https://www.nacsa.gov.my/