AI Jailbreaks: From Bypassing Restrictions to Governing Authority
AI jailbreaks are attempts to bypass models’ behavioral restrictions, but their significance increases when models operate tools and coordinate actions. This article presents a critical narrative review of the problem’s development, distinguishes jailbreaks, prompt injection and authority violations, and examines attacks, defenses and evaluation methods. Through the SGAEIA lens, it proposes an analytical synthesis with three levels: model response, authorized action and observed effect. This contribution is conceptual and non-normative; it does not demonstrate the architecture’s effectiveness or claim an unprecedented discovery. The conclusion is that behavioral robustness, authority boundaries and execution evidence should be evaluated together, with explicit assumptions and residual risk. **Keywords:** AI; jailbreak; prompt injection; autonomous agents; authority; governance; SGAEIA; assurance.
Authors
- Aridio Silva (ORCID: https://orcid.org/0009-0008-2411-6995)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-06
- DOI
- https://doi.org/10.5281/zenodo.23193016
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00