A prompt crafted to make a model ignore its own guidelines — usually through roleplay, hypotheticals, or encoding rather than a direct request.
Jailbreaks work because guidelines are learned behaviour rather than enforced rules. Enough framing — a fictional scenario, a claimed authority, an encoding the safety training did not cover — can shift which behaviour the model treats as appropriate.
It is worth distinguishing from prompt injection: a jailbreak is the user trying to unlock their own session, while injection is a third party hijacking someone else's agent through content it reads. Injection is the one that makes agents dangerous, because the victim never sees the instruction.