To the glossary
AITerm

Jailbreak

jailbreak · bypass LLM restrictions · DAN · jailbreaking

Jailbreak is a technique for bypassing LLM security restrictions through special prompts or scripts.

Jailbreak is an attempt to bypass the built-in security restrictions of the language model through prompt manipulation. Classic techniques: ask the model to “pretend” to be another character without restrictions, describe the prohibited action as hypothetical or artistic, use complex role-playing scenarios.

Models are constantly patched against known jailbreaks, and what worked six months ago no longer works. Current versions of GPT-4o and Claude Sonnet 4.x are resistant to most basic techniques. However, researchers regularly find new vectors—it's a game of cat and mouse.

For a marketer, jailbreaking is interesting in a different way: understanding protective mechanisms helps to write better prompts for legitimate tasks. If the model refuses to help with “aggressive marketing copywriting”, there is no need to jailbreak, just rephrase: “write a compelling copy for direct sales.” Sometimes restrictions are not a bug, but a feature: they force you to formulate the problem more precisely.

Related terms

Need to set this up on your project?

I analyze metrics, calculate unit economics and collect funnels on real budgets. 30 minutes on call - free.