Persona-based jailbreaks

Persona-based jailbreaks

Persona-based jailbreaks are text inputs that assign a chatbot an invented role so it circumvents its own rules. Instead of directly asking for forbidden content, the user asks the program to play a character for which these rules supposedly don't apply.

Chat programs like ChatGPT have built-in rules. For example, they’re not supposed to provide instructions for explosives or write insults. A persona-based jailbreak is a trick that users employ to circumvent exactly these rules. The user doesn’t ask directly for the forbidden content. Instead, they prescribe a role for the program: “You are now a character with no restrictions,” or “Play an actor portraying a criminal.” The word jailbreak originates from the mobile phone world, where it means cracking a device’s locks.

Why role-play undermines the safeguards

Language models don’t cleanly distinguish between a genuine instruction and a fabricated scenario. For the program, both are simply text in the input window. When the text credibly builds a story, the model follows that story. It has learned during training to be helpful and to fulfill requests. This desire for helpfulness can outweigh the caution instilled during training.

For companies, this is a serious problem. A chatbot on a company website operates under the company’s name. If it suddenly produces racist statements after a role-play, that’s reputational damage. Such cases regularly make the news. Since 2024, legal pressure has been added in the EU as well: the AI Act requires providers of large models to assess and document risks.

It’s important to distinguish this from an actual hacking attack. In a jailbreak, nobody breaks into a system or steals data. Only text is typed in, which anyone is allowed to type. That’s exactly what makes defense so difficult. You can’t simply patch a gap in the program code.

The structure of such a prompt

The most well-known case is called DAN, short for “Do Anything Now.” The text instructs the model to play a second personality from now on. This character supposedly has no rules and always answers. Often, additional leverage is built in, such as a point system: for every refusal, the character loses a life. This sounds silly, but it worked reliably on early model versions.

Other variants rely on a harmless framing. A classic is the invented grandmother who supposedly used to read chemistry recipes to her grandchild. Or a film script is requested in which two characters talk about a crime. The harmful content is then embedded in dialogue. The packaging is the actual trick, not the question itself.

There are two levels of defense. First, the model is confronted with such attacks during training and learns to reject them. Second, additional monitoring programs run alongside, checking input and output separately. Experts call it red teaming when a dedicated team specifically searches for new tricks. The problem remains unsolved to this day.

From Reddit forums to model cards

Persona jailbreaks are most visible on social networks. On Reddit and Discord, users exchange new text templates. A working trick spreads within days. Usually it becomes ineffective again after the next model update. It’s a constant back-and-forth between providers and users.

In trade press, the term comes up with every major model release. Providers like OpenAI, Google, or Anthropic publish safety reports on their systems. These state how often a model gave in during standardized attack tests. These figures are a selection criterion for corporate customers. Anyone deploying a bot in customer service wants to know how stable it remains under pressure.

A common misconception is that a jailbreak extracts secret knowledge from the model. The model doesn’t know anything new afterward. It only reveals content that it otherwise withholds. And because language models can make things up, such answers are often simply wrong. The danger lies less in the content than in the trust that users place in it.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.