Glossary

Cortexa AI Glossary · Trust, fakes, and safety

What is a jailbreak?

From Cortexa Learn, by Cortexa Consulting. Last checked .

"Chatbot goes rogue" makes a good headline. Read on, and there's usually a person doing the pushing.


The headline

Every so often you'll see a headline saying a chatbot went rogue and said something it shouldn't. It sounds alarming. But read a little further and there's usually a person in the story, someone who spent a long time coaxing the chatbot there on purpose. That coaxing has a name. It's called a jailbreak, and once you know how it works, those headlines read very differently.

Past the rules

A jailbreak is a trick a person uses to get a chatbot past its safety rules, so it produces something it was built to decline. The word is borrowed from phones, where jailbreaking meant getting around the limits a manufacturer set. International Business Machines (IBM) describes artificial intelligence (AI) jailbreaks as people exploiting weak spots to make a system bypass its guidelines, often with role-play scenarios.1

The rules it learned

Where do those rules come from? Much of it is training. People rated a model's answers, and it learned to decline some requests and handle others with care. Extra filters often sit around the model too. A jailbreak tries to get past them by keeping the request the same and changing how it's dressed.

Why a trick can work

The usual disguise is a story, a game, or a what-if. A chatbot is trained to be helpful and to follow the context of a conversation. It writes what seems to fit what came before. So a long, carefully built setup can pull it off course, one small step at a time, until a refusal no longer seems to fit the scene. IBM points to the same thing: chatbots are trained to be helpful and to understand context, and tricksters take advantage of that.1

Jailbreak or injection

It's easy to mix this up with prompt injection, so here's the difference.

  • In a jailbreak, the person typing is the one trying to break the rules.
  • In an indirect prompt injection, the person typing is trusted, and the attack hides in something the tool reads, like a web page or an email.

Anthropic's guidance for developers draws the line the same way.2

Patch and test

Labs treat jailbreaks like security holes. When one becomes known, they collect examples, train the model to resist that pattern, and add filters that spot it. Developers building on a model can add their own screens, and stop users who keep trying. When a chatbot declines something you ask, part of what you're seeing is that ongoing work. The tricks keep changing, so the patching keeps going too. The National Institute of Standards and Technology, a United States agency, lists jailbreaks among the known attacks on generative AI, with the defenses researchers use against them.23

Breaking it on purpose

Before a release, labs also hire people to try jailbreaks themselves. It's called red teaming: testing a system the way an attacker would, so problems are found and fixed before the public meets them. The next time you see "went rogue" in a headline, try asking: who was pushing it, and for how long?

Works cited

  1. IBM, "AI jailbreak: Rooting out an evolving threat." (checked )
  2. Claude docs, "Mitigate jailbreaks and prompt injections." (checked )
  3. NIST, "AI 100-2 E2025, Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations" (2025-03-24) (checked )