Glossary

Cortexa AI Glossary · Trust, fakes, and safety

What is red teaming?

From Cortexa Learn, by Cortexa Consulting. Last checked .

The people paid to make a model misbehave, so you don't meet the problem first.


Paid to break in

Banks and other companies often hire security testers to try to break into their own systems. If there's a hole, they would much rather a friendly tester find it than a criminal. Artificial intelligence (AI) labs do something similar with their models, before you ever get to use them. The testers try to make a model fail or misbehave, so the problems get fixed first. That's called red teaming, and the testers are a red team.

Where the name comes from

The name is older than computers. During the Cold War, the United States military ran practice exercises where a "blue" team played its own side and a "red" team played the Soviet side. The point was to learn to think like the opponent. Computer security borrowed the idea for testing networks and software, and researchers at International Business Machines (IBM) describe AI as its newest field.1

What they look for

Red teaming an AI model means trying to get it to say or do the things it was trained not to. Testers try to trick it, push it past what it handles well, and ask for things it should refuse. Then they note what goes wrong: harmful or biased answers, made-up facts, and private information that leaks out. They also look for biases its makers didn't know were there.1

Who does it

Some red teamers work inside the company that built the model. Others are outside experts, brought in for what they know about a particular risk. And sometimes the public joins in. In August 2023, at a hacker conference in Las Vegas, more than 2,200 people spent two and a half days probing models from several AI companies, in an event supported by the White House. It was billed as the largest public exercise of its kind.3

What happens next

Finding a problem is half the job. When a red team turns something up, the builders fix it, often by training the model on new examples that show the right response, or by strengthening the guardrails around it. Then the testers go again, because a fix needs checking too. IBM's security team describes the work as a cycle that keeps going before and after a model is released.12

What it can't promise

A red team can show that problems exist. It can't prove that none are left. Testers find what they think to look for, in the time they have. Georgetown University's Center for Security and Emerging Technology calls red teaming useful and far from a silver bullet. It notes there are no agreed best practices yet, and that a single exercise can't track how a model changes over time. So it's one layer of safety work among several.4

What to look for

When a company releases a new model, it often publishes a report on its safety testing, sometimes called a system card, and red teaming is usually part of it. That's a good sign someone tried hard to break the model before you met it. Next time a big model launches, look for that report, and see what its red team found and what was done about it.56

Works cited

  1. IBM Research, "What is red teaming for generative AI?" (checked )
  2. IBM, "Testing the limits of generative AI: How red teaming exposes vulnerabilities in AI models." (checked )
  3. AI Village and partners, "More Than 2200 Participants Exchange More Than 165000 Messages With Leading Artificial Intelligence Large Language Models During the Generative Red Team Challenge" (Business Wire, 2023-08-29) (checked )
  4. Georgetown CSET, "Revisiting AI Red-Teaming." (checked )
  5. OpenAI, "GPT-4o System Card" (2024) (checked )
  6. Anthropic, "System Card: Claude Opus 4 & Claude Sonnet 4" (2025) (checked )