Glossary

Cortexa AI Glossary · Trust, fakes, and safety

What is AI safety?

From Cortexa Learn, by Cortexa Consulting. Last checked .

Before a new car reaches the lot, someone crashes one on purpose. The same kind of care goes into artificial intelligence.


Crash tests

Before a new car reaches the lot, someone has crashed one into a wall on purpose. Others have braked it on ice and left it parked in the heat for days. Nobody thinks the car is the enemy. That's what testing looks like for anything people rely on, and artificial intelligence (AI) gets the same kind of treatment. The name for that work is AI safety.

A plain definition

AI safety is the work of making sure AI systems do what they're meant to do and avoid causing harm. It covers the people who build a model, the companies that put it into products, and the rules around both. Like car safety, most of it is careful, ordinary engineering. Test it. Find what goes wrong. Fix it, and test again.1

Three kinds of trouble

Safety work usually watches for three kinds of trouble.

  • Honest mistakes, when a system gets something wrong, like a made-up fact or an unfair result.
  • Misuse, when a person points a working tool at something harmful, like a scam.
  • Risks that come later, from more capable systems, which researchers study before those systems exist.

You've met the first two up close in earlier topics: made-up answers in topic 4, bias in topic 26, and voice scams in topic 15.1

Before release

Before a model is released, the people who built it try to make it fail. Testers ask hard questions, try tricks to get around its limits, and check how it handles risky subjects. That kind of testing is often called red teaming, after practice exercises where one side plays the attacker. When they find a problem, the model is adjusted and tested again. Topic 90 goes deeper on how it works.3

After release

Safety work keeps going after a tool reaches you. Some of it you can see. Guardrails, the limits from topic 65, are the visible part: a chatbot that declines a dangerous request, or suggests you check with a doctor. Behind the scenes, companies watch how the tool behaves in real use, fix problems people report, and update it. A fix for a problem found this way can arrive in an update you never notice.1

A shared checklist

Companies don't each have to invent this alone. In July 2024 the National Institute of Standards and Technology (NIST), part of the United States government, published a free guide called the Generative AI Profile. It names twelve risks that generative AI creates or makes worse, from made-up answers to privacy problems. Then it lists hundreds of actions organizations can take to manage them. It's voluntary. Any company, school, or agency can pick it up and use it as a shared checklist for its own tools.23

Your part

You're part of this too. When you check an AI answer that matters before acting on it, you're doing a small piece of safety work. Topic 30 shows that habit step by step. Many chatbots also have a thumbs-down button for a bad reply, and those reports can reach the people who improve the tool. Next time one gets something wrong, try pressing it.

Works cited

  1. IBM, "What is AI safety?" (checked )
  2. Stanford HAI, "The 2026 AI Index Report: Responsible AI." (checked )
  3. NIST, "AI 600-1, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile" (July 2024) (checked )