Cortexa AI Glossary · Trust, fakes, and safety
What are guardrails?
From Cortexa Learn, by Cortexa Consulting. Last checked .
If a chatbot ever told you "I can't help with that," you've met one. What guardrails are, and why there's never just one.
The rail on the mountain road
On a mountain road, the guardrail doesn't steer the car. You hardly notice it. It's there for the moment something goes wrong. Artificial intelligence (AI) tools have their own version, and you've probably bumped into one. If a chatbot has ever said it can't help with something, or an app blanked out part of an answer, that was a guardrail doing its job.
Checks around the model
Guardrails are the checks built around an AI system to catch bad inputs, bad outputs, and risky actions before they cause harm. Some live inside the model, in what it learned during training. Many sit outside it, as separate filters and rules that run before and after the model does its work. Think of the model in the middle, with checks on the way in and on the way out. None of this is exotic. It's the same kind of care that goes into the brakes on a car or the smoke alarm in a kitchen.1
On the way in
The first checks look at what arrives. A screen can flag a request that's clearly asking for harm, or a document carrying hidden instructions meant to trick the assistant. Developers often run a small, fast model just for this, sorting each request before the main model ever sees it. Most of your requests pass straight through. You'd never know the screen was there.2
In the training
Some of the guardrail is inside the model itself. During training, people rated its answers, and it learned to decline certain requests and to answer others more carefully. That's why a chatbot may refuse to give dangerous instructions even when no outside filter is involved. Training alone isn't enough, though. It can be talked around, which is why the outside layers exist.
On the way out, and on actions
After the model writes, more checks can scan the answer before it reaches you. They look for private details like phone numbers, harmful content, or claims a company doesn't allow. And when the AI can take actions, permissions limit what it can touch, and approval steps make it ask before sending, paying, or deleting. That last kind matters more each year, as assistants do more than talk.3
Slices with holes
No single check catches everything. Picture each one as a slice of Swiss cheese, holes and all. Something bad can slip through a hole in one slice. Stack several, with the holes in different places, and very little gets all the way through. The National Institute of Standards and Technology, a United States government agency, lists checks like filtering inputs and outputs and human review among its suggested ways to manage the risks of generative AI.34
The last layer
For anything that really matters, the final slice is usually a person. Someone reads it before it goes out, or approves the step before it happens. Topic 60 calls that keeping a human in the loop. So the next time an assistant says it can't help, try asking yourself which layer you just met. Was it the training, or a filter around it?
Works cited
- IBM, "What are AI guardrails?" (checked )
- Claude docs, "Mitigate jailbreaks and prompt injections." (checked )
- OWASP Gen AI Security Project, "LLM01:2025 Prompt Injection." (checked )
- NIST, "AI 600-1, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile" (July 2024) (checked )