Cortexa AI Glossary · Trust, fakes, and safety
What is data poisoning?
From Cortexa Learn, by Cortexa Consulting. Last checked .
How a few planted examples can change what a model learns, and how teams catch them.
Fake reviews
A handful of fake reviews can drag a good restaurant's rating down, or prop up a bad one. A rating feels trustworthy because it comes from lots of people, and that's exactly what makes it worth faking. The data that artificial intelligence (AI) learns from can be tampered with in a similar way, by someone who plants bad examples where a model is likely to pick them up. It's called data poisoning.
What it is
A model learns its behavior from its training data, and a large language model learns from a huge amount of text, much of it gathered from the public web. Nobody can read all of it by hand. Data poisoning is someone slipping bad examples into that data on purpose, to change how the model behaves. International Business Machines (IBM) describes it as a type of cyberattack aimed at the data a model is trained on.1
Two kinds of harm
An attacker usually wants one of two things.
- To make the model worse for everyone, with wrong or junk examples that drag down its quality.
- To plant a hidden trigger, so the model behaves normally until it meets a particular word or phrase, and then does something else.
Security researchers call the second kind a backdoor. It can be harder to spot, because the model passes ordinary tests.13
Fewer than you'd expect
You might expect an attacker would need to control a big slice of the data. A 2025 study suggests otherwise. Anthropic, working with the United Kingdom's AI Security Institute and the Alan Turing Institute, found that about 250 planted documents were enough to install a hidden trigger in every model they tested. That held from the smallest, with 600 million parameters, to the largest, with 13 billion, even though the largest had learned from far more clean data.2
Two different attacks
Poison gets in wherever data is gathered without a close look: pages collected from the web, a shared dataset downloaded from the internet, or extra examples used to tune a model for one job. All of that happens while the model learns. A close cousin, prompt injection, works later on. It hides instructions in something a model reads while it's doing a task.3
The defenses
The defenses come in layers. Teams track where their data comes from, so a suspicious source can be traced and removed. They filter data before training, looking for examples that don't fit. And after training, they test the finished model for strange behavior, including red teams, testers who go looking for hidden triggers on purpose. No single check catches everything, which is why teams stack them.13
Where you come in
For most people, the defense is the same habit that protects you from fake reviews: notice where something came from. Where data comes from is a security question as well as a quality one. If your workplace trains or tunes its own AI tools, it's reasonable to ask where the training data came from and who checked it. Who at your work would know the answer?