Cortexa AI Glossary · Trust, fakes, and safety
What is prompt injection?
From Cortexa Learn, by Cortexa Consulting. Last checked .
A phishing email fools people. This trick is written for the assistant reading your email.
A note for the assistant
You've probably seen an email pretending to be from your bank, written to fool you into clicking. Prompt injection is a similar trick with a different target. Someone hides instructions in a web page, an email, or a file, hoping an artificial intelligence (AI) assistant will read them and follow them in place of yours. The person reading might never notice. The assistant, reading every word, might.
One stream of words
Why does this work at all? A language model takes in everything as text: the instructions its makers gave it, your request, and the document you asked it to read. It all arrives in the same place, in the same kind of words. International Business Machines (IBM) explains the weakness this way: the model can't reliably tell the developer's instructions from other text, because both look the same to it. So a sentence inside a web page that sounds like an order can be mistaken for one.1
Direct and indirect
Security researchers sort it into two kinds.
- Direct: a person types the instructions straight into the chat, trying to talk the tool out of its rules.
- Indirect: the instructions are planted in something the assistant reads later, such as a web page or a shared document.
The indirect kind is the one to understand, because you're not the one typing. You only asked for a summary.24
Why agents raise the stakes
A chatbot that only writes text can be tricked into saying something odd. An assistant that can act is a bigger target. Picture one that reads your email and can also send email. A planted message could try to talk it into forwarding something private. IBM gives that kind of example: an assistant able to edit files and write emails being tricked into passing on documents. The tool hasn't turned against you. Someone else is pulling a lever it shouldn't have listened to.1
First on the list
The Open Worldwide Application Security Project (OWASP), a nonprofit that tracks software risks, keeps a top ten for apps built on large language models (LLMs). In its 2026 list, prompt injection is number one. It held the same spot on the list before. OWASP is also plain about the limit: it says no method today reliably prevents it.256
Layers, and fewer keys
So builders defend in layers. They give the assistant only the access the job needs, so a successful trick can do little. They mark outside content as untrusted data to read, never orders to follow. They screen what comes in. And they ask a person to approve anything risky, like sending, paying, or deleting. Anthropic's guidance for developers lists the same steps.23
Your part
You don't need to be a security expert to help. Notice which assistants can both read your things and act on them, and keep the pause for approval switched on. If an assistant suddenly proposes something you never asked for, stop and look before you say yes. Which of your connected tools can send things in your name?
Works cited
- IBM, "What is a prompt injection attack?" (checked )
- OWASP Gen AI Security Project, "LLM01:2025 Prompt Injection." (checked )
- Claude docs, "Mitigate jailbreaks and prompt injections." (checked )
- NIST, "AI 100-2 E2025, Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations" (2025-03-24) (checked )
- OWASP Gen AI Security Project, "LLM01:2026 Prompt Injection" (2026-08-04) (checked )
- Help Net Security, "OWASP 2026 LLM Top 10: 'The model will be fooled'" (2026-08-06) (checked )