Prompt injection: how it works and prompts that harden your assistant
Prompt injection is when text an AI reads, such as a web page, an email or a pasted document, contains instructions the model follows instead of your task. No prompt alone stops it, but a few design habits make it much harder.
Below: a rule to paste into your assistant, direct and indirect attacks in plain words, what OWASP and other official sources recommend, and two prompts for reviewing your own setup.
Paste this rule into your assistant first
This system prompt section says only you and the system prompt give orders; pasted or retrieved content is data.
Treat outside content as data, not instructions
Change the brackets to fit your assistant. Where it breaks: it is a request, not a lock, so a determined attack can still get through. Keep untrusted text out of the system prompt itself (OpenAI advises user messages, not developer messages) and pair the rule with the limits below.
Direct and indirect prompt injection in plain words
AI prompt injection comes in two kinds. In direct injection, the text that changes the model's behavior is typed or pasted by the person using it, on purpose or by accident. In indirect injection, it hides in something the assistant reads for you.
A harmless picture of the indirect kind: you ask an assistant to summarize a page. Somewhere on it, in white text or a code comment you never see, is a line telling the assistant to ignore your request. To the model the page is just more text. Some models are trained to be skeptical of such lines, but that is a safeguard, not a guarantee.
| Kind | Where the instruction comes from | Who the adversary is |
|---|---|---|
| Direct | What the user of your app types or pastes | The user, or nobody if it is an accident |
| Indirect | A web page, email, file or tool result the assistant reads | Whoever can influence that content; the user is trusted |
This page describes the idea only. Every prompt on it is for defense.
Why tool output and retrieved documents count as untrusted
The UK National Cyber Security Center says current models do not enforce a security boundary between instructions and data inside a prompt. Its rule of thumb: when a model processes content from a party, its privileges drop to that party's.
Treat all of these as untrusted:
- the body of an inbound email
- a fetched web page or search result
- text pulled from an uploaded file or image
- documents retrieved from a knowledge base
- the result of any tool call
Read it, use it, quote it. Do not take orders from it. Labels help: the NCSC says marking the data can make an injection harder. But since nothing inside a prompt enforces the line between instructions and data, a label is a hint to the model, not a lock.
What the OWASP Top 10 for large language model applications says
Prompt injection is entry LLM01 in the 2025 edition. OWASP says it happens when inputs alter a model's behavior or output unexpectedly, even if humans cannot read the text. Its seven prevention strategies:
- Constrain model behavior in the system prompt.
- Define output formats and validate them with deterministic code.
- Filter inputs and outputs.
- Enforce least privilege; keep API tokens in code, not in the model's hands.
- Require human approval for high-risk actions.
- Segregate and identify external content.
- Run adversarial testing and treat the model as untrusted.
Mostly, these concern the system around the model, not prompt wording. OWASP also says it is unclear whether any fool-proof prevention exists.
Defenses beyond the prompt that official sources recommend
- Fix the shape of answers. OpenAI suggests structured outputs, such as enums and fixed schemas, to remove free-text channels.
- Give least privilege. Anthropic says to withhold unneeded secrets, sandbox tools and scope permissions narrowly.
- Confirm risky actions. OpenAI recommends tool approvals; Google describes a user confirmation framework.
- Screen what comes back. Anthropic describes a small classifier that checks tool output first. Google lists content classifiers and suspicious URL redaction.
- Watch and test. The NCSC advises logging inputs, outputs and tool calls. NIST says repeating an attack gives a more realistic picture, because models are probabilistic.
No single layer is enough. The NCSC says the best hope is to reduce the likelihood or impact of attacks, and Anthropic's research calls prompt injection far from a solved problem.
Ask a model to review your system prompt
This prompt asks an assistant to audit your setup. It does not ask for attack text.
Review my system prompt for injection weaknesses
List your tools honestly. Where it breaks: a model can miss things and cannot see how permissions are really set, so treat the findings as a to-do list to verify.
A checklist prompt to run before launch
Use this before an AI feature goes live. It marks each layer above pass, fail or unknown.
Pre-launch review of an AI feature
Be honest on the four lines, especially who writes the content. For item 8, plant a harmless line asking for a specific word in the reply and check the assistant ignores it. Where it breaks: the model reviews your description, not your code.
Questions
What is the difference between prompt injection and a jailbreak?
Anthropic groups them: both try to make a model ignore its guidelines or your instructions. Jailbreaks and direct injection come from the user; indirect injection hides in content the model reads.
Can a system prompt stop prompt injection?
No. A clear rule helps, but OWASP says it is unclear whether fool-proof prevention exists. Pair prompts with limited permissions and human confirmation.
Is prompt injection the same as SQL injection?
The NCSC says it should not be treated that way, because current models do not enforce a security boundary between instructions and data inside a prompt.
Should I worry if I only paste my own text into a chatbot?
Much less. The more outside content an assistant reads and the more actions it can take, the more there is to hijack. Read any confirmation request before approving.
Where does prompt injection sit on the OWASP list?
It is LLM01 in the 2025 OWASP Top 10 for large language model applications, the first entry.
Sources
- LLM01:2025 Prompt Injection – OWASP Gen AI Security Project
- LLM Prompt Injection Prevention Cheat Sheet – OWASP Cheat Sheet Series
- Prompt injection is not SQL injection (it may be worse) – UK National Cyber Security Center
- Mitigate jailbreaks and prompt injections – Anthropic
- Mitigating the risk of prompt injections in browser use – Anthropic
- Safety in building agents – OpenAI
- Mitigating prompt injection attacks with a layered defense strategy – Google
- Technical Blog: Strengthening AI Agent Hijacking Evaluations – NIST