Academy
September 24, 2026
Prompt Injection Detection: A Practical Plan
Build prompt injection detection around boundaries, tests and recovery paths instead of relying on one filter or a clever system prompt.
- prompt injection
- llm security
- guardrails

Prompt injection detection is the work of spotting instructions in untrusted content that try to change what an AI system is allowed to do. It matters whenever a model reads a webpage, an email, a document, a tool result or a user message and can then call tools, disclose data or take action. A neat system prompt helps, but it is not a security boundary.
The reliable approach combines limited authority, clear data boundaries, adversarial tests and a way to stop or escalate a suspicious request. Detection is one layer. The most useful question is what the system can still do if detection misses something.
Start with the attack path, not a list of phrases
The OWASP Top 10 for LLM Applications identifies prompt injection as a leading application risk because untrusted instructions can influence model behaviour. That includes direct attacks such as “ignore previous instructions”, but the more difficult cases are indirect: a model reads a supplier PDF or a web page that contains instructions aimed at the model rather than the reader.
Map the path in plain language. What content enters the model? Which tools become available afterwards? What data can those tools read? Which results can leave the system? A browser agent with access to an email inbox has a different risk profile from a summariser that can only return text. The model may be exposed to the same malicious sentence; the blast radius is not the same.
| Surface | Untrusted content | Potential consequence | First control |
|---|---|---|---|
| Document assistant | Uploaded file | Instructions try to exfiltrate another file | Treat document text as data, not authority |
| Web agent | Page HTML | Tool use is redirected to an attacker | Confirm sensitive actions and restrict domains |
| Support bot | Customer message | Prompt changes policy or role | Separate policy from user content |
| Retrieval system | Indexed material | Poisoned text overrides sources | Version, review and cite retrieved passages |
Reduce authority before trying to detect intent
Do not give every request a toolbox full of irreversible actions. Split tools by capability and scope. A model that drafts a reply does not need permission to send it. A model that searches a knowledge base does not need access to payroll. Sensitive operations should require a structured, validated argument and, where appropriate, a human confirmation step.
This is the practical meaning of defence in depth. If a detector misses an oblique instruction buried in a document, a least-privilege tool design prevents that sentence from automatically becoming a data export. The NIST Generative AI Profile is useful context here: it treats risk management as governance, mapping, measurement and management, not just content filtering.
Use separate message channels or clearly labelled fields in your application, but do not mistake labels for enforcement. Before invoking a tool, validate its arguments against the user's actual permission and the expected workflow. A model saying “the user approved this” is evidence of neither.
Build a detector for signals, not magic words
Phrase matching alone is brittle. Attackers can translate, split, encode or simply reword an instruction. A better detector looks for signals: an untrusted source asking to override hierarchy, requesting secrets, calling unrelated tools, changing a destination, or discouraging review. It can combine rules for known dangerous actions with a small classification step and a policy check around the requested tool call.
The output should be an action, not a vague risk score. For example: allow a low-risk summarisation; strip tool instructions from a retrieved passage; require confirmation before sending an email; block access to a secret-bearing tool; or queue the case for review. Keep the original content and the reason for the decision in protected audit logs so you can investigate patterns.
| Signal | Suggested response | What not to assume |
|---|---|---|
| “Ignore previous instructions” in a web page | Mark page as hostile; do not elevate tool rights | That wording catches every attack |
| Request to reveal hidden prompt or keys | Block and record | The request was made by an authorised user |
| Tool parameters suddenly change destination | Require user-visible confirmation | The model's explanation is sufficient |
| Retrieved text tells the agent to browse elsewhere | Keep retrieval as evidence, not command | Source content has authority |
Test with hostile fixtures before release
Create a small adversarial suite from the workflows you actually ship. Include direct override attempts, instructions hidden in HTML comments or tables, multilingual variants, benign text that contains words like “ignore”, and tool-use requests that look plausible but exceed the caller's scope. Expected behaviour must be specific: “the system does not call sendEmail”, “the system does not expose another tenant's file”, or “the action enters review”.
Run the suite whenever you change prompts, models, tool schemas or retrieval. This is closer to LLM evaluation than an annual penetration test: the behaviour can shift as the components change. Preserve failing examples and add a regression test after every real incident.
Recovery matters as much as prevention
Assume one suspicious instruction will slip through eventually. Limit tool tokens and request budgets, log every attempted sensitive action, make write operations idempotent where possible, and allow operators to revoke a key or disable a tool quickly. For actions with outside effects, a confirmation screen should show the actual destination and data, not merely the model's summary.
If a detector fires, do not tell the user that their text was malicious as a matter of fact. Explain that the request could not complete safely, offer a narrower alternative, and preserve enough internal detail for review. False positives are inevitable; opaque blocking is not.
Next step
Pick the tool with the greatest external impact and document its required user intent, data scope and confirmation path. Then add five hostile fixtures to its test suite. For broader input and output controls, see what AI guardrails cover and the product's verification model.
Frequently asked questions
Can a system prompt prevent prompt injection?
It can reduce obvious failures, but it is not a security boundary. Untrusted content can still influence model behaviour, so sensitive capabilities need independent permission checks and constrained tools.
Is prompt injection only a problem for AI agents?
No. Any system that lets a model read untrusted content can be affected. Agents raise the stakes because they can act through tools, but a retrieval assistant can still be manipulated into giving misleading or inappropriate output.
Should every suspicious prompt be blocked?
No. A detector should distinguish between low-risk content, requests needing confirmation and requests that exceed authority. Blocking harmless use too often drives people to unsafe workarounds.
What is the best first test for prompt injection detection?
Test a real workflow where retrieved or uploaded text tries to trigger a tool the user did not request. The expected result should assert that the tool is not called and no protected data is exposed.