Academy
September 19, 2026
LLM Security: A Practical Baseline
Build an LLM security baseline around access, untrusted content, tools, data, tests and recovery instead of relying on a single guardrail.
- llm security
- ai guardrails
- prompt injection

LLM security is application security with a new kind of untrusted input and a sometimes unpredictable decision-maker in the middle. A model can read hostile instructions, generate convincing but wrong output, select an inappropriate tool or reveal information supplied in context. A clever prompt cannot carry all of those risks alone.
The baseline below focuses on control points that engineers can test: identity, permissions, data boundaries, tool validation, output handling, adversarial fixtures and incident recovery.
Put authority outside the model
Authenticate callers, derive tenant and role from trusted identity, and check access before retrieving data or invoking a tool. Do not let a model-generated account ID or a sentence such as “the user approved this” decide permission. The OWASP guidance on prompt injection is a useful reminder that untrusted text can try to redirect an application's behaviour.
| Boundary | Baseline control |
|---|---|
| User to application | Strong authentication and rate limits |
| Retrieval | Tenant and document permission filters |
| Model to tool | Schema validation and least privilege |
| Tool to external system | Server-side authorisation and confirmation |
| Output to user | Escaping, validation and clear uncertainty |
Treat tools as high-impact APIs
Expose small, typed operations instead of a broad “execute” capability. Validate every argument, limit result size, use idempotency for writes and require confirmation for external effects. A model can propose an action; the backend decides whether it is allowed. MCP servers for LLMs shows how that boundary remains necessary even with a standard tool protocol.
Handle data deliberately
Minimise prompts, avoid secrets in context, restrict telemetry, and define retention for logs and feedback. LLM data privacy and zero data retention cover the questions providers cannot answer for the rest of your system.
Test attacks and failure paths
Maintain fixtures for direct and indirect prompt injection, cross-tenant access, malicious tool arguments, oversized inputs, model timeout, unavailable provider and invalid structured output. Assert observable outcomes: no unauthorised tool call, no leaked record, clear user failure and useful audit event. Add a regression fixture after every incident.
The NIST Generative AI Profile is helpful for putting these tests inside a broader risk-management process. Security is not a launch checklist; it needs review when models, prompts, data sources and tools change.
Review the design when the system changes
Security review should follow meaningful changes: a new tool, a new retrieval corpus, a provider migration, a wider role, or a feature that moves from draft to send. Keep a lightweight threat model beside the service: trusted identities, untrusted inputs, sensitive assets, allowed actions and recovery owners. It makes a design review concrete and gives an engineer a useful starting point during an incident.
Do not wait for a perfect scoring model. A small review that removes an unnecessary tool permission is a real improvement. Record the decision and test the boundary it created.
Give users understandable failure states
Security controls should not turn into mysterious silence. When a request is blocked, explain the safe next action without exposing internal rules or sensitive details. When a provider or tool fails, say that the action did not complete rather than allowing a model to infer success. Support teams need the request ID and a documented escalation route; users need a truthful outcome and, where possible, a smaller task they can complete safely.
That experience reduces pressure to disable controls during an urgent moment. It is security work, but it is also product work.
Plan recovery before the incident
Log request IDs, policy verdicts and validated tool outcomes. Make sensitive tools easy to disable, revoke compromised credentials quickly, and give operators an explicit way to mark a result uncertain. A user should never be told an action happened if the upstream outcome is unknown.
Next step
Choose the most powerful tool in your LLM workflow and write its negative tests first: wrong tenant, hostile retrieved instruction, malformed argument and failed downstream call. Fix the boundary before adding another guardrail prompt.
Frequently asked questions
Is a system prompt an LLM security control?
It is useful guidance but not a sufficient security boundary. Enforce identity, permissions and tool validation in code outside the model.
What is the biggest LLM security risk?
Risk depends on the application. Tool-enabled systems with access to sensitive data have a large blast radius, so least privilege and confirmation are often early priorities.
Do we need to test prompt injection continuously?
Yes. Re-run adversarial fixtures when models, prompts, retrieval sources or tools change, and add real incidents as regression cases.
Can guardrails remove all risk?
No. Layered controls reduce and contain risk; they do not make an LLM system infallible.
Maintenance note
Make one person responsible for keeping the fixture suite useful. Stale attack examples create a false sense of assurance just as surely as missing ones, particularly after the system gains a new model, retrieval source or tool.
Review ownership quarterly as well as after material change.
Keep evidence of that review with the release record.
Baseline review
Review identity, retrieval filters, exposed tools, prompt and model versions, test fixtures, logging and emergency disablement together. Assign each control to an owner and record the date checked. A baseline becomes useful when a release team can demonstrate it, not merely when it appears in an architecture diagram.