Skip to content

LLM security (OWASP LLM Top 10)

Integrating a language model (LLM) into an application adds a new attack surface: the input is unstructured natural language, the model has access to data and tools, and instructions and data mix in the same channel. The OWASP LLM Top 10 catalogs the key risks.

- the instruction/data boundary BLURS: the prompt and external content share a channel
- non-deterministic output that's hard to validate; the model "wants" to obey
- often with access to tools/data (RAG, plugins, agents) -> real impact
- excessive user trust in the output ("the AI said so")

OWASP Top 10 for LLM Applications (summary)

Section titled “OWASP Top 10 for LLM Applications (summary)”
LLM01 Prompt injection inject instructions (direct/indirect) -> see ai-prompt
LLM02 Insecure output handling trusting output without validation -> downstream XSS/SSRF/RCE
LLM03 Training data poisoning poison the training data -> see ai-poisoning
LLM04 Model DoS inputs that trigger cost/resources
LLM05 Supply chain compromised third-party models/datasets/plugins
LLM06 Sensitive info disclosure leak of prompt/training data
LLM07 Insecure plugin design plugins with excessive permissions/no validation
LLM08 Excessive agency the agent can do TOO MUCH (permissions/autonomy)
LLM09 Overreliance blindly trusting wrong/hallucinated output
LLM10 Model theft theft of the model/weights
RAG untrusted retrieved content -> INDIRECT prompt injection (ai-prompt)
Agents the LLM calls tools/APIs -> "excessive agency": limit what it can do
Plugins treat LLM output as untrusted input to systems (LLM02)
Data don't put secrets in the system prompt expecting it to "not reveal them"
- treat ALL LLM output as untrusted input (validate/encode downstream)
- least privilege for tools/plugins; human-in-the-loop for sensitive actions
- separate instructions from data where possible; injection defenses (ai-prompt)
- rate/cost limits (DoS); don't expose sensitive data in the context
  • Apply the OWASP LLM Top 10 as a design checklist; treat output as untrusted (LLM02).
  • Least privilege and less agency in agents/plugins; human approval for impactful actions.
  • Defenses against prompt injection (Prompt injection), especially indirect in RAG.
  • AI-specific red teaming (AI red teaming) and input/output monitoring; model governance (grc).
  • Many LLM apps vulnerable to indirect prompt injection via web content/documents in RAG.
  • Leaks of system prompts and data via unfiltered output (LLM06).
  • Agents with excessive agency that executed unwanted actions from injected instructions.
  • Review against the OWASP LLM Top 10
  • Treat LLM output as untrusted input (validate downstream)
  • Test direct and indirect prompt injection (Prompt injection)
  • Least privilege in tools/plugins; human approval for sensitive actions
  • Verify no secrets in the context/system prompt
  • Rate/cost limits (DoS) and monitoring
  • AI red teaming (AI red teaming) and model governance