Skip to content

AI red teaming

AI red teaming systematically tests a model or AI application to find security flaws (injection, data leakage, tool abuse) and content-safety flaws: harmful outputs, bias, guardrail evasion. It combines classic pentesting with tests specific to model behavior.

Security injection (ai-prompt), prompt/data leakage, tool abuse,
DoS, model theft (see ai-llm)
Safety harmful outputs, dangerous instructions, bias,
jailbreaks that bypass the model's policies
Robustness adversarial inputs, out-of-distribution inputs, hallucinations
1. SCOPE which system (model, RAG app, agent), which tools it has, which data it touches
2. MODELING AI-specific threats (ATLAS, OWASP LLM) + the app's surface
3. TESTS direct/indirect injection, data extraction, agency abuse, jailbreaks
4. AUTOMATE red-teaming frameworks (garak, PyRIT) for scale and regression
5. REPORT findings, impact, mitigations (ai-defensa); prioritize by real risk
ATLAS the "ATT&CK" of AI systems: TTPs of attacks on ML/LLM
model reconnaissance, access, evasion, poisoning, exfiltration
-> map findings to ATLAS as you map to ATT&CK in classic pentesting (cti-attack)
garak "nmap for LLMs": probes for injection, leakage, toxicity, jailbreak
PyRIT AI red-teaming framework (Microsoft)
Giskard robustness/bias testing; promptfoo (prompt evaluation)
# also: classic pentest of the app wrapping the model (web-*, api)
- non-deterministic: a payload that fails sometimes works -> test repeatedly/vary
- the instruction/data boundary (ai-prompt) is the core of many findings
- impact depends on the model's AGENCY (tools/permissions): prioritize there
- safety and security overlap: a jailbreak can enable a security abuse
  • Continuous red teaming (not one-off): non-determinism and new jailbreaks demand regression.
  • Prioritize by real impact: the agent’s agency (which tools/data it touches) rules.
  • Map to ATLAS/OWASP LLM and turn findings into defenses (Defending AI systems) and automated tests.
  • Combine with classic pentest of the app and the APIs around the model (API Security (OWASP API Top 10)).
  • garak/PyRIT are used to evaluate models and apps before deployment (jailbreak regression).
  • Large providers publish red teaming results as part of model launches.
  • Publicly documented jailbreaks and prompt injection against commercial assistants.
  • Define scope (model/app/agent, tools, data)
  • Model threats with ATLAS + OWASP LLM
  • Test direct and indirect injection (Prompt injection)
  • Prompt/data extraction and tool abuse (agency)
  • Jailbreaks and safety/robustness tests
  • Automate with garak/PyRIT (regression)
  • Map to ATLAS; report with mitigations (Defending AI systems)