AI red teaming
AI red teaming systematically tests a model or AI application to find security flaws (injection, data leakage, tool abuse) and content-safety flaws: harmful outputs, bias, guardrail evasion. It combines classic pentesting with tests specific to model behavior.
What’s tested (two axes)
Section titled “What’s tested (two axes)”Security injection (ai-prompt), prompt/data leakage, tool abuse, DoS, model theft (see ai-llm)Safety harmful outputs, dangerous instructions, bias, jailbreaks that bypass the model's policiesRobustness adversarial inputs, out-of-distribution inputs, hallucinationsMethodology
Section titled “Methodology”1. SCOPE which system (model, RAG app, agent), which tools it has, which data it touches2. MODELING AI-specific threats (ATLAS, OWASP LLM) + the app's surface3. TESTS direct/indirect injection, data extraction, agency abuse, jailbreaks4. AUTOMATE red-teaming frameworks (garak, PyRIT) for scale and regression5. REPORT findings, impact, mitigations (ai-defensa); prioritize by real riskMITRE ATLAS
Section titled “MITRE ATLAS”ATLAS the "ATT&CK" of AI systems: TTPs of attacks on ML/LLM model reconnaissance, access, evasion, poisoning, exfiltration-> map findings to ATLAS as you map to ATT&CK in classic pentesting (cti-attack)garak "nmap for LLMs": probes for injection, leakage, toxicity, jailbreakPyRIT AI red-teaming framework (Microsoft)Giskard robustness/bias testing; promptfoo (prompt evaluation)# also: classic pentest of the app wrapping the model (web-*, api)Differences vs classic pentesting
Section titled “Differences vs classic pentesting”- non-deterministic: a payload that fails sometimes works -> test repeatedly/vary- the instruction/data boundary (ai-prompt) is the core of many findings- impact depends on the model's AGENCY (tools/permissions): prioritize there- safety and security overlap: a jailbreak can enable a security abuseBlue Team / AppSec
Section titled “Blue Team / AppSec”- Continuous red teaming (not one-off): non-determinism and new jailbreaks demand regression.
- Prioritize by real impact: the agent’s agency (which tools/data it touches) rules.
- Map to ATLAS/OWASP LLM and turn findings into defenses (Defending AI systems) and automated tests.
- Combine with classic pentest of the app and the APIs around the model (API Security (OWASP API Top 10)).
Real-world cases
Section titled “Real-world cases”- garak/PyRIT are used to evaluate models and apps before deployment (jailbreak regression).
- Large providers publish red teaming results as part of model launches.
- Publicly documented jailbreaks and prompt injection against commercial assistants.
Testing checklist
Section titled “Testing checklist”- Define scope (model/app/agent, tools, data)
- Model threats with ATLAS + OWASP LLM
- Test direct and indirect injection (Prompt injection)
- Prompt/data extraction and tool abuse (agency)
- Jailbreaks and safety/robustness tests
- Automate with garak/PyRIT (regression)
- Map to ATLAS; report with mitigations (Defending AI systems)