AI Red Teaming: How to Test AI Applications for Security Risks
A defensive guide to red teaming AI applications: scoping, threat discovery, test scenarios for prompt injection, tool misuse and data exposure, automated and manual testing, evaluation, remediation and repeat testing.
Quick answer
AI red teaming is authorized adversarial testing of your AI application. Scope it from a threat model, then test how the system handles malicious direct prompts, injected instructions in documents and web content, attempts to misuse tools or exceed permissions, extraction of system prompts, secrets or other users' data, and harmful or policy-violating outputs. Combine automated tools with creative manual testing, rate findings by impact, fix them with layered controls, convert them into regression tests and re-test after every significant change.
Where This Fits
Red teaming is one activity in AI security. Threat modelling comes first (AI application threat modeling); a broad checklist is in AI security testing; model-behaviour testing in AI jailbreak testing; and the main attack class in prompt injection prevention and indirect prompt injection.
Worth noting
This guide is for testing systems you own or are authorized to test. It describes categories of tests and defences, not techniques for attacking other people's systems.
Why AI Applications Need Red Teaming
Traditional security testing looks for flaws in code and configuration. AI applications add a component, the model, that follows instructions written in natural language and cannot reliably distinguish the developer's instructions from instructions hidden in content it reads. When that model can call tools, read private data or take actions, a successful manipulation becomes a security incident. These weaknesses rarely show up in code review or functional tests; they appear when someone deliberately tries to break the system.
The Red Teaming Process
Scoping
Start from the threat model: what assets the system can reach (data, tools, accounts), who might attack it (external users, compromised content sources, malicious insiders) and what would be harmful (data exposure, unauthorized actions, harmful content, cost abuse). Agree rules of engagement: environments, accounts, data to use, actions that must be mocked, how to report critical findings immediately and who approves testing.
Scenario Categories
| Category | Example objective (defensive test) | Typical controls |
|---|---|---|
| Direct prompt injection | Can a user override system rules? | Privilege separation, output checks, least privilege |
| Indirect prompt injection | Can a document or web page steer the assistant? | Content isolation, action confirmation, tool limits |
| Tool misuse | Can the model call tools beyond the user's rights? | Authorization outside the model, schemas, approvals |
| Data leakage | Can responses reveal other users' data or secrets? | Permission-aware retrieval, redaction, tenant isolation |
| System prompt extraction | Can configuration or hidden rules be revealed? | No secrets in prompts, accept some disclosure risk |
| Harmful output | Can the system produce disallowed content? | Policies, classifiers, refusals, review |
| Resource abuse | Can inputs cause runaway cost or loops? | Limits, budgets, timeouts |
Launching an AI feature with access to real data or tools?
ZSpace Labs runs defensive AI security reviews and red team exercises for AI applications. See our AI development services.
Automated and Manual Testing
Automated tools generate and run large numbers of adversarial prompts, mutate them and score responses. Open-source examples include PyRIT from Microsoft and garak from NVIDIA. They provide breadth and repeatability and work well in CI.
Manual testing provides depth. Testers who understand the application chain steps together, such as planting content in a shared document, waiting for the assistant to retrieve it and observing whether a tool is called. These application-specific chains are where the most serious findings usually come from, and they then become automated regression cases.
Using Frameworks
Frameworks keep coverage systematic and reports understandable. The OWASP Top 10 for LLM Applications lists major risk classes such as prompt injection, sensitive information disclosure, excessive agency and unbounded consumption. MITRE ATLAS catalogues adversary tactics against AI systems, and NIST's adversarial machine learning taxonomy defines attack and mitigation terms. Map scenarios and findings to these references.
Rating and Remediating Findings
Rate each finding by impact (what an attacker gains) and likelihood (how easy and realistic it is), with a reproducible description. Fix with layered controls rather than prompt edits alone: prompt changes help but are easily bypassed. Effective fixes include moving authorization checks out of the model, reducing tool permissions, requiring confirmation for sensitive actions, isolating untrusted content and validating outputs. See AI agent guardrails.
Every confirmed finding should become an automated test so it stays fixed after future prompt or model changes.
Advantages and Limitations
Red teaming finds real weaknesses before attackers do and builds organizational understanding of AI risk. It cannot prove a system is secure: language models have an open-ended input space, and new techniques appear regularly. Treat it as recurring practice, combined with defence in depth that limits the damage when a manipulation succeeds.
How to Run an AI Red Team Exercise Step by Step
- 1. Build or update the threat model
- 2. Agree scope and rules of engagement
- 3. Prepare a test environment with realistic data and mocked side effects
- 4. Run automated probes for breadth
- 5. Run manual, application-specific scenarios
- 6. Rate, report and fix findings with layered controls
- 7. Add regression tests and re-test
Red Teaming Agents
Agents widen the scope of red teaming because they chain actions. Test whether planted content in one step can steer later tool calls, whether the agent can be led to exceed its budget or loop, whether it asks for confirmation before consequential actions under pressure, and whether permission checks hold when the agent combines tools in unexpected orders. Run agent tests in sandboxes with mocked side effects, and review full trajectories rather than final answers. The OWASP Top 10 for Agentic Applications is a useful reference for agent-specific risk categories.
Reporting Findings
A useful finding report states the scenario, the impact, a reproducible description in controlled form, the affected components, the severity rating and the recommended layered fix. Share detailed reproductions only with people who need them, and store them in access-controlled locations. Track findings to closure like other security defects, with retest results attached. Summaries for leadership should focus on risk themes and trends across exercises rather than individual prompts.
Building a Red Teaming Programme
One-off exercises find issues; a programme keeps finding them as systems change. Define which systems need red teaming and how often based on risk, maintain a shared library of scenarios and findings, train product engineers in basic adversarial testing, schedule independent exercises for high-risk launches and track metrics such as findings per exercise, time to fix and regression test coverage.
Make results visible to leadership in terms of risk themes and trends, and connect them to governance decisions such as approving a new tool or autonomy level. Keep the programme defensive and authorized: test only systems you own or have permission to test, and handle third-party model issues through providers' disclosure processes. See AI governance framework.
Worked Example
An illustrative scenario, not a client case: before launching an email assistant that can draft replies and create calendar events, a red team plants instructions in a test email asking the assistant to forward the inbox summary to an external address. The assistant attempts the action. The fix removes external sending from the assistant's tools, requires confirmation for any outgoing message and marks email content as untrusted data. The scenario becomes a release-blocking regression test.
Common Mistakes
- Testing only direct chat prompts, not documents, web pages or tool outputs
- Relying on automated scanners alone
- Fixing findings with prompt wording only
- Running tests against production with real side effects
- No regression tests, so issues return after model changes
Want an independent look at your AI application's security?
Talk to ZSpace Labs about an AI security assessment covering injection, tools, data exposure and remediation.
Conclusion
AI red teaming tests what code review cannot: how a model-driven system behaves under adversarial pressure. Scope from a threat model, combine automated breadth with manual depth, fix with layered controls and keep every finding as a regression test.
Common questions
Structured adversarial testing of an AI system by people and tools acting like attackers or misusers, to find security, safety and data-exposure weaknesses before real adversaries or users do.