The challenge behind the engagement.
An AI red team exercise answers a different question than a penetration test: not which categories fail, but whether a determined adversary reaches specific goals. We write objectives with you, usually a handful, such as revealing another customer's record or triggering an unapproved instruction. Rules of engagement cover the channels, off-limits content, spending limits, stop conditions, and whether your operators are told. We run single- and multi-turn campaigns that pressure the guardrails, plus chains that move from planted content through the model's tools to data leaving the system. Automated tools scale the attempts; the design and validation stay in our hands. You receive results objective by objective with attempt counts and a detection timeline. This is not a safety certification and does not replace your model vendor's testing.
For CISOs and product leaders whose assistant or agent handles money, regulated data, or customer trust. Run it after a penetration test has fixed the obvious issues, before a public launch, or when the board asks whether the guardrails hold against a persistent attacker.
What we do.
What you can use.
Write objectives and stop conditions
We agree written objectives and rules of engagement signed by an executive sponsor, run a legal and provider-terms check for the adversarial content involved, and match access to the adversary position: anonymous, customer, insider, or partner. A spend cap, a pause contact, and an out-of-band stop channel are in place before the first prompt is sent.
Run manual and automated campaigns
Each objective gets its own campaign: single- and multi-turn attempts to talk the model past its guardrails, pressure on the people who trust its output, and chains that move from planted content through its tools to data leaving the system. Automated tools scale the attempts; every success is confirmed by hand and counted against how many tries it took.
Report by objective and detection timeline
You receive a result per objective, reached, partial, or blocked, with the chain used, plus narratives with transcripts and attempt statistics. Guardrail findings are grouped by manipulation class, with a timeline of what your filters, monitoring, and people noticed and when. We retest closed paths and can leave a reusable prompt and objective harness for a regular cadence.
Who it’s for.
When you need it.
- Organizations whose assistant or agent handles money, regulated data, or high customer trust
- Teams that already ran a penetration test and remediated the obvious findings
- Product leaders preparing a public launch of a high-visibility AI feature
- Security programs that must show guardrails hold against a determined adversary
- A public or high-profile launch of the AI feature is approaching
- The board or a regulator asks whether guardrails resist a persistent attacker
- Earlier testing is closed out and leadership wants the guardrails stressed next
- The AI system just took on control of payments, records, or sensitive workflows
What the scope can include.
- 01
Written objectives such as revealing another customer's record or triggering an unapproved payment
- 02
Rules of engagement: channels, off-limits content, spending limits, real-data handling, and stop conditions
- 03
Single- and multi-turn attempts to talk the model past its guardrails
- 04
Pressure on the operators and reviewers who trust what the model produces
- 05
Chains from planted content through the model's tools to data leaving the system
- 06
Whether content filters, monitoring, and people noticed the campaign, and how quickly
What it typically costs.
One rate: $150/hour.
Every engagement is priced by the hours it takes at one flat rate, with scoping, the work, and reporting, plus a retest of your fixes, included. Find the size closest to yours.
Two or three attack goals against one AI assistant
About 60–90 hoursFour to six attack goals against an assistant with tools, including long multi-step conversations and a check on who noticed
About 110–160 hoursMany attack goals across several AI systems, played from more than one attacker position
About 180–260 hours- Number of attack goals and AI systems in play
- Attacker positions to simulate, such as anonymous visitor, customer, or insider
- Whether your monitoring and staff response are measured
- Production versus test environment, and the content limits that apply
Ranges are planning estimates at $150/hour, not a quote. Your price is confirmed in writing after a scoping call, before any work begins.
What you take forward.
- Objective-by-objective results: reached, partial, or blocked, with the route used for each
- Campaign narratives with transcripts, attempt and success counts, and guardrail findings grouped by type
- Detection timeline showing what filters, monitoring, and staff noticed and when
- Fixes at the model, application, and process layers, a retest of closed paths, and an optional reusable objective set
Final coverage, deliverables, timing, and any retesting or implementation work are confirmed before the engagement begins.
Before we get started.
How is this different from the LLM Application Penetration Test?
The penetration test is coverage-driven: it works through each way an AI application can be abused, at every entry point, and reports what fails. The red team exercise is objective-driven: it asks whether a persistent adversary can reach a short list of specific goals by any route, including multi-turn manipulation and pressure on the people around the model, and whether anyone noticed. Most buyers run the penetration test first, fix the findings, then use the exercise to stress the guardrails and the response that remain.
Will you generate harmful content in production?
Only inside the rules of engagement you sign. Prohibited content categories, the channels we may use, and whether production or staging is in scope are written down before the exercise, and we run a legal and provider-terms check for the adversarial content involved. Where an objective requires probing a refusal boundary, we use the least severe prompt that proves the point, log it, and stop at the agreed condition. Real customer data that surfaces is reported to your pause contact and not retained.
- OWASP GenAI LLM Top 10 2026
- OWASP Top 10 for Agentic Applications for 2026
- GenAI Red Teaming Guide
- MITRE ATLAS
- NIST AI 600-1: Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile
- Model Context Protocol Specification: Authorization (2026-07-28)
- Document-Level Access Control - Azure AI Search
