AI Security Testing / FOCUSED SERVICE

AI Red Team Exercise

AI red teaming tests prompt injection, guardrail bypasses, and multi-turn manipulation against agreed objectives, with documented results and fixes.

WHAT THIS SERVICE ADDRESSES

The challenge behind the engagement.

An AI red team exercise answers a different question than a penetration test: not which categories fail, but whether a determined adversary reaches specific goals. We write objectives with you, usually a handful, such as revealing another customer's record or triggering an unapproved instruction. Rules of engagement cover the channels, off-limits content, spending limits, stop conditions, and whether your operators are told. We run single- and multi-turn campaigns that pressure the guardrails, plus chains that move from planted content through the model's tools to data leaving the system. Automated tools scale the attempts; the design and validation stay in our hands. You receive results objective by objective with attempt counts and a detection timeline. This is not a safety certification and does not replace your model vendor's testing.

WHEN THIS IS THE RIGHT FIT

For CISOs and product leaders whose assistant or agent handles money, regulated data, or customer trust. Run it after a penetration test has fixed the obvious issues, before a public launch, or when the board asks whether the guardrails hold against a persistent attacker.

THE WORK BEHIND THE SERVICE

What we do.
What you can use.

Write objectives and stop conditions

We agree written objectives and rules of engagement signed by an executive sponsor, run a legal and provider-terms check for the adversarial content involved, and match access to the adversary position: anonymous, customer, insider, or partner. A spend cap, a pause contact, and an out-of-band stop channel are in place before the first prompt is sent.

Run manual and automated campaigns

Each objective gets its own campaign: single- and multi-turn attempts to talk the model past its guardrails, pressure on the people who trust its output, and chains that move from planted content through its tools to data leaving the system. Automated tools scale the attempts; every success is confirmed by hand and counted against how many tries it took.

Report by objective and detection timeline

You receive a result per objective, reached, partial, or blocked, with the chain used, plus narratives with transcripts and attempt statistics. Guardrail findings are grouped by manipulation class, with a timeline of what your filters, monitoring, and people noticed and when. We retest closed paths and can leave a reusable prompt and objective harness for a regular cadence.

IS THIS THE RIGHT ENGAGEMENT?

Who it’s for.
When you need it.

BEST SUITED FOR
  • Organizations whose assistant or agent handles money, regulated data, or high customer trust
  • Teams that already ran a penetration test and remediated the obvious findings
  • Product leaders preparing a public launch of a high-visibility AI feature
  • Security programs that must show guardrails hold against a determined adversary
WHEN IT’S TIME TO ENGAGE
  • A public or high-profile launch of the AI feature is approaching
  • The board or a regulator asks whether guardrails resist a persistent attacker
  • Earlier testing is closed out and leadership wants the guardrails stressed next
  • The AI system just took on control of payments, records, or sensitive workflows
AGREED AROUND YOUR ENVIRONMENT

What the scope can include.

  • Written objectives such as revealing another customer's record or triggering an unapproved payment

  • Rules of engagement: channels, off-limits content, spending limits, real-data handling, and stop conditions

  • Single- and multi-turn attempts to talk the model past its guardrails

  • Pressure on the operators and reviewers who trust what the model produces

  • Chains from planted content through the model's tools to data leaving the system

  • Whether content filters, monitoring, and people noticed the campaign, and how quickly

TRANSPARENT PRICING

What it typically costs.
One rate: $150/hour.

Every engagement is priced by the hours it takes at one flat rate, with scoping, the work, and reporting, plus a retest of your fixes, included. Find the size closest to yours.

Small
$9,000–$13,500

Two or three attack goals against one AI assistant

About 60–90 hours
Mid-size
$16,500–$24,000

Four to six attack goals against an assistant with tools, including long multi-step conversations and a check on who noticed

About 110–160 hours
Large
$27,000–$39,000

Many attack goals across several AI systems, played from more than one attacker position

About 180–260 hours
WHAT MOVES THE PRICE
  • Number of attack goals and AI systems in play
  • Attacker positions to simulate, such as anonymous visitor, customer, or insider
  • Whether your monitoring and staff response are measured
  • Production versus test environment, and the content limits that apply
TYPICAL TIMELINE

3–8 weeks

Get a fixed quote for your scope

Ranges are planning estimates at $150/hour, not a quote. Your price is confirmed in writing after a scoping call, before any work begins.

TANGIBLE DELIVERABLES

What you take forward.

  • Objective-by-objective results: reached, partial, or blocked, with the route used for each
  • Campaign narratives with transcripts, attempt and success counts, and guardrail findings grouped by type
  • Detection timeline showing what filters, monitoring, and staff noticed and when
  • Fixes at the model, application, and process layers, a retest of closed paths, and an optional reusable objective set

Final coverage, deliverables, timing, and any retesting or implementation work are confirmed before the engagement begins.

SERVICE-SPECIFIC QUESTIONS

Before we get started.

How is this different from the LLM Application Penetration Test?

The penetration test is coverage-driven: it works through each way an AI application can be abused, at every entry point, and reports what fails. The red team exercise is objective-driven: it asks whether a persistent adversary can reach a short list of specific goals by any route, including multi-turn manipulation and pressure on the people around the model, and whether anyone noticed. Most buyers run the penetration test first, fix the findings, then use the exercise to stress the guardrails and the response that remain.

Will you generate harmful content in production?

Only inside the rules of engagement you sign. Prohibited content categories, the channels we may use, and whether production or staging is in scope are written down before the exercise, and we run a legal and provider-terms check for the adversarial content involved. Where an objective requires probing a refusal boundary, we use the least severe prompt that proves the point, log it, and stop at the agreed condition. Real customer data that surfaces is reported to your pause contact and not retained.

REFERENCE POINTS
START AT THE SOURCE

Let’s find your next move.

A focused conversation. A clear scope. A practical path to stronger security.

Let’s talk security