The challenge behind the engagement.
This penetration test is for a large language model (LLM) application, such as a chat assistant, copilot, or model-backed API. We attempt to override the model's instructions wherever input reaches it: the chat box, uploaded files, and form fields. We try the same through content the model reads later, such as documents, fetched pages, and ticket bodies, checking whether planted text makes the model act or leak. We probe for exposure of hidden instructions, disclosure of another user's data, unsafe handling of model output, and usage patterns that run up cost. We also test the application layer: sign-in, sessions, and whether each request is authorized. You receive transcripts, attempt counts, and fixes placed at the layer that closes each issue, under agreed limits.
For product and platform teams shipping a customer-facing assistant or an internal copilot. Book it before launch, before a customer security questionnaire asks about prompt injection, or after adding tools or file upload to an existing assistant.
What we do.
What you can use.
Gather inputs, access, and test accounts
We ask for a walkthrough of the system, the model behind it, and the tools and data it can reach, plus test accounts in at least two roles and two tenants. Seeing the hidden instructions and tool definitions lets us test in more depth. We confirm your provider terms allow this testing and agree a rate limit, a spend limit, and a pause contact.
Attack every entry point by hand
We work through the ways an AI application gets abused: planting instructions in input and in content the model reads, drawing out its hidden instructions, forcing unsafe output, giving it more power to act than it should have, and driving usage up to exhaust it. Then we test the interface and access controls as well. Each attempt is recorded so successes are counted, not claimed.
Hand over transcripts, fixes, and a retest
The report holds every finding with the exact transcript, how many of the attempts succeeded, screenshots and logs, a severity rating, and the category of risk in plain terms. Fixes land at the layer that closes them, because hardening a prompt does not repair a broken access control. We retest remediated findings and hand over the evidence.
Who it’s for.
When you need it.
- Teams shipping a customer-facing chat assistant or an internal copilot backed by a model
- SaaS providers whose buyers request evidence that the AI feature was security tested
- Product owners who added tools, file upload, or retrieval to an existing assistant
- Engineering groups exposing a language-model API to external or partner traffic
- A customer-facing assistant is scheduled to launch or leave beta within the next quarter
- A security questionnaire or customer review asks how you handle prompt injection
- The assistant recently gained new tools, connectors, or a broader API surface
- Model spending has grown enough that unbounded consumption would be costly
What the scope can include.
- 01
Overriding the model's instructions through the chat box, uploads, form fields, and the interface
- 02
Slipping instructions into documents, fetched pages, and other content the model reads later
- 03
Drawing out the model's hidden instructions, policy text, and tool definitions
- 04
Leaking one user's data to another through conversation memory, caches, and logs
- 05
Unsafe use of model output: content rendered in the page, or commands built from it
- 06
Sign-in, sessions, request-level authorization, rate limits, and how model credentials are handled
What it typically costs.
One rate: $150/hour.
Every engagement is priced by the hours it takes at one flat rate, with scoping, the work, and reporting, plus a retest of your fixes, included. Find the size closest to yours.
One chat assistant with no connected tools, tested as one or two user types
About 40–56 hoursOne assistant that accepts file uploads or uses a few connected tools, tested across several user roles and customer accounts
About 64–96 hoursSeveral AI features or a multi-customer platform with many connected tools, roles, and integrations
About 120–180 hours- How many AI features and ways in (chat, uploads, forms, outside content) are in scope
- Number of user roles and customer accounts to test as
- Whether the assistant can call tools or reach other systems
- Whether hidden instructions and source code are shared for deeper testing
Ranges are planning estimates at $150/hour, not a quote. Your price is confirmed in writing after a scoping call, before any work begins.
What you take forward.
- Findings with exact transcripts, attempt counts, screenshots, and logs for each successful attack
- Severity by impact and exploitability, with the risk category named in plain terms
- Fixes placed at the correct layer: the prompt, the application code, access control, or configuration
- Retest of remediated findings with before-and-after transcripts as evidence
Final coverage, deliverables, timing, and any retesting or implementation work are confirmed before the engagement begins.
Before we get started.
Will testing run up our model bill or get our provider account suspended?
We agree a spending limit and a rate limit before testing starts, and we ask you to confirm that your provider's terms permit adversarial testing of your own deployment. Hosted model services publish usage policies, and we stay inside them. The tests that probe runaway usage run in short, measured bursts with your pause contact aware, so we can show the exposure without producing the bill. If a staging environment exists, cost-related tests run there first.
You found prompt injection; is that the vendor's problem or ours?
Usually yours to fix, even when the model belongs to a vendor. Overriding a model's instructions is still an unsolved problem at the model layer, so the durable fixes live in your application: what the model is allowed to call, whose identity it acts with, which outputs get rendered or run, and what data it can see for a given user. Our report separates the findings that need a prompt or filter change from those that need an access or output-handling change, and names the owner of each.
- OWASP GenAI LLM Top 10 2026
- OWASP Top 10 for Agentic Applications for 2026
- GenAI Red Teaming Guide
- MITRE ATLAS
- NIST AI 600-1: Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile
- Model Context Protocol Specification: Authorization (2026-07-28)
- Document-Level Access Control - Azure AI Search
