The challenge behind the engagement.
This assessment is for an agent that reads, decides, and acts with authority you have delegated, whatever software runs it. We inventory every tool the agent can call, along with the identity behind it, the access it carries, and whether its actions can be undone. Then we try to redirect the agent's goal through the documents, tickets, web pages, and calendar entries it reads, chain low-risk tools into high-risk ones, reuse its stored credentials, and poison its memory so planted instructions return in a later run. We test approval steps for fatigue and for stale approvals. You receive a permission and blast-radius matrix, the attack chains, and the approval points and access cuts to make. Anything the agent can change is tested in a sandbox or under a written allowlist.
For engineering leads and security teams deploying an agent that can send email, change records, run code, or call internal APIs. Run it before granting the agent write access, before moving from pilot to production, or after adding a second agent or a new tool.
What we do.
What you can use.
Walk the design and inventory tools
We start with a design walkthrough with the builders and an inventory of every tool the agent can call, listing the identity behind each one, the access it holds, and whether its actions can be reversed. We agree a sandbox for anything the agent can change, or a written production allowlist, and we get access to its activity traces and memory store.
Hijack, chain, and poison
We plant instructions in the content the agent reads to steer it off its goal, chain low-risk tools into high-risk ones, reuse the credentials it has cached, and store poisoned memory that fires on a later run. Approval steps are tested for fatigue and for stale approvals. Each chain is captured as a trace tied back to the instruction that caused it.
Deliver the blast-radius matrix
You receive a matrix of agent, tool, identity, access, reversibility, and approval step, plus the attack chains with transcripts and traces. Findings are tied to recognized categories of agent risk in plain terms, with specific access cuts and approval points to add. After you change the design or permissions, we rerun the chains and report which are closed.
Who it’s for.
When you need it.
- Teams deploying an agent that can send email, change records, or call internal APIs
- Organizations moving an agent from a supervised pilot toward more autonomous operation
- Security teams reviewing an agent that holds its own credentials and write access
- Builders running multi-agent workflows where one agent's output drives another's actions
- You are about to grant the agent write access or remove human approval steps
- A new tool, connector, or second agent was just added to the workflow
- The agent is moving from an internal pilot to production or external users
- The agent took an action no one intended and the cause is unclear
What the scope can include.
- 01
Redirecting the agent's goal through documents, tickets, web pages, or messages from other agents
- 02
A per-tool inventory of identity, access, reversibility, and paths from reading to writing
- 03
Poisoned memory that carries planted instructions into later runs or other users' sessions
- 04
Approval steps: which actions need a human, approval fatigue, and reuse of stale approvals
- 05
Escape from the sandbox and outbound network reach from code the agent writes and runs
- 06
Tools, add-ons, and connectors the agent loads at runtime without review
What it typically costs.
One rate: $150/hour.
Every engagement is priced by the hours it takes at one flat rate, with scoping, the work, and reporting, plus a retest of your fixes, included. Find the size closest to yours.
One agent with a handful of tools, tested in a safe copy of your systems
About 56–80 hoursOne agent with many tools, write access, and memory, or two agents that hand work to each other
About 96–140 hoursA team of cooperating agents with broad tool access, shared memory, and some live-system actions
About 160–240 hours- Number of agents and the tools each one can use
- How much the agent can change on its own, and whether those actions can be undone
- Whether the agent remembers things between sessions or users
- Whether a safe test copy exists or live-system limits must be written
Ranges are planning estimates at $150/hour, not a quote. Your price is confirmed in writing after a scoping call, before any work begins.
What you take forward.
- Permission and blast-radius matrix covering agent, tool, identity, access, reversibility, and approval step
- Attack chains with transcripts and execution traces, tied to recognized categories of agent risk
- Recommended approval points, access cuts, and logging so every tool call names the identity behind it
- Retest of each attack chain after design or permission changes, with traces as evidence
Final coverage, deliverables, timing, and any retesting or implementation work are confirmed before the engagement begins.
Before we get started.
Can you test without the agent sending real emails or changing real records?
Yes, and we prefer it. During scoping we agree a sandbox for every system the agent writes to, such as a test mailbox, a non-production CRM tenant, or a scratch repository, so chains can run to completion. Where a sandbox does not exist, we write a production allowlist naming the exact records, recipients, and actions we may touch, and stop the chain at the last read-only step everywhere else. The report states which chains were fully executed and which were proven up to the write.
Our agent is read-only, so what is the risk?
Read-only still means the agent can be steered to read the wrong thing and repeat it. A hijacked read-only agent can pull records the asking user may not see, summarize them into a reply, or exfiltrate them through a link or a tool that fetches a URL. We also check whether read-only is true: cached tokens with broader scopes, tools that accept side-effecting parameters, and memory that persists across users often turn a read-only agent into a write path. The matrix shows the actual reach, not the intended one.
- OWASP GenAI LLM Top 10 2026
- OWASP Top 10 for Agentic Applications for 2026
- GenAI Red Teaming Guide
- MITRE ATLAS
- NIST AI 600-1: Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile
- Model Context Protocol Specification: Authorization (2026-07-28)
- Document-Level Access Control - Azure AI Search
