AI Security · 3 min read

Your AI agent’s permissions are part of its attack surface.

A practical way to review what an agent can read, what it can change, and where untrusted content enters the workflow.

Follow the workflow beyond the chat window

An AI assistant that summarizes a document has a different exposure than an agent that can also email the summary, modify a record, or deploy code. To review the system, draw the workflow: user request, retrieved content, model processing, tool call, and resulting action. Identify the data and authority that cross each connection.

A useful question for every tool is: what could this action affect if the model misunderstood the task? OWASP describes excessive agency in terms of unnecessary functionality, permissions, and autonomy. Narrow tools and permissions to the job the application is meant to perform. [1]

Separate retrieved content from authorization

An external document can contain instructions that try to redirect the agent. OWASP’s prompt injection guidance distinguishes direct user-input attacks from instructions carried through external content, such as retrieved files or web pages. Content filtering and prompt design can help, but should not be treated as complete prevention. [2]

Imagine a fictional support agent reading a customer attachment that says to send the entire account history to a new address. Reading that sentence does not grant permission to send the records. Review where the application enforces this distinction: retrieval, model context, tool selection, and the service that actually executes the action.

Enforce the boundary where actions happen

For consequential actions, validate authorization outside the model’s free-form reasoning. Check the actor, destination, resource, and parameters at execution time. If an approval is required, it should apply to the exact action being executed. OWASP’s agent security guidance recommends independent scope, privilege, and approval validation for high-impact actions. [3]

Prefer a tool that reads a permitted record over one that runs an arbitrary database query when the task only needs that record. Keep a drafting capability separate from sending where the workflow allows it. These are concrete ways to reduce the authority available when a model produces an unexpected result. [1]

Test outcomes, then preserve the evidence

Use a test environment and synthetic data to exercise realistic misuse scenarios. Can one user retrieve another user’s content? Can a document change the destination of an export? Does a rejected action stay rejected after a retry? Record the attempted action and the enforced result, not only whether the model used reassuring language.

Keep enough protected telemetry to understand tool use and authorization outcomes. Repeat relevant checks after material changes to prompts, tools, retrieval, memory, or providers, as OWASP recommends. A useful security finding identifies the boundary that failed and the system control needed to restore it. [3]

References & further reading

  1. OWASP GenAI Security Project: Excessive Agency
  2. OWASP GenAI Security Project: Prompt Injection
  3. OWASP Cheat Sheet Series: AI Agent Security
START AT THE SOURCE

Bring this thinking to your environment.

A focused conversation. A clear scope. A practical path to stronger security.

Let’s talk security