Your AI Agent Needs a Brake Pedal: 3 Safety Controls
Explore three engineering patterns—Approval Gates, Scoped Permissions, and Dry Runs—to mitigate Excessive Agency risks in AI agents.

Stock photo for illustration only, not from the actual event
- AI agents now perform critical tasks like opening PRs and sending emails.
- Three core engineering patterns limit blast radius: Gates, Permissions, Dry Runs.
- Scoped permissions offer the strongest defense against prompt injection attacks.
- These mitigations directly address 'Excessive Agency' in the OWASP Top 10 for LLMs.
AI agents are no longer confined to answering simple user prompts. Today, they actively open pull requests, execute database queries, and dispatch emails on behalf of users. As these autonomous workflows expand, developer forums and regulators are increasingly questioning who bears responsibility when an agent executes an unauthorized or destructive action.
Regardless of how legal and regulatory frameworks evolve, the engineering consensus remains clear: grant real capabilities to the agent, but strictly bound its blast radius. Three core design patterns accomplish the majority of this safety work.

Stock photo for illustration only, not from the actual event
The first pattern is Approval Gates. While the agent continues to reason and plan autonomously, any consequential tool execution pauses at the final step, waiting for human intervention and explicit approval before proceeding.
The second pattern is Scoped Permissions. While a gate relies on a human catching a flawed proposal, scoped permissions ensure that dangerous actions are fundamentally impossible to execute in the first place. This represents standard least privilege enforcement and serves as the strongest defense against prompt injection, because malicious instructions cannot invoke capabilities the agent lacks.
Implementing scoped permissions is critical as Large Language Models (LLMs) gain broader tool-use capabilities. Constraining agent agency directly mitigates the 'Excessive Agency' vulnerability vector, preventing attackers from exploiting prompt injection to manipulate connected enterprise systems.
The third pattern is a Dry Run, which executes the planning logic without triggering external side effects. For instance, a dry run preview message might explicitly state:
"This would delete 3 rows: id IN (812, 813, 977) ."
Dev.to
An approval gate answers the question "Should this action happen?" whereas a dry run answers "What exact changes will occur?" These two mechanisms work best when combined. Approving a concrete, detailed dry-run output is vastly more meaningful and secure than authorizing a vague request to "delete some records."
Together, these practices serve as direct mitigations for "Excessive Agency" as categorized in the OWASP Top 10 for LLM applications—representing the fastest-growing risk category as agent capabilities scale upward.
Source: Dev.to
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment