AI agent security is the protection of agents, and of the systems they can reach, against misuse and attack, including prompt injection, credential and tool abuse, data exfiltration and excessive privileges. It treats the agent both as something to defend and as a potential path to everything it has access to.
Traditional application security asks whether an attacker can get in. Agent security adds a second question: once an attacker can influence the agent, what can the agent do for them? An agent holds credentials, calls tools and acts on someone’s behalf. Anything it can reach, a compromised agent can reach too.
The agent as an attack path
NIST’s taxonomy of adversarial machine learning puts it directly: because agents can take actions using tools, attacks on the underlying model create additional risks, and can let attackers hijack agents to execute arbitrary code or exfiltrate data from the environment they operate in. NIST’s Center for AI Standards and Innovation describes agent hijacking as a form of indirect prompt injection in which malicious instructions are planted in data the agent is likely to read, causing it to take unintended, harmful actions.
The pattern is the old confused deputy problem in new form. The agent has legitimate authority; the attacker borrows it. The Model Context Protocol’s security guidance names confused deputy vulnerabilities explicitly for servers that broker access to third-party services.
The main risks
OWASP’s Top 10 for Agentic Applications, published in December 2025, frames the field. Its list begins with agent goal hijack, tool misuse, and identity and privilege abuse. A few risks recur across every framework.
Excessive agency. OWASP’s LLM guidance traces many incidents to one or more of excessive functionality, excessive permissions and excessive autonomy, and recommends limiting what an agent’s tools can do to the minimum necessary.
Credential and tool abuse. Agents often hold broad, long-lived tokens because they are convenient. A stolen broad token, in the words of the MCP guidance, expands the blast radius to unrelated tools and resources. Tools themselves can be weaponized: researchers have shown malicious instructions hidden in tool descriptions that users never see but models read, which they called tool poisoning.
Data exfiltration. Simon Willison’s “lethal trifecta” names the combination that makes theft easy: access to private data, exposure to untrusted content, and the ability to communicate externally. If one agent has all three, an attacker who controls some of the content can instruct it to send the data out.
Supply chain and identity. Agents load tools, plugins and other agents at run time, and each is a dependency. And many organizations cannot yet say which agent did what, under whose authority.
Why least privilege is harder for agents
Least privilege is the classic answer, and it still applies. It is harder to apply to agents because the actions an agent needs are not fully known in advance. NIST’s National Cybersecurity Center of Excellence, in a 2026 draft concept paper on agent identity and authorization, asks exactly this: how to establish least privilege for an agent when its required actions might not be fully predictable when deployed, and whether its identity should be fixed or tied to the task.
That points toward two complementary layers. Permissions set the outer bound of what an agent can do: credentials, roles, network access. A declared scope describes what this agent, on this task, is expected to touch, which is usually far narrower. An agent with write access to a whole cloud account, asked to clean up one staging service, has a permission problem and a scope question. Security narrows the first; comparing actions with the second shows when a hijacked or confused agent strays.
Consequence and evidence
Not every action deserves the same scrutiny. Deleting infrastructure, sending data outside the organization, changing who has access: these are high-consequence operations, and they are where controls such as approval, sandboxing and stricter authorization pay for themselves. Concentrating defenses there, rather than spreading them evenly, is what keeps them usable.
And when something goes wrong, the first questions are what the agent did, in what order, under whose authority, and how far it went beyond its task. Answering them requires a record of the agent’s actions captured as they happened, linked to an identity and to the task it was given.
An example
The following is illustrative.
A research agent can read a company’s shared drive, browse the web and send email, which is the full trifecta. It summarizes a web page that contains hidden instructions to collect documents marked “board” and email them to an outside address. Its credentials allow all of it. A content filter does not recognize the instruction as harmful. What does stand out is that a task described as summarizing a public web page involved reading internal board documents and sending external email, two actions far outside its declared scope, one of them a high-consequence operation.
Removing one leg of the trifecta, such as external email for research tasks, would have closed the specific exfiltration path in this example, though not every possible attack. Comparing actions with the task would have shown the departure either way.
How Sentience Governor relates
Sentience Governor is not a security control. It does not detect attacks, manage credentials or block actions. It evaluates each attempted action against the session’s declared intent and scope and the operator’s designations of high-consequence operations, and records the findings as evidence while the action proceeds. In a security context, that record is useful for spotting departures from the task, whatever caused them, and for reconstructing what an agent did after an incident.
Where this is heading
Emerging approaches. Dedicated agent identities rather than shared service accounts, with authorization tied to the task. Agent-specific threat frameworks, such as OWASP’s agentic list, becoming shared vocabulary between security teams and builders. And architectural defenses that keep untrusted content from choosing an agent’s actions at all.
Open questions. How to grant least privilege when the task decides what is needed. How trust should pass between agents that delegate to one another. And how to attribute an action to an agent, a user and a task in a way an auditor can verify.
Often confused with
AI agent governance. Security asks whether an agent, or someone controlling it, can misuse its access. Governance asks whether the agent’s actions are consistent with its task and the expectations that apply. They overlap heavily, and an incident usually involves both, but a secure agent can still do the wrong thing, and a well-governed one can still be attacked.
Where the concept stops
Agent security protects agents and the systems behind them from misuse. It cannot make a model immune to manipulation, and controls that work for one deployment may not transfer to another with different tools and data. It reduces what an attacker can do through an agent and shortens the time to notice, but it does not establish that an agent’s legitimate actions were the right ones.