Policy as code is the practice of writing policies in a machine-readable form that is version-controlled, tested and evaluated automatically, instead of keeping them only as documents. It makes a policy’s application repeatable and its changes reviewable.
A policy written in a document depends on someone reading it, interpreting it and remembering to apply it. A policy written as code is applied the same way every time, and every change to it leaves a record. That shift, from guidance people follow to rules a system evaluates, is what policy as code means.
Where it came from
The practice grew up in infrastructure and cloud operations, where the volume of changes outran what people could review by hand. NIST’s guidance on DevSecOps describes it as codifying policies and running them as part of the delivery pipeline, versioned, documented and access-controlled like application source code.
Several tools now define the category.
- Open Policy Agent decouples policy decisions from enforcement. Policies are written in a declarative language, Rego, and a decision is the answer to a query: given this input, what does the policy say?
- HashiCorp Sentinel describes policy as code as writing code to manage and automate policies, kept as plain text files under version control.
- Cedar, an open-source language from AWS, writes authorization policies that are separate from application code, and was designed to be analyzable: its authors used a proof assistant to establish properties of the language, and found and fixed bugs in the process.
- Kubernetes and Kyverno express admission policies as declarative resources inside the cluster.
What it gives you
Repeatability. The same input produces the same decision, whoever is on call.
Reviewability. A policy change is a diff. It can be discussed, approved and rolled back like any other change.
Testability. Policies can carry their own tests. OPA, for example, runs test rules written in the same language as the policy, so a change that breaks an expectation fails before it ships.
Staged enforcement. Most policy engines separate the rule from what happens when it fails. Sentinel has three levels: advisory, where a failure is only reported; soft mandatory, where it can be overridden; and hard mandatory, where it cannot. Kubernetes admission policies can deny a request or only record the failure in the audit log. That separation lets a new rule run in a reporting mode, be measured against real traffic, and only then be made to block anything. Enforcement can be proportional to confidence and consequence, instead of all or nothing.
An example
The following illustrative rule, in Rego, allows an agent to send email only to the organization’s own domain:
package agent.tools
default allow := false
allow if {
input.tool == "send_email"
endswith(input.recipient, "@example.com")
}
The rule is short, testable and versioned. It also shows the hard part. The policy is only as good as the input it receives: here, a tool name and a recipient. It cannot see why the agent is sending the email, whether that serves the task it was given, or whether the message contains data it should not.
What changes when the subject is an agent
Policy as code was built for subjects whose requests are well described: a deployment, a network call, an API request with a known schema. Agents strain that model in three ways.
Meaning, not just names. An agent’s actions arrive as tool calls, shell commands and free text. A rule keyed to a tool name cannot tell a harmless call from a destructive one made through the same tool. The facts a policy evaluates need to describe what an action does and what it affects, not only what it is called.
Context, not just the request. Whether an action is appropriate often depends on what the agent was asked to do. A policy that evaluates each request in isolation cannot tell the same database query in a reporting task from the same query in a cleanup task. Agent policy needs the declared objective and scope as part of its input.
Which policy, and which version. When many agents run under different rules, the first question is which policy applied to a given run. Policy resolution answers it, and recording the answer, including the exact version of the policy, is what lets a finding be traced to the rules that produced it.
Early work is appearing. Amazon Bedrock AgentCore, for example, evaluates agent tool requests against Cedar policies before allowing tool access. The broader pattern is still forming: policy defined independently of the agent, evaluated against what the agent actually does, with the result kept as evidence.
How Sentience Governor applies it
In Sentience Governor, an operator’s expectations for an agent live in a governance profile: a short configuration file that can be versioned, reviewed and changed like other operational configuration. Profiles are resolved per agent, and the record identifies which policy governed each session, so every finding can be traced to the exact policy that produced it. Governor evaluates profiles to record findings; it does not block or deny actions.
Where this is heading
Emerging approaches. Authorization languages such as Cedar applied to agent tool access. Policy inputs enriched with the meaning of actions and the declared intent of the session. And the staged rollout infrastructure already uses, from audit to enforce, applied to agent policy.
Open questions. Who writes agent policy, and how it is tested when the agent’s behavior is not deterministic. How a policy can say “consistent with the task” in a machine-readable way. And how policy written for one agent framework can apply to another that names the same actions differently.
Often confused with
AI guardrails. Guardrails are controls in the path of a model or agent: filters, validators, checks. Policy as code is a way of writing and managing rules. A guardrail can be implemented as policy as code, and many are written some other way.
Configuration management. Both are versioned and automated. Configuration says how a system is set up; policy says what is allowed or expected of it.
Where the concept stops
Policy as code makes rules explicit, consistent and reviewable. It does not make them right: a well-tested policy can encode the wrong expectation, and it decides only on the inputs it is given. For agents, the hardest part is rarely the policy language. It is giving the policy facts that describe what an action means and what the agent was asked to do.