A flag identifies a condition or behavior that warrants attention during an AI agent’s execution, while a policy violation indicates that an action meets the conditions of a rule defining prohibited or noncompliant behavior. The distinction separates observations about agent behavior from judgments against applicable policy, without necessarily determining what response should follow.
Any system that evaluates agent actions produces findings, and a finding answers one of two different questions. The first is descriptive: what about this action is worth noticing? It left the declared scope, it deletes something, it is the first write after a long run of reads. The second is normative: which rule did it break? The two have different audiences and age differently. A description stays true as long as the record does. A judgment depends on which rules were in force, and rules change.
Why one alert stream is not enough for agents
A single agent session can produce hundreds of actions. Collapse everything an evaluation notices into one stream of “alerts” and the stream becomes too long to read, so people stop reading it. The record also loses the ability to answer basic questions later: would this run look different under a revised policy? Was this pattern ever the subject of a rule at all?
Human attention is the scarcest resource in agent governance. It should go to decisions that need judgment, not to sorting a flood of undifferentiated warnings. Keeping observation separate from judgment is what makes that possible. Observations can be broad and cheap, because nobody has to act on each one. Judgments can be narrow and specific, because each names a rule someone chose to care about. And the decision about what deserves a person’s time can be made deliberately, on top of both.
Observation, judgment, response
Evaluation systems in many fields separate three layers, even when they name them differently.
| Layer | The question | Typical form |
|---|---|---|
| Observation | What happened, and what about it is notable? | A finding, a signal, a flag |
| Judgment | Does it breach a rule, and which one? | A violation, a failed check |
| Response | What happens because of it? | Record, warn, ask a person, deny, roll back |
The pattern is well established. SARIF, the OASIS standard for static analysis output, separates a rule, the criterion an analysis checks, from a result, a condition found in an artifact, and gives each result its own level. Kubernetes lets the same admission policy deny a request, return a warning or add it to the audit log, depending on how it is bound. Azure Policy can audit non-compliance without stopping a request, or deny it outright. In each case, what is noticed, what is judged and what is done about it are separate decisions, often made by different people at different times.
Flags and violations are the first two layers. They are deliberately silent about the third.
What each one says
An advisory flag is a description, not an accusation. It records that a defined condition was observed on an action: the target was outside the declared scope, a write was attempted before any objective was declared, the action matched an operation the operator marked as high-consequence. A flag can mark a legitimate next step as easily as a problem. Several can attach to one action. And an action with no flags is not thereby approved; it is an action on which none of these particular conditions was observed.
A policy violation is a claim that a named rule matched. It is more specific than a flag and also more contingent. Change the rules and the same action would be judged differently, while what was noticed about it would not change.
When the two diverge
Most violations travel with a flag describing the underlying condition. The cases where they part are the instructive ones.
Picture an agent asked to remove an unused staging load balancer, with a declared scope of cloud infrastructure. It deletes the load balancer: the most consequential action of the session, and exactly what it was asked to do. That action carries a high-consequence flag and no violation, because it was inside the agent’s declared boundaries and no rule forbids it. Later, in the same session, it tidies up a local build directory. That action is routine and harmless, and it breaks a rule, because it reached outside the declared scope.
Read together, the two show why a single severity score misleads. The most consequential action broke no rule, and the action that did break one was mundane. Consequence and rule-breach are different axes, and a record that merges them cannot show either clearly.
Neither is a decision to stop
A finding is not a response. A violation written to the record is not a denial, a warning returned to the agent or a request for approval. That separation is a choice about where the response layer lives, not a claim that responses do not matter.
It also holds for systems that do enforce. “Denied by rule X” and “noticed, and allowed” are different facts, and keeping findings separate from responses lets an organization change how it responds without losing the history of what was observed.
How Sentience Governor applies it
Sentience Governor evaluates each attempted action against the session’s declared intent and scope, and against the operator’s governance profile. It records both kinds of finding side by side on each event in the Sentience Agent Execution Record: advisory flags for conditions it noticed, and policy violations for defined rules that matched. Neither interrupts the agent. Every action proceeds, and the record says so.
Where this is heading
Emerging approaches. Infrastructure policy already separates a check from its enforcement action. Gatekeeper, the Open Policy Agent project for Kubernetes, and Google Cloud’s organization policies can both run a rule in dry-run mode, reporting violations without denying anything, so a rule can be measured before it is made to deny. The same staged path suits agent policy: record findings first, learn what the rules would have caught and how often they would have been wrong, then decide which findings deserve a stronger response or a person’s attention. That only works if the record keeps observations and judgments apart.
Open questions. How findings should be graded beyond a binary split, for example by confidence or by consequence. How a finding produced under one version of a policy should be re-read under a later one. And how findings from many agents, governed by different profiles, can be compared at all.
Often confused with
An alert. An alert is a notification sent to someone. A flag is a fact written to the record whether or not anyone is notified.
A guardrail result. An AI guardrail typically decides, in the path of a call, whether something passes. A flag or a violation decides nothing about the call; it records a finding beside it.
Severity. A flag is not a low-severity violation. A high-consequence flag with no violation can matter more than a violation of a reporting rule.
Where the concept stops
A flag establishes that a defined condition was observed on an attempted action. A violation establishes that a named rule matched under the rules in force. Neither establishes that the action was harmful, that it completed, or that it went against what the person who assigned the task wanted. An action with no findings is not thereby appropriate: an action inside the declared scope that does nothing for the task produces no finding. What happens because of a finding is a separate decision, made outside the record, by a person or by a system built to make it.