Industry concept · Evidence and observability

AI agent monitoring

Last reviewed
2026-10-09

AI agent monitoring is the ongoing observation of agents in operation against expected behavior, using measures such as error rates, latency, cost and task outcomes, usually with alerts when they move out of range. It tells operators that something has changed, while diagnosing why usually falls to tracing and observability.

The discipline is older than agents. Google’s site reliability engineering book describes monitoring as collecting and displaying real-time quantitative data about a system, built around four golden signals: latency, traffic, errors and saturation. A monitoring system, it says, should answer two questions: what is broken, and why. The first is the symptom, and monitoring is good at it. The second usually needs more detail than a dashboard holds.

Agents inherit all of that and add problems of their own.

Why agents are harder to monitor

A web service that returns a 200 has, in most cases, done its job. An agent that returns a confident answer may not have. Anthropic’s guidance on agent evaluation makes the point with a booking agent: the reply says the flight is booked, but the outcome that matters is whether a reservation exists in the database. Request success and task success are different measurements, and only the first comes for free.

Three more properties make the job harder:

  • Non-determinism. The same prompt can produce different runs. Anthropic’s account of its multi-agent research system notes that agents are non-deterministic between runs, even with identical prompts, so a single good run proves little.
  • Compounding. Agents are stateful and take many steps, so an early mistake shapes everything after it. A run can look healthy at every step and still end somewhere it should not be.
  • Cost as behavior. Token usage, tool calls and retries are not just bills. A sudden rise in tool calls per task is often the first visible sign that an agent is looping or has lost the thread.

What gets monitored

In practice, agent monitoring watches three layers.

Operational health. Latency, error rates, token consumption and cost, the agent-era version of the golden signals. OpenTelemetry’s emerging conventions for generative AI give these vendor-neutral names: the duration of an agent invocation, the number of tool calls it makes, the input and output tokens it uses. The conventions are still marked as in development, which says something about how young the practice is.

Outcomes. Whether tasks succeed, measured by checking results rather than replies, often by evaluating a sample of production traffic. Cloud providers increasingly offer this as scheduled or sampled evaluation, partly to catch drift as models, prompts and data change.

Behavior against intent. Whether the agent is doing what it was asked to do. This layer is the least developed, and it is the one autonomy makes most important.

The signal that is easy to miss

An agent can be fast, error-free, cheap and successful by its own account while working on the wrong thing. Operational metrics cannot see that, because nothing failed. Outcome checks catch some of it, after the fact and only for the tasks someone thought to check.

Catching it as it happens requires a reference point: what the agent was supposed to be doing. When each action is evaluated against the agent’s declared objective and scope, departures become countable. The share of actions outside the declared scope, or the number of high-consequence operations per session, can be watched over time like any other metric. Those flags and violations are exactly the kind of signal monitoring is built to track: cheap to collect, meaningful in aggregate, and worth an alert when the rate changes.

A scope-departure rate is a governance signal, not a complete measure of task success. An agent can stay entirely within its declared scope and still fail: by doing the right kind of work badly, stopping short, or getting the answer wrong. Outcome checks remain part of monitoring for exactly that reason.

That framing also helps with attention. Monitoring that alerts on every unusual action trains people to ignore alerts. Monitoring that alerts when the rate of departures from declared intent changes sends a person to look only when there is something to judge.

An example

The following is illustrative.

A team runs a support agent that answers billing questions. After a prompt change, every operational chart stays flat: latency normal, error rate near zero, cost steady. Customer satisfaction scores do not move for a week. But the share of sessions in which the agent attempts an action outside its declared scope, such as editing account settings instead of reading invoices, rises from about one percent to nine. An alert on that rate fires on the first day. The traces show why: the new prompt encouraged the agent to “resolve the customer’s issue fully”, and it took that as permission.

No error was raised, and nothing was blocked. The monitoring worked because it was watching a measure of intent, not just a measure of health.

How Sentience Governor applies it

Sentience Governor contributes the governance side of that signal. It evaluates each attempted action against the session’s declared intent and scope, and records the advisory flags and policy violations it finds as governance evidence in the execution record, without blocking the agent. Those findings are structured facts that can be counted and watched over time alongside operational metrics. Governor records them; alerting on them is the job of whatever monitoring an organization already runs.

Where this is heading

Emerging approaches. Shared conventions, such as OpenTelemetry’s generative AI semantic conventions, are making agent telemetry portable between tools. Sampled online evaluation is becoming a standard part of production monitoring. And risk frameworks are pushing in the same direction: NIST’s AI Risk Management Framework asks that AI systems be monitored in production and that post-deployment monitoring plans be put in place.

Open questions. How to measure task success when outcomes are open-ended. How to set thresholds for behavior that is non-deterministic by design. And how monitoring should treat multi-agent systems, where the unit that succeeds or fails is a chain of agents rather than one.

Often confused with

Agent observability. Monitoring watches known measures and says when they move. Observability makes execution inspectable so that someone can find out why. The two share data, and most teams need both.

Agent evaluation. Evaluation tests an agent against defined tasks, usually before deployment. Monitoring watches it in production. Anthropic’s guidance puts the trade-off plainly: production monitoring catches issues that synthetic evaluations miss, but it is reactive, so problems reach users before anyone knows.

Where the concept stops

Monitoring tells an operator that a measure has moved. It does not explain why, it does not judge whether a given action was appropriate, and it can only watch what it was set up to measure. An agent doing the wrong thing quietly, within every threshold, is invisible to monitoring until a measure of intent is part of what it watches.

Often confused with

Agent evaluationAgent observability

Related entries

Bridges

See in practice

Sources

  1. Google, Site Reliability Engineering, Monitoring Distributed Systems
  2. OpenTelemetry, Semantic conventions for generative AI metrics (in development)
  3. OpenTelemetry, Semantic conventions for GenAI agent and framework spans (in development)
  4. NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0)
  5. Anthropic, How we built our multi-agent research system (2025)
  6. Anthropic, Demystifying evals for AI agents (2026)
  7. Microsoft Learn, Observability in generative AI (Microsoft Foundry)
Return to the glossary