Industry concept · Security

AI guardrails

Last reviewed
2026-10-09

AI guardrails are controls placed around an AI system to constrain or check its inputs, outputs or actions, such as content filters, schema validation, allow-lists and rule-based checks. They apply limits that someone defined in advance, and sit at different points depending on the product: before the model, after it, or around tool use.

The word is borrowed from roads, and the metaphor is accurate in one respect: a guardrail does not steer. It marks an edge and stops, or at least slows, whatever goes over it. In AI systems, guardrails are the checks that sit in the path of a request and decide whether something may pass.

Where guardrails sit

Products differ in vocabulary but agree on the positions.

  • Before the model. Input checks screen what goes in: harmful requests, attempts to override instructions, sensitive data that should not reach a model. NVIDIA’s NeMo Guardrails calls these input rails; Microsoft’s Prompt Shields analyzes user prompts and documents before content is generated.
  • After the model. Output checks screen what comes out: unsafe content, leaked secrets, answers in the wrong format. Amazon Bedrock Guardrails evaluates both prompts and completions against filters an administrator has defined.
  • Around tool use. For agents, the most consequential position. OpenAI’s Agents SDK lets tool guardrails validate or block a tool call before and after it runs; NeMo calls these execution rails. Prompt Shields also scans tool responses, because what a tool returns can carry an attack back into the agent.

Structured output is a quieter form of the same idea. Requiring a response to match a schema, and validating it with deterministic code, turns “the model should answer in this shape” into something that can be checked. OWASP recommends exactly this as a mitigation for prompt injection.

What guardrails are good at

Guardrails are fast, predictable in where they apply, and easy to explain: this check, at this point, with this outcome. They are the right tool for limits that can be stated in advance. Never return a credit card number. Never call the payments tool with an amount over a threshold. Always return valid JSON. Refuse requests in categories the product does not serve.

They also produce a clean signal. When a guardrail trips, something specific happened at a specific point, and the record of it is useful in its own right.

Where they stop

Guardrails have three limits that come from what they are.

They check what was defined in advance. A content filter recognizes the categories it was built for. A rule recognizes the conditions written into it. Bedrock’s documentation describes filtering based on predefined harmful content categories, and detection of sensitive information as probabilistic. Anything outside the definition passes, and detection near the edge is uncertain in both directions: legitimate requests get blocked, and some harmful ones do not.

They can be bypassed. Research on adaptive attacks is sobering. A 2025 study by Nasr, Carlini, Tramèr and colleagues bypassed twelve recently published defenses against jailbreaks and prompt injection, with attack success above 90 percent for most, although most had originally reported near-zero success. NIST’s taxonomy of adversarial machine learning concludes that current mitigations do not offer full protection.

Passing is not the same as serving the task. Some guardrail systems keep state across a conversation or use context in their decisions. But a guardrail is still built to check conditions defined in advance, and passing those checks does not establish that an action serves the agent’s overall declared objective.

Guardrails and governance

Guardrails answer a narrow question very well: does this input, output or action cross a line someone drew in advance? Governance asks a broader one: is what the agent is doing consistent with what it was asked to do, under the expectations that apply to this context?

The two work best together. An agent asked to summarize a support ticket might pass every guardrail, no harmful content, no secrets, valid output, while updating a customer record it had no reason to touch. No line was crossed; the action was simply outside the task. Evaluating actions against the agent’s declared scope catches that kind of departure, and recording it as a finding rather than an error lets someone judge it. Distinguishing between a flag and a violation matters here: a guardrail typically produces a decision in the path, while an evaluation against declared intent produces evidence beside it.

Where the expectations live matters too. Guardrail settings are usually configured per product or per application. A governance profile keeps an operator’s expectations for a class of agents in one reviewable place, independent of any single tool, so that the same intent can apply wherever the agent runs.

An example

The following is illustrative.

A coding agent has three guardrails: a secret scanner on its output, a deny rule on force-pushing to the main branch, and a check that its commands do not touch files outside the repository. Asked to update a dependency, it also rewrites a CI configuration file, inside the repository, with no secrets involved, without force-pushing. Every guardrail passes. The change disables a security check in the pipeline.

Nothing in the guardrails was wrong. They were written for specific edges, and the agent did not cross any of them. Noticing that a dependency update reached into CI configuration requires comparing the action with the task.

How Sentience Governor relates

Sentience Governor is not a guardrail. It does not sit in the path of a call or decide whether it may proceed. It evaluates each attempted action against the session’s declared intent and scope and the operator’s governance profile, and records the result as evidence while the action proceeds. In practice it runs alongside whatever guardrails a team already has, recording what they are not designed to judge.

Where this is heading

Emerging approaches. Guardrails are moving from text to actions, as tool-call checks become a standard feature of agent frameworks. Checks increasingly use models to judge models, trading determinism for coverage. And NIST’s generative AI profile recommends regularly reviewing safety and security guardrails, particularly when a system is used in new circumstances, which treats guardrails as something to govern rather than something that governs.

Open questions. How to evaluate guardrails against adaptive attackers rather than fixed test sets. How to balance false positives, which erode trust in the product, against false negatives, which erode trust in the guardrail. And how to express limits that depend on context, which is where most of an agent’s risk lives.

Often confused with

AI agent governance. Governance is the broader discipline of defining expectations for agents, evaluating behavior against them and keeping evidence. Guardrails are one kind of control within it.

Policy as code. Policy as code is a way of writing and managing rules. Many guardrails are implemented as code; many are not, such as model-based content classifiers.

Where the concept stops

A guardrail establishes that a defined check ran at a defined point and what it decided. It does not establish that what passed was appropriate, only that it did not trip that check. Passing it does not establish that an action served the agent’s declared objective, and it is as good as the conditions someone thought to define.

Often confused with

Related entries

Bridges

See in practice

Sources

  1. OpenAI Agents SDK documentation, Guardrails
  2. NVIDIA NeMo Guardrails documentation, Guardrail Types
  3. AWS, Detect and filter harmful content by using Amazon Bedrock Guardrails
  4. Microsoft Learn, Prompt Shields in Azure AI Content Safety
  5. OWASP GenAI Security Project, LLM01:2025 Prompt Injection
  6. Nasr et al., The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt Injections
  7. NIST AI 100-2 E2025, Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations
  8. NIST AI 600-1, Generative Artificial Intelligence Profile
Return to the glossary