Glossary · term

AI Guardrails

AI guardrails are explicit controls that constrain, inspect, or redirect an AI application's behavior at runtime. Depending on the system, they can evaluate user input, retrieved context, dialogue state, model output, or proposed actions; block or transform content; redact sensitive information; require a refusal or escalation; and record policy results. Guardrails complement a model's built-in training and alignment. They are application controls, not a promise that every unsafe or incorrect behavior will be prevented.

Safety2023-04-25Wave 2 · 2024Maturity: 3/5

Origin and context

NVIDIA's April 2023 NeMo Guardrails release used the term for programmable topical, safety, and security boundaries around language-model applications. AWS independently made Amazon Bedrock Guardrails generally available in April 2024, with denied topics, harmful-content thresholds, word filters, and sensitive-information controls that could be applied across models. NIST's Generative AI Profile places such controls in a wider risk-management cycle of defining tolerance, evaluating performance, documenting decisions, and monitoring effectiveness.

Sources: s1, s2, s3

Why it matters

A base model's generic policy cannot fully represent every application's users, data permissions, regulated topics, or consequences. Guardrails provide a place to express context-specific rules and to apply them consistently across models. They also make policy outcomes observable: teams can measure blocks, false alarms, escalations, and changes after a model or prompt update. This supports defense in depth, but deterministic authorization and business rules should remain outside the model and its natural-language instructions.

Sources: s1, s2, s3

Example

A benefits assistant may screen an incoming message for disallowed abuse and sensitive identifiers, retrieve only records the authenticated user may access, check whether the draft answer is supported by approved policy text, and redact protected fields before display. A proposed account change is validated by deterministic permission rules and may require human confirmation. Each control is tested separately and as part of the full workflow, with thresholds selected for the use case rather than copied unchanged from a vendor default.

Sources: s1, s2, s3

How it differs

Prompt injection

Prompt injection is an attack class in which untrusted content changes an application's intended instructions or behavior. Guardrails are a broader set of preventive, detective, and response controls. An input classifier or tool-call policy can reduce injection impact, but calling it a guardrail does not make prompt injection impossible.

Groundedness

Groundedness asks whether claims are supported by a specified context. A groundedness evaluator can be used as one output or retrieval rail, but guardrails also cover content policy, privacy, dialogue flow, permissions, and actions. A response can be grounded in a source that is itself wrong or unauthorized.

Maturity and evidence

Maturity is rated 3. NVIDIA and AWS provide independent, configurable implementations, while NIST supplies organization-level guidance for testing and monitoring controls. The practice is established in production platforms, but terminology, coverage, interfaces, evaluation sets, and acceptable error rates remain use-case dependent rather than standardized.

Sources: s1, s2, s3

Limits and open questions

Guardrails can miss harmful cases and block legitimate ones; their performance changes with language, modality, context, model versions, and adversarial behavior. A model-based judge may share blind spots with the model it checks. Static word filters cannot understand every context, and natural-language rules must not carry secrets or enforce access control. Teams should threat-model the full application, use least privilege and deterministic authorization, evaluate each rail against representative and adversarial cases, monitor drift, retain an appeal or escalation path, and document residual risk. Passing a guardrail is evidence from one control, not proof of safety or compliance.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-09-04

In the Skills Atlas

This term is also covered in the Skills Atlas as ai guardrails skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as nemo guardrails skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as owasp top 10 for llm applications skill.