Safety & Security · Data artifact
Guardrail Policy
Data artifactSafety & SecuritySafety, Security & Governancearc:GuardrailPolicy
A declarative, version-controlled configuration of rail definitions — canonical user intents, predefined bot responses, dialogue flows, enabled rails, thresholds and fact-checking method — that programs guardrail behaviour independently of the LLM.
Responsibility. Specifies which safety checks apply and under what conditions.
Also known as: Rails configuration, Colang flows, Guardrail configuration, Brand safety policy, config.yml, rails.co, Fairness rail configuration, Colang flow definitions, Executable constitutional principles
Relationships
configures structural
is configured by structural
is audited by assurance
is evaluated by assurance
Design guidance
- SHOULD enable only the rails the threat model requires; enabling rails is an architecture decision.
- SHOULD enable all six rails with strict policies in high-security environments (healthcare, finance), accepting 50-150ms latency as a compliance cost.
- SHOULD keep guardrail logic in version-controlled text files to enable code review, audit trails and rollback.
- SHOULD keep thresholds (e.g., approval limits) in auditable configuration rather than scattered through tool implementations.
- SHOULD be validated against adversarial prompts in development before production deployment.
- SHOULD keep policy declarative and visible so compliance officers can review active protections without tracing code.
- SHOULD be updated rapidly as new attack techniques and production bypass attempts are discovered.
Classification
- Patterns
- Programmable, composable guardrailsPolicy as codeCanonical forms (define user)Predefined bot messages (define bot)Flows with execute/await statementsColang 1.0 and 2.0 syntaxCanonical user intent + bot refusal message + flow linking them
- Technologies
- ColangNVIDIA NeMo Guardrails
- Quality attributes
- Maintainability (ISO/IEC 25010)Transparency and accountability (NIST AI RMF: accountable and transparent)
Sources
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
- Ch7.1B: T. Nguyen, "Nvidia NIM and Colang," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1B. ISBN: 9798244538229.
- Ch8.2B: T. Nguyen, "NeMo Guardrails Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.2B. ISBN: 9798244538229.
- Ch9.1: T. Nguyen, "Output Filtering and Content Moderation," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.1. ISBN: 9798244538229.
- Ch9.4: T. Nguyen, "Fairness and Bias Mitigation," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.4. ISBN: 9798244538229.
- Ch9.5: T. Nguyen, "Constitutional AI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.5. ISBN: 9798244538229.
- Ch10.4: T. Nguyen, "Human-in-the-Loop," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.4. ISBN: 9798244538229.
- Ch10.5: T. Nguyen, "Human-over-the-Loop," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.5. ISBN: 9798244538229.
- Ref7.03: NVIDIA, "Overview," NVIDIA NeMo Guardrails Library Developer Guide. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nemo/guardrails/about-nemo-guardrails-library/overview
- Ref9.04: "Safety Guardrails Implementation for Agent Systems," unpublished reference note (04-Safety-Guardrails-Implementation.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note