Safety & Security · Data artifact

Violation Response Policy

Data artifactSafety & SecuritySafety, Security & GovernanceVariation point (abstract)arc:ViolationResponsePolicy

An abstract configuration stating what the system does when an output is judged to violate, or plausibly violate, a principle.

Responsibility. Determines the system's reaction to principle violations.

Also known as: Failure response policy, Enforcement strictness

configuresconfiguresis specialized byis specialized byis specialized byis specialized byis specialized byAction Policy Engine: configuresAction Policy EngineOutput Rail: configuresOutput RailBlocking Violation Response Policy: is specialized byBlocking Violation Respo…Human-Escalation Violation Response Policy: is specialized byHuman-Escalation Violati…Content-Modification Violation Response Policy: is specialized byContent-Modification Vio…Warning-Label Violation Response Policy: is specialized byWarning-Label Violation …Monitor-Only Violation Response Policy: is specialized byMonitor-Only Violation R…
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
Blocking Violation Response PolicyChoose for clearly prohibited categories (e.g., illegal activity) where erring on the side of safety outweighs refusing some acceptable requests.
Content-Modification Violation Response PolicyChoose when the violating portion can be removed or rewritten so the remaining response stays useful.
Human-Escalation Violation Response PolicyChoose for ambiguous cases with possible legitimate justification (fiction, security research) or high-stakes dilemmas, accepting slower responses for greater accuracy.
Monitor-Only Violation Response PolicyChoose for borderline outputs approaching policy boundaries where immediate blocking is not required; use blocking mode for inviolable legal, ethical or safety constraints.
Warning-Label Violation Response PolicyChoose when preserving user autonomy matters and the concern (e.g., possible misinformation, low confidence) warrants flagging rather than blocking.

Relationships

configures structural

Design guidance

Classification

Quality attributes
Safety (ISO/IEC 25010 | NIST AI RMF: safe)Interaction capability (ISO/IEC 25010)

Sources

  1. Ch9.5: T. Nguyen, "Constitutional AI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.5. ISBN: 9798244538229.
  2. Ch10.4: T. Nguyen, "Human-in-the-Loop," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.4. ISBN: 9798244538229.