Safety & Security · Data artifact
Violation Response Policy
Data artifactSafety & SecuritySafety, Security & GovernanceVariation point (abstract)arc:ViolationResponsePolicy
An abstract configuration stating what the system does when an output is judged to violate, or plausibly violate, a principle.
Responsibility. Determines the system's reaction to principle violations.
Also known as: Failure response policy, Enforcement strictness
Variants
| Variant | When to choose |
|---|---|
| Blocking Violation Response Policy | Choose for clearly prohibited categories (e.g., illegal activity) where erring on the side of safety outweighs refusing some acceptable requests. |
| Content-Modification Violation Response Policy | Choose when the violating portion can be removed or rewritten so the remaining response stays useful. |
| Human-Escalation Violation Response Policy | Choose for ambiguous cases with possible legitimate justification (fiction, security research) or high-stakes dilemmas, accepting slower responses for greater accuracy. |
| Monitor-Only Violation Response Policy | Choose for borderline outputs approaching policy boundaries where immediate blocking is not required; use blocking mode for inviolable legal, ethical or safety constraints. |
| Warning-Label Violation Response Policy | Choose when preserving user autonomy matters and the concern (e.g., possible misinformation, low confidence) warrants flagging rather than blocking. |
Relationships
configures structural
Design guidance
- SHOULD be chosen explicitly as a judgment balancing user experience, autonomy and safety assurance.
Classification
- Quality attributes
- Safety (ISO/IEC 25010 | NIST AI RMF: safe)Interaction capability (ISO/IEC 25010)
Sources
- Ch9.5: T. Nguyen, "Constitutional AI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.5. ISBN: 9798244538229.
- Ch10.4: T. Nguyen, "Human-in-the-Loop," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.4. ISBN: 9798244538229.