Safety & Security · Data artifact
Moderation Threshold Policy
Data artifactSafety & SecuritySafety, Security & Governancearc:ModerationThresholdPolicy
A configuration of classifier score thresholds that separate auto-allow, human-review and auto-block bands for moderated content, tuned per use case.
Responsibility. Defines the score bands that decide allow, review or block.
Also known as: Toxicity threshold, Review/block thresholds
Relationships
configures structural
Design guidance
- SHOULD define a review band between an allow threshold and a block threshold so that borderline cases reach humans.
- SHOULD be tuned against labeled datasets reflecting the deployment's users and risk tolerance.
Quantitative guidance
As stated by the sources; verify before use.
- Healthcare example: review threshold 0.5, auto-block threshold 0.8; a response scored 0.65 is routed to a human moderator (Ch9.1).
Sources
- Ch9.1: T. Nguyen, "Output Filtering and Content Moderation," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.1. ISBN: 9798244538229.
- Ref9.04: "Safety Guardrails Implementation for Agent Systems," unpublished reference note (04-Safety-Guardrails-Implementation.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note