Safety & Security · Data artifact

Moderation Threshold Policy

Data artifactSafety & SecuritySafety, Security & Governancearc:ModerationThresholdPolicy

A configuration of classifier score thresholds that separate auto-allow, human-review and auto-block bands for moderated content, tuned per use case.

Responsibility. Defines the score bands that decide allow, review or block.

Also known as: Toxicity threshold, Review/block thresholds

configuresconfiguresToxicity Classifier: configuresToxicity ClassifierModeration Triage Router: configuresModeration Triage Router
Direct neighbourhood (hover for relationship types)

Relationships

configures structural

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Sources

  1. Ch9.1: T. Nguyen, "Output Filtering and Content Moderation," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.1. ISBN: 9798244538229.
  2. Ref9.04: "Safety Guardrails Implementation for Agent Systems," unpublished reference note (04-Safety-Guardrails-Implementation.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note