Knowledge & Data · Data artifact

Deduplication Threshold Configuration

Data artifactKnowledge & DataKnowledge & Dataarc:DeduplicationThresholdConfig

A configuration of fuzzy and semantic similarity thresholds that trades duplicate recall against false-positive removal of legitimately different documents.

Responsibility. Sets the similarity thresholds at which documents are treated as duplicates.

Also known as: Similarity thresholds

configuresconfiguresContent Deduplicator: configuresContent DeduplicatorNear-Duplicate Detector: configuresNear-Duplicate Detector
Direct neighbourhood (hover for relationship types)

Relationships

configures structural

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Conservative vs aggressive thresholding
Quality attributes
Maintainability (ISO/IEC 25010)
Risks mitigated
False-positive duplicate removalMissed near-duplicates

Sources

  1. Ch6.4: T. Nguyen, "Data Quality Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.4. ISBN: 9798244538229.
  2. Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.