Knowledge & Data · Data artifact
Deduplication Threshold Configuration
Data artifactKnowledge & DataKnowledge & Dataarc:DeduplicationThresholdConfig
A configuration of fuzzy and semantic similarity thresholds that trades duplicate recall against false-positive removal of legitimately different documents.
Responsibility. Sets the similarity thresholds at which documents are treated as duplicates.
Also known as: Similarity thresholds
Relationships
configures structural
Design guidance
- SHOULD be set according to whether coverage or precision of deduplication is prioritised.
- SHOULD use higher similarity thresholds for conversational transcripts, whose patterns overlap more than written documents.
Quantitative guidance
As stated by the sources; verify before use.
- Conservative: fuzzy 0.95 / semantic 0.98; aggressive: fuzzy 0.80 / semantic 0.85 (Ch6.4).
Classification
- Patterns
- Conservative vs aggressive thresholding
- Quality attributes
- Maintainability (ISO/IEC 25010)
- Risks mitigated
- False-positive duplicate removalMissed near-duplicates
Sources
- Ch6.4: T. Nguyen, "Data Quality Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.4. ISBN: 9798244538229.
- Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.