Knowledge & Data · Software component
Near-Duplicate Detector
Software componentKnowledge & DataKnowledge & Dataarc:NearDuplicateDetector
A content deduplicator that uses locality-sensitive signatures to detect near-identical documents above a configurable similarity threshold.
Responsibility. Removes near-duplicate documents that differ only slightly.
Also known as: Fuzzy deduplication, Fuzzy deduplicator, MinHash deduplicator
Variant of Content Deduplicator abstract
When to choose. Choose when content varies slightly (whitespace, minor edits, reformatting) so exact hashing misses duplicates.
Relationships
deployed on structural
is configured by structural
receives data from dynamic
sends data to dynamic
is orchestrated by control
alternative to variability
Design guidance
- SHOULD run after exact deduplication to avoid O(N^2) pairwise comparison over billions of documents.
Quantitative guidance
As stated by the sources; verify before use.
- Fuzzy deduplication of 50M documents: 38 hours on 64 CPU cores vs 2.8 hours on 4x A100 (13.6x) (Ch6.3B).
- Example uses 5-character n-grams, 128 hash functions and Jaccard similarity >0.85 as the duplicate threshold (Ch7.5).
Classification
- Patterns
- MinHashSimHashMinHash LSHCharacter n-gram shingling
- Technologies
- NVIDIA NeMo Curator
- Risks mitigated
- Near-duplicates differing only in dates, formatting or boilerplate evading exact deduplication
Sources
- Ch6.3A: T. Nguyen, "ETL Pipeline Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.3A. ISBN: 9798244538229.
- Ch6.3B: T. Nguyen, "ETL Worked Example - Load Phase & Pipeline Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.3B. ISBN: 9798244538229.
- Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.