Knowledge & Data · Software component
Fuzzy Text Deduplicator
Software componentKnowledge & DataKnowledge & Dataarc:FuzzyTextDeduplicator
A deduplicator that compares documents with a length-normalized string distance and removes those whose similarity exceeds a configured threshold.
Responsibility. Removes near-duplicate documents with minor textual differences.
Also known as: Fuzzy matching, Near-duplicate detection
Variant of Content Deduplicator abstract
When to choose. Choose when near-duplicates differ by minor edits, dates or formatting (versioning, format conversion); tune the threshold between conservative and aggressive.
Relationships
is orchestrated by control
alternative to variability
Design guidance
- SHOULD calibrate the similarity threshold through manual review of edge cases.
Quantitative guidance
As stated by the sources; verify before use.
- Conservative threshold 0.95 vs aggressive 0.80; 30-60 s for ~6,000 documents, catching ~30% more duplicates (Ch6.4).
Classification
- Patterns
- Normalized Levenshtein similarity
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)
- Risks mitigated
- Near-duplicates missed by hashing
Sources
- Ch6.4: T. Nguyen, "Data Quality Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.4. ISBN: 9798244538229.