Knowledge & Data · Software component

Fuzzy Text Deduplicator

Software componentKnowledge & DataKnowledge & Dataarc:FuzzyTextDeduplicator

A deduplicator that compares documents with a length-normalized string distance and removes those whose similarity exceeds a configured threshold.

Responsibility. Removes near-duplicate documents with minor textual differences.

Also known as: Fuzzy matching, Near-duplicate detection

Variant of Content Deduplicator abstract

When to choose. Choose when near-duplicates differ by minor edits, dates or formatting (versioning, format conversion); tune the threshold between conservative and aggressive.

specializesis target of alternativeTois orchestrated byalternative toContent Deduplicator: specializesContent DeduplicatorExact Hash Deduplicator: is target of alternativeToExact Hash DeduplicatorCascading Deduplicator: is orchestrated byCascading DeduplicatorSemantic Deduplicator: alternative toSemantic Deduplicator
Direct neighbourhood (hover for relationship types)

Relationships

is orchestrated by control

alternative to variability

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Normalized Levenshtein similarity
Quality attributes
Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)
Risks mitigated
Near-duplicates missed by hashing

Sources

  1. Ch6.4: T. Nguyen, "Data Quality Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.4. ISBN: 9798244538229.