Knowledge & Data · Software component
Cascading Deduplicator
Software componentKnowledge & DataKnowledge & Dataarc:CascadingDeduplicator
A deduplicator that runs exact, fuzzy and semantic deduplication sequentially so each more expensive level only processes survivors of the cheaper ones.
Responsibility. Sequences deduplication levels from cheapest to most expensive.
Also known as: Three-level deduplication, Multi-level deduplication pipeline
Variant of Content Deduplicator abstract
When to choose. Choose when a corpus contains exact, near and semantic duplicates together (typical in production); orders cheap levels first to minimise cost.
Relationships
receives data from dynamic
is orchestrated by control
orchestrates control
alternative to variability
Design guidance
- SHOULD run fast methods first and expensive semantic comparison last.
Quantitative guidance
As stated by the sources; verify before use.
- Reduces deduplication of 10,000 documents from ~15 min (semantic-only) to ~4 min; 10-15x speedup at 100,000 documents (Ch6.4).
Classification
- Patterns
- Cost-ordered cascade
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)
- Risks mitigated
- Systematic duplicates left by single-level approaches
Sources
- Ch6.4: T. Nguyen, "Data Quality Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.4. ISBN: 9798244538229.