Model Adaptation · Software component
Perplexity Filter
Software componentModel AdaptationModelsarc:PerplexityFilter
A curation filter that computes each document's perplexity under a reference language model and discards documents exceeding a threshold as random, spammy or corrupted text.
Responsibility. Removes documents whose text is far from the distribution of clean natural language.
Also known as: Perplexity-based quality filter
Relationships
hosts structural
is configured by structural
receives data from dynamic
sends data to dynamic
is orchestrated by control
Design guidance
- SHOULD fine-tune the reference model on clean domain data if the filter removes too many valid domain documents.
Quantitative guidance
As stated by the sources; verify before use.
- Typical threshold >1500 perplexity; generic web text scores 200-400 and domain jargon 500-800 (Ch7.5).
- Domain recalibration uses 10,000-100,000 clean domain examples (Ch7.5).
Classification
- Patterns
- Two-stage filtering (generic model then domain-adapted model)
- Technologies
- NVIDIA NeMo Curator
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)
- Risks mitigated
- OCR errors, spam and malformed HTML-to-text conversions in training data
Sources
- Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.