Model Adaptation · Software component

Language Identification Filter

Software componentModel AdaptationModelsarc:LanguageIdentificationFilter

A curation filter that predicts each document's language with a pretrained classifier and retains only documents in the configured target languages.

Responsibility. Removes documents whose language does not match the deployment's target languages.

Also known as: Language ID stage, Language filter

is orchestrated bysends data toreceives data fromhostsData Curator: is orchestrated byData CuratorDocument Quality Filter: sends data toDocument Quality FilterPartitioned Dataset Reader: receives data fromPartitioned Dataset ReaderLanguage Identification Model: hostsLanguage Identification …
Direct neighbourhood (hover for relationship types)

Relationships

hosts structural

receives data from dynamic

sends data to dynamic

is orchestrated by control

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Embarrassingly parallel per-document classification
Technologies
NVIDIA NeMo CuratorfastText
Quality attributes
Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Performance efficiency (ISO/IEC 25010)
Risks mitigated
Dilution of target-language training densityModel capacity spent on unused languages

Sources

  1. Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.