Model Adaptation · Software component
Data Curator
Software componentModel AdaptationModelsarc:DataCurator
A model-adaptation component that filters, deduplicates and quality-scores raw trajectories or demonstrations, retaining only examples that satisfy outcome-based quality criteria.
Responsibility. Admits only high-quality examples into training and demonstration data.
Also known as: Rejection fine-tuning filter, Trajectory quality filter, Noisy example filtering, Data curation pipeline, GPU-accelerated data curation toolkit, Training corpus curation pipeline, Bias prefiltering, Flywheel data processing
Relationships
deployed on structural
invokes dependency
reads dependency
writes dependency
receives data from dynamic
is constrained by control
is guarded by control
is orchestrated by control
orchestrates control
produces lifecycle
Design guidance
- SHOULD retain only trajectories meeting objective success criteria such as passing tests or confirmed resolution.
- SHOULD align filtering criteria with the performance profile the fine-tuned model must exhibit, using composite metrics for multiple objectives.
- SHOULD tighten filtering criteria in each training iteration.
- SHOULD prioritize training-data quality over quantity and ensure data represents domain variation (Ref6.01).
- SHOULD treat curation as an iterative loop: curate with initial thresholds, train a small model, evaluate on held-out benchmarks, analyze failures, then adjust thresholds.
- SHOULD keep the core stages (language ID, quality filtering, deduplication, PII redaction) fixed and adapt thresholds and domain classifiers per data source.
- SHOULD run curation stages on GPUs so that several threshold configurations can be tested empirically rather than guessed.
Quantitative guidance
As stated by the sources; verify before use.
- Filtering noisy few-shot examples raised intent-classification accuracy from 59.6% to over 80% (Cleanlab, Ch3.5).
- Training on only the top ~30% of trajectories by quality often outperforms training on the full dataset (Ch3.5).
- Three progressive filter-train-regenerate iterations raised software-agent task success from ~40% to ~75% (Ch3.5).
- GPU-accelerated curation delivers 16x speedup over CPU alternatives for document processing and 89x for video curation (Ch7.1A).
- GPU-accelerated curation stages run 16-89x faster than CPU-based alternatives (Ch7.5).
- Financial-services example reduced 5 billion documents (10TB) to 800 million, 17% of original size, in under 24 hours vs. ~2 weeks on CPUs (Ch7.5).
- Curation costs tens of GPU-hours vs. thousands for training; ~1% of budget can improve final model quality by 20-40% (Ch7.5).
- Call-transcript example reduced 10 million raw transcripts to ~2.5 million curated examples (75% reduction) (Ch7.5).
Classification
- Patterns
- Rejection fine-tuning (RFT)Composite quality-metric filteringProgressive iterative refinementComposable multi-stage curation pipelineIterative curate-train-evaluate-refine loopData flywheel
- Technologies
- NVIDIA NeMo CuratorCleanlab
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Performance efficiency (ISO/IEC 25010)Maintainability (ISO/IEC 25010)
- Risks mitigated
- Learning mistakes from low-quality trajectoriesNoisy demonstrations degrading few-shot accuracyOverfitting on limited or low-quality fine-tuning dataModel learning noise, spam and OCR corruption from unfiltered web scrapesMemorization from duplicated training dataPII memorization and regurgitationBiased patterns entering learned representations
Sources
- Ch3.5: T. Nguyen, "Prompt Optimization, Few-Shot Learning, Fine-Tuning," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.5. ISBN: 9798244538229.
- Ch4.3: T. Nguyen, "Container Orchestration and Edge Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.3. ISBN: 9798244538229.
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
- Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.
- Ch9.1: T. Nguyen, "Output Filtering and Content Moderation," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.1. ISBN: 9798244538229.
- Ch10.1: T. Nguyen, "Conversational UI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.1. ISBN: 9798244538229.
- Ch10.2: T. Nguyen, "Proactive Agents," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.2. ISBN: 9798244538229.
- Ch10.5: T. Nguyen, "Human-over-the-Loop," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.5. ISBN: 9798244538229.
- Ref6.01: S. Schürch, "How to Make Your LLM More Accurate with RAG & Fine-Tuning," Towards Data Science, Mar. 11, 2025. [Online]. Available: https://towardsdatascience.com/how-to-make-your-llm-more-accurate-with-rag-fine-tuning/