Model Adaptation · Software component

Partitioned Dataset Reader

Software componentModel AdaptationModelsarc:PartitionedDatasetReader

A curation-stage loader that lazily streams a large line-delimited document dataset from disk-backed storage in batches sized to fit accelerator memory, validating record schema on load.

Responsibility. Feeds dataset records to downstream curation stages in memory-bounded batches.

Also known as: Lazy dataset loader, Document dataset loader

is orchestrated bysends data toreadsis configured byData Curator: is orchestrated byData CuratorLanguage Identification Filter: sends data toLanguage Identification …Raw Document Corpus: readsRaw Document CorpusSource Document Schema: is configured bySource Document Schema
Direct neighbourhood (hover for relationship types)

Relationships

is configured by structural

reads dependency

sends data to dynamic

is orchestrated by control

Design guidance

Classification

Patterns
Lazy batch streaming
Technologies
NeMo Curator DocumentDataset
Quality attributes
Performance efficiency (ISO/IEC 25010)
Risks mitigated
Out-of-memory failure when loading multi-terabyte datasets

Sources

  1. Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.