Model Adaptation · Software component
Partitioned Dataset Reader
Software componentModel AdaptationModelsarc:PartitionedDatasetReader
A curation-stage loader that lazily streams a large line-delimited document dataset from disk-backed storage in batches sized to fit accelerator memory, validating record schema on load.
Responsibility. Feeds dataset records to downstream curation stages in memory-bounded batches.
Also known as: Lazy dataset loader, Document dataset loader
Relationships
is configured by structural
reads dependency
sends data to dynamic
is orchestrated by control
Design guidance
- SHOULD process documents in batches that fit GPU memory rather than loading the whole dataset into RAM, to scale to petabyte datasets.
Classification
- Patterns
- Lazy batch streaming
- Technologies
- NeMo Curator DocumentDataset
- Quality attributes
- Performance efficiency (ISO/IEC 25010)
- Risks mitigated
- Out-of-memory failure when loading multi-terabyte datasets
Sources
- Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.