Model Adaptation · Data artifact
Raw Document Corpus
Data artifactModel AdaptationModelsarc:RawDocumentCorpus
An uncurated collection of documents or transcripts in line-delimited JSON, each record holding text plus metadata such as source URL, extraction timestamp and document type, awaiting curation.
Responsibility. Supplies raw candidate training text and its provenance metadata to the curation pipeline.
Also known as: Raw scraped dataset, JSON Lines dataset, Raw call transcripts
Relationships
is read by dependency
is produced by lifecycle
Quantitative guidance
As stated by the sources; verify before use.
- Financial example starts at 5 billion documents / 10TB mixing earnings reports, spam, OCR-corrupted PDFs and duplicated press releases (Ch7.5).
Classification
- Technologies
- JSON Lines
- Quality attributes
- Transparency and accountability (NIST AI RMF: accountable and transparent)
Sources
- Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.