Model Adaptation · Software component
Continued Pretrainer
Software componentModel AdaptationModelsarc:ContinuedPretrainer
A training component that further pretrains a base language model on a domain text corpus with the next-token prediction objective before task-specific fine-tuning.
Responsibility. Internalizes domain knowledge into a base model before supervised fine-tuning.
Also known as: Continuous pretraining (CPT), Domain-adaptive pretraining, Domain-adaptive pretraining (DAPT)
Relationships
reads dependency
triggers dynamic
- Fine-Tuning Pipeline abstract Ch3.5
is orchestrated by control
trains lifecycle
Design guidance
- SHOULD precede supervised fine-tuning when agents need substantial domain knowledge, so SFT can focus on task reasoning.
Quantitative guidance
As stated by the sources; verify before use.
- CPT on JP1 middleware documentation followed by SFT gave a 16% improvement on domain certification exams versus SFT only (Hitachi, Ch3.5).
- CPT typically runs for days to weeks depending on corpus scale (Ch3.5).
Classification
- Patterns
- Continued pretraining then SFT (CPT+SFT)
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Performance efficiency (ISO/IEC 25010)
- Risks mitigated
- Agents misinterpreting domain concepts in trajectories
Sources
- Ch3.5: T. Nguyen, "Prompt Optimization, Few-Shot Learning, Fine-Tuning," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.5. ISBN: 9798244538229.
- Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.
- Ch10.5: T. Nguyen, "Human-over-the-Loop," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.5. ISBN: 9798244538229.