Model Adaptation · Data artifact
Instruction Demonstration Dataset
Data artifactModel AdaptationModelsarc:InstructionDemonstrationDataset
A curated training dataset of instructions paired with high-quality responses demonstrating the desired style, tone and approach, used for supervised fine-tuning before preference optimization.
Responsibility. Supplies direct positive examples of ideal instruction-following behaviour to the SFT stage.
Also known as: SFT demonstration set, Demonstrations of ideal behavior
Relationships
is read by dependency
- Fine-Tuning Pipeline abstract Ch10.3
receives data from dynamic
Design guidance
- MAY be used alongside preference comparisons; demonstrations are particularly valuable early in training.
Quantitative guidance
As stated by the sources; verify before use.
- InstructGPT used approximately 13,000 carefully crafted demonstrations (Ch10.3).
Sources
- Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.