Model Adaptation · Software component
Training Pipeline Orchestrator
Software componentModel AdaptationModelsarc:TrainingPipelineOrchestrator
A model-adaptation orchestrator that sequences data curation, continued pretraining, supervised fine-tuning, reward modeling and preference optimization stages into a reproducible, automated pipeline.
Responsibility. Coordinates execution of model-adaptation stages as one reproducible pipeline.
Also known as: Customization pipeline, RLHF pipeline orchestration, Data flywheel
Relationships
reads dependency
writes dependency
is triggered by dynamic
triggers dynamic
orchestrates control
- AI-Feedback Preference Labeler Ch9.5
- Candidate Response Sampler Ch10.3
- Continued Pretrainer Ch3.5
- Critique-Revision Generator Ch9.5
- Data Curator Ch3.5 Ch10.1 +1
- Fine-Tuning Pipeline abstract Ch3.5 Ch9.5 +3
- Rollout Manager abstract Ch10.2
- Preference Optimizer abstract Ch3.5
- RLHF Policy Optimizer Ch9.5 Ch10.3
- Regression Gate Ch10.2
- Reward Model Trainer Ch3.5 Ch9.5 +1
Design guidance
- SHOULD run supervised fine-tuning before preference learning so the policy is competent before preference optimization.
- MUST keep training data pipelines separate from production customer data.
- SHOULD coordinate generation (memory-bound), reward scoring (compute-bound, batchable), KL computation and PPO updates (synchronization-heavy) as distinct stages.
Quantitative guidance
As stated by the sources; verify before use.
- Klarna feeds feedback into weekly retraining; repeat inquiry rates declined 25% year-over-year (Ch10.1).
- Feedback implementation cycle: collect wk1-2, analyze wk2-3, design wk3-4, implement wk5-8, test wk8-9, release wk9, measure wk10+ (Ref10.03).
- AT&T customer service: steady improvement over 18 months across dozens of flywheel cycles, improving satisfaction and inference cost (Ch10.2).
Classification
- Patterns
- Three-stage RLHF (SFT, reward model, RL)Continuous improvement loopTwo-phase Constitutional AI (SL-CAI then RL-CAI)Iterative alternation of RLAIF and RLHFDifferential privacy during trainingData flywheelWeekly retraining cycle
- Technologies
- NVIDIA NeMo CustomizerMLOps platforms
- Quality attributes
- Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Performance efficiency (ISO/IEC 25010)
- Risks mitigated
- Error-prone manual multi-stage training workflowsInability to remove individual training examples from learned weights
Sources
- Ch3.5: T. Nguyen, "Prompt Optimization, Few-Shot Learning, Fine-Tuning," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.5. ISBN: 9798244538229.
- Ch4.3: T. Nguyen, "Container Orchestration and Edge Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.3. ISBN: 9798244538229.
- Ch9.5: T. Nguyen, "Constitutional AI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.5. ISBN: 9798244538229.
- Ch9.6: T. Nguyen, "Value Alignment Frameworks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.6. ISBN: 9798244538229.
- Ch9.7: T. Nguyen, "GDPR and Data Protection Regulations," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.7. ISBN: 9798244538229.
- Ch10.1: T. Nguyen, "Conversational UI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.1. ISBN: 9798244538229.
- Ch10.2: T. Nguyen, "Proactive Agents," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.2. ISBN: 9798244538229.
- Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.
- Ref10.03: "User Feedback and Iterative Improvement," unpublished reference note (03-User-Feedback-Iteration.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note