Model Adaptation · Software component

Training Pipeline Orchestrator

Software componentModel AdaptationModelsarc:TrainingPipelineOrchestrator

A model-adaptation orchestrator that sequences data curation, continued pretraining, supervised fine-tuning, reward modeling and preference optimization stages into a reproducible, automated pipeline.

Responsibility. Coordinates execution of model-adaptation stages as one reproducible pipeline.

Also known as: Customization pipeline, RLHF pipeline orchestration, Data flywheel

orchestratesorchestratesorchestratesorchestratesis triggered byorchestratesis triggered byorchestratesorchestratesreadsorchestratesorchestratesorchestratesreadsorchestratestriggerswritesRollout Manager: orchestratesRollout ManagerFine-Tuning Pipeline: orchestratesFine-Tuning PipelineData Curator: orchestratesData CuratorRegression Gate: orchestratesRegression GateQuality Drift Detector: is triggered byQuality Drift DetectorRLHF Policy Optimizer: orchestratesRLHF Policy OptimizerAlignment Drift Monitor: is triggered byAlignment Drift MonitorAI-Feedback Preference Labeler: orchestratesAI-Feedback Preference L…Critique-Revision Generator: orchestratesCritique-Revision Genera…Synthetic Dataset: readsSynthetic DatasetCandidate Response Sampler: orchestratesCandidate Response SamplerPreference Optimizer: orchestratesPreference OptimizerContinued Pretrainer: orchestratesContinued PretrainerQuality Improvement Backlog: readsQuality Improvement Back…Reward Model Trainer: orchestratesReward Model TrainerImprovement Announcer: triggersImprovement AnnouncerTraining Data Lineage Store: writesTraining Data Lineage St…
Direct neighbourhood (hover for relationship types)

Relationships

reads dependency

writes dependency

is triggered by dynamic

triggers dynamic

orchestrates control

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Three-stage RLHF (SFT, reward model, RL)Continuous improvement loopTwo-phase Constitutional AI (SL-CAI then RL-CAI)Iterative alternation of RLAIF and RLHFDifferential privacy during trainingData flywheelWeekly retraining cycle
Technologies
NVIDIA NeMo CustomizerMLOps platforms
Quality attributes
Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Performance efficiency (ISO/IEC 25010)
Risks mitigated
Error-prone manual multi-stage training workflowsInability to remove individual training examples from learned weights

Sources

  1. Ch3.5: T. Nguyen, "Prompt Optimization, Few-Shot Learning, Fine-Tuning," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.5. ISBN: 9798244538229.
  2. Ch4.3: T. Nguyen, "Container Orchestration and Edge Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.3. ISBN: 9798244538229.
  3. Ch9.5: T. Nguyen, "Constitutional AI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.5. ISBN: 9798244538229.
  4. Ch9.6: T. Nguyen, "Value Alignment Frameworks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.6. ISBN: 9798244538229.
  5. Ch9.7: T. Nguyen, "GDPR and Data Protection Regulations," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.7. ISBN: 9798244538229.
  6. Ch10.1: T. Nguyen, "Conversational UI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.1. ISBN: 9798244538229.
  7. Ch10.2: T. Nguyen, "Proactive Agents," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.2. ISBN: 9798244538229.
  8. Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.
  9. Ref10.03: "User Feedback and Iterative Improvement," unpublished reference note (03-User-Feedback-Iteration.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note