Model Adaptation · Software component
Fine-Tuning Pipeline
Software componentModel AdaptationModelsVariation point (abstract)arc:FineTuningPipeline
A software component that adapts model weights to a domain or task using curated training data.
Responsibility. Adapts model weights to a domain or task.
Also known as: Supervised fine-tuning (SFT), Trajectory fine-tuning, SL-CAI supervised phase, Supervised Fine-Tuning (SFT) stage
Variants
| Variant | When to choose |
|---|---|
| Full-Parameter Fine-Tuner | Choose when substantial multi-GPU infrastructure is available and all model weights are to be modified. |
| LoRA Fine-Tuner abstract | Choose when GPU memory is constrained and near-full-fine-tuning quality is needed by training only a small adapter. |
Relationships
reads dependency
- Agent Trajectory Dataset Ch3.5
- Annotated Trace Dataset Ch3.6
- Critique-Revision Dataset Ch9.5
- Curated Training Corpus Ch7.5 Ch10.5
- Domain Text Corpus Ref6.01
- Instruction Demonstration Dataset Ch10.3
- Labeled Decision Case Dataset Ch10.4 Ref10.01
- Output Correction Dataset Ch10.5
- Production Feedback Dataset Ch10.1
- Synthetic Dataset Ch3.6 Ref1.05
is triggered by dynamic
receives data from dynamic
- Feedback Collector abstract Ref7.14
triggers dynamic
- Preference Optimizer abstract Ch3.5 Ch10.3
is orchestrated by control
trains lifecycle
- Constitutionally Aligned Model Ch9.5
- Cross-Modal Reranking Model Ch2.7
- Fine-Tuned Agent Model Ch3.5 Ch10.1 +3
- Foundation LLM abstract Ch10.3 Ref7.14 +1
- Intent Classifier Model Ch10.1
- NER Model Ch1.7A
- Reference Policy Model Ch10.3
- Tool Call Verifier Model Ch3.7
- Toxicity Classification Model Ch9.1
- Verifier Model Ch3.6
Design guidance
- SHOULD be used only after prompt engineering and RAG fail to resolve behavioral-consistency problems.
- SHOULD start with conservative learning rates and stop training when held-out validation performance peaks.
- SHOULD be treated as an ongoing process with periodic retraining as production distributions drift.
- SHOULD verify through small pilots that fine-tuning improves on prompt engineering and RAG before scaling.
- SHOULD be chosen when substantial domain data exists and the domain is stable; knowledge is hard to update and catastrophic forgetting is a risk (Ref6.01).
- SHOULD include explicit alignment maintenance when fine-tuning aligned foundation models on domain data, since the new objective can override constitutional principles.
- SHOULD precede preference optimization with SFT on curated instruction-response demonstrations; it is optional but makes the subsequent RL phase substantially more efficient.
Quantitative guidance
As stated by the sources; verify before use.
- Conservative learning rates of 1e-5 to 5e-5 are industry practice for LLM fine-tuning (Ch3.5).
- Illustrative hierarchy: prompt engineering ~70% resolution, adding RAG ~80%, remaining ~20% failures motivate fine-tuning (Ch3.5).
- InstructGPT fine-tuned GPT-3 on approximately 13,000 demonstrations of correct instruction-following (Ch10.3).
- Fine-tuning a large model on demonstrations might complete in days, versus weeks or months for full RLHF training (Ch10.3).
Classification
- Patterns
- Domain fine-tuningSupervised fine-tuning on agent trajectoriesEarly stopping on validation performanceData parallelism (distributed data loading, centralized parameter updates)Supervised fine-tuning on instruction demonstrations (RLHF phase one)Data flywheel (feedback-driven retraining)Supervised fine-tuning
- Technologies
- NVIDIA NeMo CustomizerNVIDIA NeMo Framework
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
- Risks mitigated
- Behavioral inconsistency unresolved by prompting or RAGOverfittingCatastrophic forgettingFine-tuning drift eroding alignment
Sources
- Ch1.7A: T. Nguyen, "Relational Reasoning with Knowledge Graphs - The Fundamentals, Integration, and Extraction," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.7A. ISBN: 9798244538229.
- Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.
- Ch3.5: T. Nguyen, "Prompt Optimization, Few-Shot Learning, Fine-Tuning," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.5. ISBN: 9798244538229.
- Ch3.6: T. Nguyen, "Trace Analysis and Execution Debugging," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.6. ISBN: 9798244538229.
- Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.
- Ch9.1: T. Nguyen, "Output Filtering and Content Moderation," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.1. ISBN: 9798244538229.
- Ch9.5: T. Nguyen, "Constitutional AI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.5. ISBN: 9798244538229.
- Ch10.1: T. Nguyen, "Conversational UI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.1. ISBN: 9798244538229.
- Ch10.2: T. Nguyen, "Proactive Agents," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.2. ISBN: 9798244538229.
- Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.
- Ch10.4: T. Nguyen, "Human-in-the-Loop," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.4. ISBN: 9798244538229.
- Ch10.5: T. Nguyen, "Human-over-the-Loop," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.5. ISBN: 9798244538229.
- Ref6.01: S. Schürch, "How to Make Your LLM More Accurate with RAG & Fine-Tuning," Towards Data Science, Mar. 11, 2025. [Online]. Available: https://towardsdatascience.com/how-to-make-your-llm-more-accurate-with-rag-fine-tuning/
- Ref7.14: "NVIDIA Agentic AI Platform Ecosystem Integration," unpublished reference note (14-NVIDIA-Ecosystem-Integration.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.17: "Scaling Agentic AI Systems: Patterns and Strategies," unpublished reference note (17-Scalability-Patterns.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.18: "Chapter 7 Summary: NVIDIA Platform Implementation," unpublished reference note (18-Chapter-7-Summary.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref8.06: "Error Troubleshooting and Incident Response for Agent Systems," unpublished reference note (06-Error-Troubleshooting-Incident-Response.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref10.01: "Human-in-the-Loop Systems for Agent Interactions," unpublished reference note (01-Human-in-the-Loop-Systems.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note