Model Adaptation · Model asset

Reference Policy Model

Model assetModel AdaptationModelsarc:ReferencePolicyModel

A frozen copy of the supervised-fine-tuned model whose token distributions anchor preference optimization, against which divergence of the evolving policy is measured and penalized.

Responsibility. Provides the fixed reference distribution for KL-divergence regularization or DPO's implicit reward.

Also known as: Reference model, Frozen SFT model

Variant of Foundation LLM abstract

is trained byspecializesdeployed ondeployed onFine-Tuning Pipeline: is trained byFine-Tuning PipelineFoundation LLM: specializesFoundation LLMRLHF Policy Optimizer: deployed onRLHF Policy OptimizerDirect Preference Optimizer: deployed onDirect Preference Optimi…
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

is trained by lifecycle

Classification

Patterns
KL-divergence regularization
Risks mitigated
Reward hackingCatastrophic forgettingLoss of generation coherence

Sources

  1. Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.