Model Serving · Model asset
Distilled Draft Model
Model assetModel ServingModelsarc:DistilledDraftModel
A separate small draft model trained by knowledge distillation to minimize KL divergence from the target model's logits, learning the target's preferences rather than accuracy.
Responsibility. Provides high-acceptance speculative drafts aligned to the target distribution.
Variant of Speculative Draft Model abstract
When to choose. Choose for production where improved acceptance justifies 50-100 GPU-hours of training; use online distillation at scale to track target shifts.
Relationships
is trained by lifecycle
alternative to variability
Quantitative guidance
As stated by the sources; verify before use.
- alpha=0.7-0.85, 2-3x speedup; online distillation sustains alpha=0.75-0.90 (Ch7.1A).
- Distilled drafts outperform general LLM drafts by 40-111%; 3B distilled achieves 2.5x vs 1.5x for independent 3B (Ch7.1A).
Classification
- Patterns
- Knowledge distillationOnline distillation
Sources
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.