Model Serving · Model asset
Draft Token Proposer
Model assetModel ServingModelsVariation point (abstract)arc:DraftTokenProposer
Model weights that propose K speculative future tokens for a target model to verify in one parallel forward pass, selected for distributional alignment with the target rather than task accuracy.
Responsibility. Supplies speculative tokens whose acceptance multiplies decode throughput.
Also known as: Draft model, Draft strategy
Variants
| Variant | When to choose |
|---|---|
| Self-Speculative Decoding Head | Choose when memory is constrained or batch size must be preserved (e.g., edge devices, high-batch serving). |
| Speculative Draft Model abstract | — |
Relationships
deployed on structural
is evaluated by assurance
Design guidance
- SHOULD be chosen by minimizing total variation distance to the target on 1000-2000 production queries, not by benchmark accuracy.
Quantitative guidance
As stated by the sources; verify before use.
- Acceptance alpha is approximately 1 - TVD; a 60% MMLU 3B draft reached alpha=0.80 while a 72% MMLU 7B draft reached alpha=0.65 (Ch7.1A).
Sources
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.