Model Serving · Model asset
Speculative Draft Model
Model assetModel ServingModelsVariation point (abstract)arc:SpeculativeDraftModel
A small, fast language model that proposes multi-token candidate continuations for a larger target model to verify in one forward pass.
Responsibility. Drafts candidate tokens for target-model verification.
Also known as: Draft model, Two-model draft
Variant of Draft Token Proposer abstract
Variants
| Variant | When to choose |
|---|---|
| Distilled Draft Model | Choose for production where improved acceptance justifies 50-100 GPU-hours of training; use online distillation at scale to track target shifts. |
| Independent Draft Model | Choose for rapid prototyping where zero training cost and immediate deployment outweigh modest acceptance. |
Relationships
deployed on structural
is configured by structural
alternative to variability
Quantitative guidance
As stated by the sources; verify before use.
- Typically 5-10x smaller than the target with 60-80% token-level agreement (Ch4.4).
- 3B draft reached only 48% acceptance; 13B draft overhead negated benefits; 8B (8.7x smaller) was near-optimal for a 70B target (Ch4.4).
- Draft 7B + target 70B needs ~77GB FP16 (154GB with KV cache/activations); on A100 80GB draft reduces target batch from 16 to 8-10 (Ch7.1A).
Classification
- Technologies
- Llama-3.1-8B (draft for 70B)Llama-3.2-3B (draft for 8B)
Sources
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.