Model Serving · Model asset

Speculative Draft Model

Model assetModel ServingModelsVariation point (abstract)arc:SpeculativeDraftModel

A small, fast language model that proposes multi-token candidate continuations for a larger target model to verify in one forward pass.

Responsibility. Drafts candidate tokens for target-model verification.

Also known as: Draft model, Two-model draft

Variant of Draft Token Proposer abstract

deployed onspecializesis specialized byis target of alternativeTois configured byis specialized byInference Server: deployed onInference ServerDraft Token Proposer: specializesDraft Token ProposerDistilled Draft Model: is specialized byDistilled Draft ModelSelf-Speculative Decoding Head: is target of alternativeToSelf-Speculative Decodin…Speculative Decoding Configuration: is configured bySpeculative Decoding Con…Independent Draft Model: is specialized byIndependent Draft Model
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
Distilled Draft ModelChoose for production where improved acceptance justifies 50-100 GPU-hours of training; use online distillation at scale to track target shifts.
Independent Draft ModelChoose for rapid prototyping where zero training cost and immediate deployment outweigh modest acceptance.

Relationships

deployed on structural

is configured by structural

alternative to variability

Quantitative guidance

As stated by the sources; verify before use.

Classification

Technologies
Llama-3.1-8B (draft for 70B)Llama-3.2-3B (draft for 8B)

Sources

  1. Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
  2. Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.