Model Serving · Model asset

Draft Token Proposer

Model assetModel ServingModelsVariation point (abstract)arc:DraftTokenProposer

Model weights that propose K speculative future tokens for a target model to verify in one parallel forward pass, selected for distributional alignment with the target rather than task accuracy.

Responsibility. Supplies speculative tokens whose acceptance multiplies decode throughput.

Also known as: Draft model, Draft strategy

is evaluated bydeployed onis specialized byis specialized byEvaluation Harness: is evaluated byEvaluation HarnessSpeculative Decoder: deployed onSpeculative DecoderSpeculative Draft Model: is specialized bySpeculative Draft ModelSelf-Speculative Decoding Head: is specialized bySelf-Speculative Decodin…
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
Self-Speculative Decoding HeadChoose when memory is constrained or batch size must be preserved (e.g., edge devices, high-batch serving).
Speculative Draft Model abstract—

Relationships

deployed on structural

is evaluated by assurance

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Sources

  1. Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.