Model Serving · Data artifact
Speculative Decoding Configuration
Data artifactModel ServingModelsarc:SpeculativeDecodingConfig
A configuration selecting the draft model, target model and speculation window length for speculative decoding.
Responsibility. Fixes draft-target pairing and speculation window.
Relationships
configures structural
is written by dependency
Quantitative guidance
As stated by the sources; verify before use.
- Production windows typically 4-8 tokens, averaging 4-5 accepted tokens per target pass (Ch4.4).
Sources
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.