Model Serving · Software component
Speculative Decoder
Software componentModel ServingModelsarc:SpeculativeDecoder
An inference-engine component that drafts candidate tokens with a small fast model and verifies them in parallel with the large target model, keeping only tokens the target accepts.
Responsibility. Reduces generation latency without changing the target model's output quality.
Also known as: Draft-Target Decoding
When to choose. Choose when latency is critical or sequences are long (Ref7.05).
Relationships
deployed on structural
hosts structural
is configured by structural
emits telemetry to dynamic
is evaluated by assurance
is monitored by assurance
Design guidance
- SHOULD be preferred over quantization when accuracy cannot be compromised but latency must improve.
- SHOULD be used for throughput-bound workloads (batch, offline summarization, synthetic data) rather than interactive streaming chat, where it raises time-to-first-token and makes emission bursty.
- SHOULD use a draft model roughly 5-10x smaller than the target and 4-8 token speculation windows.
- SHOULD estimate output predictability before deployment and validate output quality after.
- SHOULD apply speculative decoding particularly for long-sequence, high-throughput generation.
- SHOULD measure acceptance rate on the actual serving distribution and reduce K if acceptance < 0.5 or disable speculation if < 0.3.
- SHOULD account for batch-size reduction from draft memory when estimating net throughput.
- SHOULD auto-disable speculation and fall back to autoregressive decode when running acceptance drops below a threshold.
- SHOULD be applied for latency-critical (time-to-first-token) applications.
- MAY be adopted as a continuing cost-optimization step after initial caching, retrieval, output and routing optimizations.
Quantitative guidance
As stated by the sources; verify before use.
- Draft tokens accepted 60-80% of the time; 1.5-2.5x latency reduction with zero accuracy loss (Ch3.4).
- Expected tokens per target pass = 1 + window x acceptance (e.g., 1 + 5 x 0.6 = 4), minus ~10% draft overhead (Ch4.4).
- Production gains 2.5-3.5x for long-form moderately predictable generation, 4-5x for structured/template outputs (Ch4.4).
- Summarization example: 4.2 -> 11.2 req/s (2.67x), 62.7% acceptance, ROUGE-L 0.642 -> 0.641, +2 GB memory, 6.5 -> 2.4 GPU-hours/day (Ch4.4).
- Up to 3.6x throughput improvement; 3x throughput for Llama 3.3 70B (Ref4.06).
- Verification costs 1.2-1.5x a single decode step; alpha=0.8, K=4, c=0.1 gives 2.8-3.0x speedup; alpha=0.9 gives 3.2-3.5x; alpha=0.6 gives 1.8-2.2x; below alpha=0.3 underperforms baseline (Ch7.1A).
- Production optimizes K=4-6; K=8 at alpha=0.8 yields 25% more tokens at 2x cost (Ch7.1A).
- 10-30% of deployments achieve negative or negligible speedup with suboptimal configuration (Ch7.1A).
- Two-model coordination adds 5-15ms per cycle (Ch7.1A).
- 2-4x latency reduction while keeping full-model quality (Ref7.05).
- In a 70B worked example, speculative decoding cut latency from 40 to 25 ms/token (1.6x) and raised throughput from 120 to 150 tokens/s, at +2 GB per GPU and about $1/hour draft-model overhead (Ref7.15).
Classification
- Patterns
- Speculative decodingModified rejection samplingParallel verification
- Technologies
- TensorRT-LLMTensorRT-LLM SpeculativeConfig
- Quality attributes
- Performance efficiency (ISO/IEC 25010)
Sources
- Ch3.4: T. Nguyen, "Tuning Model Parameters for Production Performance," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.4. ISBN: 9798244538229.
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
- Ch8.3: T. Nguyen, "Token Economics and Architecture," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.3. ISBN: 9798244538229.
- Ref4.01: NVIDIA, "TensorRT-LLM," GitHub repository. Accessed: Sep. 27, 2026. [Online]. Available: https://github.com/NVIDIA/TensorRT-LLM
- Ref4.03: M. Zhang, J. Wyman, I. M. Bhosale, and W. Tan, "Scaling LLMs with NVIDIA Triton and NVIDIA TensorRT-LLM using Kubernetes," NVIDIA Technical Blog, Oct. 22, 2024. [Online]. Available: https://developer.nvidia.com/blog/scaling-llms-with-nvidia-triton-and-nvidia-tensorrt-llm-using-kubernetes/
- Ref7.05: S. Verma and N. Vaidya, "Mastering LLM Techniques: Inference Optimization," NVIDIA Technical Blog, Nov. 17, 2023. [Online]. Available: https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/
- Ref7.15: "Advanced Agentic AI Optimization Techniques," unpublished reference note (15-Advanced-Agentic-Optimization.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.18: "Chapter 7 Summary: NVIDIA Platform Implementation," unpublished reference note (18-Chapter-7-Summary.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note