Model Serving · Software component

Speculative Decoder

Software componentModel ServingModelsarc:SpeculativeDecoder

An inference-engine component that drafts candidate tokens with a small fast model and verifies them in parallel with the large target model, keeping only tokens the target accepts.

Responsibility. Reduces generation latency without changing the target model's output quality.

Also known as: Draft-Target Decoding

When to choose. Choose when latency is critical or sequences are long (Ref7.05).

deployed onis evaluated bydeployed onis monitored byemits telemetry todeployed onhostshostshostsis configured byis monitored byLLM Inference Service: deployed onLLM Inference ServiceEvaluation Harness: is evaluated byEvaluation HarnessInference Server: deployed onInference ServerAlert Manager: is monitored byAlert ManagerMetrics Collector: emits telemetry toMetrics CollectorLLM Generation Backend: deployed onLLM Generation BackendSmall Language Model Tier: hostsSmall Language Model TierLarge Language Model Tier: hostsLarge Language Model TierDraft Token Proposer: hostsDraft Token ProposerSpeculative Decoding Configuration: is configured bySpeculative Decoding Con…Draft Length Controller: is monitored byDraft Length Controller
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

hosts structural

is configured by structural

emits telemetry to dynamic

is evaluated by assurance

is monitored by assurance

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Speculative decodingModified rejection samplingParallel verification
Technologies
TensorRT-LLMTensorRT-LLM SpeculativeConfig
Quality attributes
Performance efficiency (ISO/IEC 25010)

Sources

  1. Ch3.4: T. Nguyen, "Tuning Model Parameters for Production Performance," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.4. ISBN: 9798244538229.
  2. Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
  3. Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
  4. Ch8.3: T. Nguyen, "Token Economics and Architecture," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.3. ISBN: 9798244538229.
  5. Ref4.01: NVIDIA, "TensorRT-LLM," GitHub repository. Accessed: Sep. 27, 2026. [Online]. Available: https://github.com/NVIDIA/TensorRT-LLM
  6. Ref4.03: M. Zhang, J. Wyman, I. M. Bhosale, and W. Tan, "Scaling LLMs with NVIDIA Triton and NVIDIA TensorRT-LLM using Kubernetes," NVIDIA Technical Blog, Oct. 22, 2024. [Online]. Available: https://developer.nvidia.com/blog/scaling-llms-with-nvidia-triton-and-nvidia-tensorrt-llm-using-kubernetes/
  7. Ref7.05: S. Verma and N. Vaidya, "Mastering LLM Techniques: Inference Optimization," NVIDIA Technical Blog, Nov. 17, 2023. [Online]. Available: https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/
  8. Ref7.15: "Advanced Agentic AI Optimization Techniques," unpublished reference note (15-Advanced-Agentic-Optimization.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
  9. Ref7.18: "Chapter 7 Summary: NVIDIA Platform Implementation," unpublished reference note (18-Chapter-7-Summary.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note