Model Serving · Model asset

Self-Speculative Decoding Head

Model assetModel ServingModelsarc:SelfSpeculativeDecodingHead

Lightweight prediction layers attached to the target model that predict several future tokens simultaneously, enabling speculation without a separate draft model.

Responsibility. Provides speculative drafts with minimal memory and no inter-model coordination.

Also known as: Medusa heads, EAGLE heads

Variant of Draft Token Proposer abstract

When to choose. Choose when memory is constrained or batch size must be preserved (e.g., edge devices, high-batch serving).

deployed onalternative tospecializesEdge GPU Device: deployed onEdge GPU DeviceSpeculative Draft Model: alternative toSpeculative Draft ModelDraft Token Proposer: specializesDraft Token Proposer
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

alternative to variability

Quantitative guidance

As stated by the sources; verify before use.

Classification

Technologies
MedusaEAGLE

Sources

  1. Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.