Skip to content

Speculative Decoding and Medusa/EAGLE

At a glance A small draft model predicts k tokens; the large model verifies them all in a single forward pass and accepts every correct one — the most elegant lossless 2–6× acceleration for the LLM decode phase. This article dissects three generations of speculative decoding, the mathematical losslessness guarantee, and the boundary conditions.

Speculative Decoding and Medusa/EAGLE ​

1. Definition: The Most Elegant Acceleration for the Decode Phase ​

Speculative Decoding is an LLM inference acceleration technique proposed by Leviathan et al. in 2022. The core idea in one sentence: let a small, fast draft model sketch k tokens, and let the large target model verify all of them in a single forward pass — accepting the correct ones and rolling back the wrong ones.

Why is this both "lossless" and "effective"? Two mathematical facts:

  1. Losslessness guarantee: by rewriting the verification procedure with rejection sampling, the final sampling distribution is exactly equivalent to sampling directly from the target model — in theory, no quality is lost at all.
  2. Why it works: the LLM decode phase is memory-bound (see The GPU Memory Hierarchy and the Bandwidth Wall and The Roofline Model and Compute Analysis) — generating each token requires reading the entire model weights + KV cache from HBM once, while performing very few FLOPs. Verifying k tokens costs k× the FLOPs of a single token, but the memory access barely changes (the weights are read once). So "verify k at once" is far cheaper than "generate k one by one."
Traditional autoregression (k=4):
  target forward × 4   each pass: read all weights + very few FLOPs = 4× the memory-bandwidth time

Speculative decoding (k=4):
  draft forward × 4    the draft is tiny, nearly free
  target forward × 1   verify 4 tokens at once, roughly the time of 1 pass
  acceptance rate 80% → average 4×0.8 + 0.2×1 = 3.4 tokens per step, a 3.4× speedup

The key fact

Speculative decoding is essentially "trading the draft's sense of direction for the target's throughput" — the more accurately the draft guesses and the more tokens get accepted, the closer the speedup gets to k×. When the draft guesses wrong, some draft compute is wasted, but the target never computes anything wrong (rejection sampling guarantees losslessness).

2. The Mathematical Foundation: Why It Is Lossless ​

Understanding the losslessness requires looking at the rejection sampling implementation. Let the target distribution be p(x) and the draft distribution be q(x):

1. The draft samples k tokens from q: x_1, x_2, ..., x_k
2. For each x_i:
   - compute the acceptance probability r = min(1, p(x_i) / q(x_i))
   - accept x_i with probability r
   - if rejected, resample from the normalized distribution max(0, p(x) - q(x))
3. After all x_i are processed, the target samples the next token

Key theorem (Leviathan 2023): the token distribution produced by the above procedure is exactly the same as sampling directly from p(x). Intuition: tokens the draft gets right are accepted outright (r=1); tokens it gets "half right" are accepted proportionally (r<1); tokens it gets wrong are always rejected and resampled "from the probability mass where p exceeds q" — which exactly makes up the difference.

Numerical intuition

If q ≈ p (draft and target highly aligned): the acceptance rate approaches 100%, giving an almost lossless k× speedup. If q ≠ p: the acceptance rate drops, but each step still generates > 1 token, so the speedup remains > 1.

3. Three Generations of Methods ​

Generation 1: Classic Speculative Decoding (Leviathan 2023) ​

draft model (small, e.g. 1B)  →  generates k tokens serially: x_1..x_k
target model (large, e.g. 70B) →  verifies x_1..x_k in one forward pass
  • Pros: simple idea, theoretically lossless.
  • Cons: requires a separately trained / chosen draft model (a "little sibling" from the same family), complicating deployment; the draft also spends time generating k tokens serially.

Typical setup: Llama-2-70B + Llama-2-7B as the draft, 2–3× speedup.

Generation 2: Medusa (Cai 2024, ICML) — Dropping the Draft Model ​

Medusa's insight: why train a separate draft from scratch? Let the target model grow extra prediction heads of its own — attach k "Medusa heads" on top of the target's last layer, where each head independently predicts the next token, the one after that, ..., up to the k-th.

hidden state h of the target model's last layer
       │
       ├── original LM head        → predict token t+1
       ├── Medusa head 1           → predict token t+2
       ├── Medusa head 2           → predict token t+3
       └── ...
       └── Medusa head k           → predict token t+k+1

One forward pass → candidate tokens at multiple positions, all at once
Build a candidate tree → tree attention verification → accept the longest correct prefix
  • Pros: no separate draft model; only lightweight fine-tuning of the Medusa heads (a few hours).
  • Cons: the heads predict independently — the second head cannot condition on "the first one being wrong" — so long-prefix acceptance rates are limited.

Performance: 2.2–2.8× speedup (as reported, Llama-2-70B).

Generation 3: EAGLE-1/2/3 (Li 2024–2025) — Feature-Level Autoregression ​

EAGLE's core innovation: draft at the feature level instead of the token level. The draft model is a lightweight transformer that takes the target's hidden states (not tokens) as input, predicts the next hidden state, and decodes it into a token with the target's LM head.

the target model provides a hidden state sequence h_1, h_2, ..., h_t
       │
       ▼
EAGLE draft model (1-layer transformer + projection)
       │
       ▼
predicts h_{t+1}, h_{t+2}, ..., h_{t+k} (autoregressive, each step conditioned on the previous)
       │
       ▼
decode h_{t+i} into a token with the target's LM head
       │
       ▼
tree attention verifies k tokens in one pass

Why is EAGLE stronger than Medusa? Because hidden states carry more information than tokens — the draft model predicts in a "high-dimensional semantic space," which is more accurate than predicting in "discrete token space."

VersionDateKey improvementSpeedup
EAGLE-12024.03Feature-level autoregression + tree attention3×
EAGLE-22024.06Dynamic draft tree (adapts to context)3.5–4×
EAGLE-32025.03"Future token" supervision during training, deeply aligning draft with target3–6.5×

EAGLE-3 in practice

On SGLang + H100:

  • Llama-3-8B: 373 tokens/s vs a 158 tokens/s baseline (2.4×)
  • Llama-3-70B: 80 tokens/s vs a 30 tokens/s baseline (2.7×)
  • MT-Bench scores barely change (lossless)

These numbers make EAGLE-3 the de facto first choice for speculative decoding in 2025.

Other Important Variants ​

MethodIdeaApplicability
Lookahead Decoding (2024)n-gram-style autoregression, no training neededUniversal, zero cost, 1.5–2× speedup
REST / BiLDRetrieval-based draft, retrieving candidates from a corpusLong context + document-dense tasks
Medusa-MultiheadMedusa + joint multi-head trainingImproves accuracy over Medusa
DeepSeek-V3 MTPTrain "multi-token prediction" heads during pretraining; the module doubles as a draftNative to DeepSeek-V3 / R1

DeepSeek-V3's MTP (Multi-Token Prediction) is an interesting paradigm: the model learns to predict multiple tokens at once during training, internalizing the draft capability — at inference time the MTP module serves as the draft with no extra training. See Classic Papers in Depth and Frontier Advances.

4. Why the Long-Prefix Acceptance Rate Is the Key Metric ​

The speculative decoding speedup ≈ 1 + (1 - α^(k+1)) / (1 - α) × β, where α is the single-step acceptance rate and β is the draft's relative speed. The intuitive version:

Acceptance rate 50%, k=4 → average 1.9 tokens/step → 1.9× speedup
Acceptance rate 70%, k=4 → average 3.0 tokens/step → 3.0× speedup
Acceptance rate 80%, k=4 → average 3.4 tokens/step → 3.4× speedup
Acceptance rate 90%, k=4 → average 3.8 tokens/step → 3.8× speedup

The acceptance rate is the core metric. EAGLE-3 pushes it to 85–92% on Llama-3 — the root cause of its 3–6.5× speedups.

5. Code Examples: Enabling It on vLLM / SGLang ​

Enabling EAGLE on vLLM ​

bash
# 1. Download the EAGLE-3 draft checkpoint
huggingface-cli download yuhuili/EAGLE3-LLaMA3-Instruct-8B

# 2. Launch vLLM (0.6+) with EAGLE
python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Meta-Llama-3-8B-Instruct \
    --speculative_model yuhuili/EAGLE3-LLaMA3-Instruct-8B \
    --use_v2_block_manager \
    --speculative_draft_tensor_parallel_size 1 \
    --num_speculative_tokens 4 \
    --ngram_prompt_lookup_max 0

Enabling EAGLE-3 on SGLang ​

python
from sglang import Runtime, gen

runtime = Runtime(
    model="meta-llama/Meta-Llama-3-8B-Instruct",
    speculative_model="yuhuili/EAGLE3-LLaMA3-Instruct-8B",
    speculative_num_draft_tokens=4,
    speculative_eagle_topk=8,
    mem_fraction_static=0.85,
)
runtime.start()

@runtime.function
def generate(text):
    out = gen(text, max_tokens=512, temperature=0.7)
    return out

print(generate("Explain the core idea of PagedAttention"))

vLLM Lookahead (Training-Free Version) ​

bash
python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Meta-Llama-3-8B-Instruct \
    --speculative_model "[ngram]" \
    --num_speculative_tokens 5 \
    --ngram_prompt_lookup_max 4 \
    --ngram_prompt_lookup_min 2

6. Performance Data: Measured Reference ​

Here is a set of baselines on H100 (see Inference Benchmarking in Practice):

MethodModelBaseline tokens/sAccelerated tokens/sSpeedup
EAGLE-3Llama-3-8B1583732.4×
EAGLE-3Llama-3-70B (2×H100)30802.7×
EAGLE-2Qwen-2-72B25753.0×
MedusaLlama-2-70B22602.7×
LookaheadLlama-3-8B1582801.8×
DeepSeek-V3 MTPDeepSeek-V335752.1×

The boundaries of these numbers

All speedups are affected by the following:

  1. Batch size: the speedup peaks at batch=1; at large batches (>32) decode is already close to compute-bound and the speedup drops sharply to 1.0–1.3×.
  2. Task type: "highly repetitive" tasks like code / math see 80–92% acceptance rates; open-ended creative writing sees 60–75%.
  3. Draft-target alignment: cross-family pairs (e.g., a Llama draft + Qwen target) perform poorly.

7. Limitations and Boundaries ​

  1. Fails at large batches: decode is memory-bound at batch=1, and speculative decoding saves bandwidth; at large batches it is already compute-bound, and verifying k tokens genuinely adds FLOPs, so the speedup vanishes. Gains are limited in high-concurrency production serving — this is why vLLM keeps speculative decoding off by default.
  2. Requires a draft model or extra training: the classic method needs a trained draft; Medusa/EAGLE need head / small-transformer training; DeepSeek-V3 MTP must be planned into pretraining.
  3. Acceptance rate drops for long context + diverse tasks: open-ended creation and long-form reasoning can fall to 50–60% acceptance, dragging the speedup down to ~1.5×.
  4. Complex implementation: tree attention and KV cache management mean major engine changes and high bug risk. It only stabilized in vLLM 0.5+, and TRT-LLM added support in 0.10+.
  5. Lossless ≠ no latency variance: the number of tokens generated per step varies a lot, increasing variance in first-token and total latency; latency-sensitive applications must evaluate this.
  6. Does not apply to the prefill phase: speculative decoding only accelerates decode; prefill time for long prompts is unchanged. See Latency, Throughput, and Concurrency.

8. Comparison with Similar Methods ​

MethodTraining requiredSpeedupDeployment difficultyApplicability
Classic spec decodingTrain a standalone draft2–3×HighLarge model + small model on two cards
MedusaTrain Medusa heads2.2–2.8×MediumModerate-compute scenarios
EAGLE-1/2/3Train an EAGLE draft3–6.5×MediumCurrent recommendation
LookaheadNone1.5–2×LowQuick experiments
DeepSeek MTPIntegrated at pretraining2–3×Extreme (requires retraining)When training a new model
Speculative Decoding + Quantization-AdditiveMediumINT4 + spec stack nicely

9. When to Use It, and When Not To ​

Use it when:

  • Low-concurrency scenarios (batch ≤ 8) where decode dominates latency
  • Tasks with high acceptance rates (code completion, dialogue continuation, math reasoning)
  • You have a suitable draft model or are willing to train EAGLE

Don't use it when:

  • High-concurrency production serving (batch > 32, already compute-bound)
  • Open-ended creation (low acceptance rate, degraded speedup)
  • Strict latency SLAs (high variance)
  • Scenarios where prefill dominates decode (e.g., RAG with long contexts)

10. Where to Go Next ​

References ​