Appearance
Speculative Decoding and Medusa/EAGLE
1. Definition: The Most Elegant Acceleration for the Decode Phase
Speculative Decoding is an LLM inference acceleration technique proposed by Leviathan et al. in 2022. The core idea in one sentence: let a small, fast draft model sketch k tokens, and let the large target model verify all of them in a single forward pass — accepting the correct ones and rolling back the wrong ones.
Why is this both "lossless" and "effective"? Two mathematical facts:
- Losslessness guarantee: by rewriting the verification procedure with rejection sampling, the final sampling distribution is exactly equivalent to sampling directly from the target model — in theory, no quality is lost at all.
- Why it works: the LLM decode phase is memory-bound (see The GPU Memory Hierarchy and the Bandwidth Wall and The Roofline Model and Compute Analysis) — generating each token requires reading the entire model weights + KV cache from HBM once, while performing very few FLOPs. Verifying k tokens costs k× the FLOPs of a single token, but the memory access barely changes (the weights are read once). So "verify k at once" is far cheaper than "generate k one by one."
Traditional autoregression (k=4):
target forward × 4 each pass: read all weights + very few FLOPs = 4× the memory-bandwidth time
Speculative decoding (k=4):
draft forward × 4 the draft is tiny, nearly free
target forward × 1 verify 4 tokens at once, roughly the time of 1 pass
acceptance rate 80% → average 4×0.8 + 0.2×1 = 3.4 tokens per step, a 3.4× speedupThe key fact
Speculative decoding is essentially "trading the draft's sense of direction for the target's throughput" — the more accurately the draft guesses and the more tokens get accepted, the closer the speedup gets to k×. When the draft guesses wrong, some draft compute is wasted, but the target never computes anything wrong (rejection sampling guarantees losslessness).
2. The Mathematical Foundation: Why It Is Lossless
Understanding the losslessness requires looking at the rejection sampling implementation. Let the target distribution be p(x) and the draft distribution be q(x):
1. The draft samples k tokens from q: x_1, x_2, ..., x_k
2. For each x_i:
- compute the acceptance probability r = min(1, p(x_i) / q(x_i))
- accept x_i with probability r
- if rejected, resample from the normalized distribution max(0, p(x) - q(x))
3. After all x_i are processed, the target samples the next tokenKey theorem (Leviathan 2023): the token distribution produced by the above procedure is exactly the same as sampling directly from p(x). Intuition: tokens the draft gets right are accepted outright (r=1); tokens it gets "half right" are accepted proportionally (r<1); tokens it gets wrong are always rejected and resampled "from the probability mass where p exceeds q" — which exactly makes up the difference.
Numerical intuition
If q ≈ p (draft and target highly aligned): the acceptance rate approaches 100%, giving an almost lossless k× speedup. If q ≠ p: the acceptance rate drops, but each step still generates > 1 token, so the speedup remains > 1.
3. Three Generations of Methods
Generation 1: Classic Speculative Decoding (Leviathan 2023)
draft model (small, e.g. 1B) → generates k tokens serially: x_1..x_k
target model (large, e.g. 70B) → verifies x_1..x_k in one forward pass- Pros: simple idea, theoretically lossless.
- Cons: requires a separately trained / chosen draft model (a "little sibling" from the same family), complicating deployment; the draft also spends time generating k tokens serially.
Typical setup: Llama-2-70B + Llama-2-7B as the draft, 2–3× speedup.
Generation 2: Medusa (Cai 2024, ICML) — Dropping the Draft Model
Medusa's insight: why train a separate draft from scratch? Let the target model grow extra prediction heads of its own — attach k "Medusa heads" on top of the target's last layer, where each head independently predicts the next token, the one after that, ..., up to the k-th.
hidden state h of the target model's last layer
│
├── original LM head → predict token t+1
├── Medusa head 1 → predict token t+2
├── Medusa head 2 → predict token t+3
└── ...
└── Medusa head k → predict token t+k+1
One forward pass → candidate tokens at multiple positions, all at once
Build a candidate tree → tree attention verification → accept the longest correct prefix- Pros: no separate draft model; only lightweight fine-tuning of the Medusa heads (a few hours).
- Cons: the heads predict independently — the second head cannot condition on "the first one being wrong" — so long-prefix acceptance rates are limited.
Performance: 2.2–2.8× speedup (as reported, Llama-2-70B).
Generation 3: EAGLE-1/2/3 (Li 2024–2025) — Feature-Level Autoregression
EAGLE's core innovation: draft at the feature level instead of the token level. The draft model is a lightweight transformer that takes the target's hidden states (not tokens) as input, predicts the next hidden state, and decodes it into a token with the target's LM head.
the target model provides a hidden state sequence h_1, h_2, ..., h_t
│
▼
EAGLE draft model (1-layer transformer + projection)
│
▼
predicts h_{t+1}, h_{t+2}, ..., h_{t+k} (autoregressive, each step conditioned on the previous)
│
▼
decode h_{t+i} into a token with the target's LM head
│
▼
tree attention verifies k tokens in one passWhy is EAGLE stronger than Medusa? Because hidden states carry more information than tokens — the draft model predicts in a "high-dimensional semantic space," which is more accurate than predicting in "discrete token space."
| Version | Date | Key improvement | Speedup |
|---|---|---|---|
| EAGLE-1 | 2024.03 | Feature-level autoregression + tree attention | 3× |
| EAGLE-2 | 2024.06 | Dynamic draft tree (adapts to context) | 3.5–4× |
| EAGLE-3 | 2025.03 | "Future token" supervision during training, deeply aligning draft with target | 3–6.5× |
EAGLE-3 in practice
On SGLang + H100:
- Llama-3-8B: 373 tokens/s vs a 158 tokens/s baseline (2.4×)
- Llama-3-70B: 80 tokens/s vs a 30 tokens/s baseline (2.7×)
- MT-Bench scores barely change (lossless)
These numbers make EAGLE-3 the de facto first choice for speculative decoding in 2025.
Other Important Variants
| Method | Idea | Applicability |
|---|---|---|
| Lookahead Decoding (2024) | n-gram-style autoregression, no training needed | Universal, zero cost, 1.5–2× speedup |
| REST / BiLD | Retrieval-based draft, retrieving candidates from a corpus | Long context + document-dense tasks |
| Medusa-Multihead | Medusa + joint multi-head training | Improves accuracy over Medusa |
| DeepSeek-V3 MTP | Train "multi-token prediction" heads during pretraining; the module doubles as a draft | Native to DeepSeek-V3 / R1 |
DeepSeek-V3's MTP (Multi-Token Prediction) is an interesting paradigm: the model learns to predict multiple tokens at once during training, internalizing the draft capability — at inference time the MTP module serves as the draft with no extra training. See Classic Papers in Depth and Frontier Advances.
4. Why the Long-Prefix Acceptance Rate Is the Key Metric
The speculative decoding speedup ≈ 1 + (1 - α^(k+1)) / (1 - α) × β, where α is the single-step acceptance rate and β is the draft's relative speed. The intuitive version:
Acceptance rate 50%, k=4 → average 1.9 tokens/step → 1.9× speedup
Acceptance rate 70%, k=4 → average 3.0 tokens/step → 3.0× speedup
Acceptance rate 80%, k=4 → average 3.4 tokens/step → 3.4× speedup
Acceptance rate 90%, k=4 → average 3.8 tokens/step → 3.8× speedupThe acceptance rate is the core metric. EAGLE-3 pushes it to 85–92% on Llama-3 — the root cause of its 3–6.5× speedups.
5. Code Examples: Enabling It on vLLM / SGLang
Enabling EAGLE on vLLM
bash
# 1. Download the EAGLE-3 draft checkpoint
huggingface-cli download yuhuili/EAGLE3-LLaMA3-Instruct-8B
# 2. Launch vLLM (0.6+) with EAGLE
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Meta-Llama-3-8B-Instruct \
--speculative_model yuhuili/EAGLE3-LLaMA3-Instruct-8B \
--use_v2_block_manager \
--speculative_draft_tensor_parallel_size 1 \
--num_speculative_tokens 4 \
--ngram_prompt_lookup_max 0Enabling EAGLE-3 on SGLang
python
from sglang import Runtime, gen
runtime = Runtime(
model="meta-llama/Meta-Llama-3-8B-Instruct",
speculative_model="yuhuili/EAGLE3-LLaMA3-Instruct-8B",
speculative_num_draft_tokens=4,
speculative_eagle_topk=8,
mem_fraction_static=0.85,
)
runtime.start()
@runtime.function
def generate(text):
out = gen(text, max_tokens=512, temperature=0.7)
return out
print(generate("Explain the core idea of PagedAttention"))vLLM Lookahead (Training-Free Version)
bash
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Meta-Llama-3-8B-Instruct \
--speculative_model "[ngram]" \
--num_speculative_tokens 5 \
--ngram_prompt_lookup_max 4 \
--ngram_prompt_lookup_min 26. Performance Data: Measured Reference
Here is a set of baselines on H100 (see Inference Benchmarking in Practice):
| Method | Model | Baseline tokens/s | Accelerated tokens/s | Speedup |
|---|---|---|---|---|
| EAGLE-3 | Llama-3-8B | 158 | 373 | 2.4× |
| EAGLE-3 | Llama-3-70B (2×H100) | 30 | 80 | 2.7× |
| EAGLE-2 | Qwen-2-72B | 25 | 75 | 3.0× |
| Medusa | Llama-2-70B | 22 | 60 | 2.7× |
| Lookahead | Llama-3-8B | 158 | 280 | 1.8× |
| DeepSeek-V3 MTP | DeepSeek-V3 | 35 | 75 | 2.1× |
The boundaries of these numbers
All speedups are affected by the following:
- Batch size: the speedup peaks at batch=1; at large batches (>32) decode is already close to compute-bound and the speedup drops sharply to 1.0–1.3×.
- Task type: "highly repetitive" tasks like code / math see 80–92% acceptance rates; open-ended creative writing sees 60–75%.
- Draft-target alignment: cross-family pairs (e.g., a Llama draft + Qwen target) perform poorly.
7. Limitations and Boundaries
- Fails at large batches: decode is memory-bound at batch=1, and speculative decoding saves bandwidth; at large batches it is already compute-bound, and verifying k tokens genuinely adds FLOPs, so the speedup vanishes. Gains are limited in high-concurrency production serving — this is why vLLM keeps speculative decoding off by default.
- Requires a draft model or extra training: the classic method needs a trained draft; Medusa/EAGLE need head / small-transformer training; DeepSeek-V3 MTP must be planned into pretraining.
- Acceptance rate drops for long context + diverse tasks: open-ended creation and long-form reasoning can fall to 50–60% acceptance, dragging the speedup down to ~1.5×.
- Complex implementation: tree attention and KV cache management mean major engine changes and high bug risk. It only stabilized in vLLM 0.5+, and TRT-LLM added support in 0.10+.
- Lossless ≠ no latency variance: the number of tokens generated per step varies a lot, increasing variance in first-token and total latency; latency-sensitive applications must evaluate this.
- Does not apply to the prefill phase: speculative decoding only accelerates decode; prefill time for long prompts is unchanged. See Latency, Throughput, and Concurrency.
8. Comparison with Similar Methods
| Method | Training required | Speedup | Deployment difficulty | Applicability |
|---|---|---|---|---|
| Classic spec decoding | Train a standalone draft | 2–3× | High | Large model + small model on two cards |
| Medusa | Train Medusa heads | 2.2–2.8× | Medium | Moderate-compute scenarios |
| EAGLE-1/2/3 | Train an EAGLE draft | 3–6.5× | Medium | Current recommendation |
| Lookahead | None | 1.5–2× | Low | Quick experiments |
| DeepSeek MTP | Integrated at pretraining | 2–3× | Extreme (requires retraining) | When training a new model |
| Speculative Decoding + Quantization | - | Additive | Medium | INT4 + spec stack nicely |
9. When to Use It, and When Not To
Use it when:
- Low-concurrency scenarios (batch ≤ 8) where decode dominates latency
- Tasks with high acceptance rates (code completion, dialogue continuation, math reasoning)
- You have a suitable draft model or are willing to train EAGLE
Don't use it when:
- High-concurrency production serving (batch > 32, already compute-bound)
- Open-ended creation (low acceptance rate, degraded speedup)
- Strict latency SLAs (high variance)
- Scenarios where prefill dominates decode (e.g., RAG with long contexts)
10. Where to Go Next
- Concept pages: Latency, Throughput, and Concurrency, The GPU Memory Hierarchy and the Bandwidth Wall, The Roofline Model and Compute Analysis, Batching and Request Scheduling, Model Quantization Fundamentals
- Case-study pages: vLLM, TensorRT-LLM, llama.cpp
- Papers: Classic Papers in Depth, Reading Paths, Frontier Advances
- Practice pages: Tuning and Performance Optimization, Inference Benchmarking in Practice, Inference Engine Comparison, Common Pitfalls and Anti-Patterns
References
- Leviathan et al. Fast Inference from Transformers via Speculative Decoding (ICML 2023) — classic spec decoding
- Chen et al. Accelerating Large Language Model Decoding with Speculative Sampling (2023) — concurrent independent work
- Cai et al. Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads (ICML 2024) — Medusa
- Li et al. EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty (ICML 2024) — EAGLE-1
- Li et al. EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees (2024) — EAGLE-2
- Li et al. EAGLE-3: Scaling up Inference Acceleration towards 1/2 Throughput (2025) — EAGLE-3
- Fu et al. Lookahead Decoding (2024) — training-free draft
- DeepSeek-AI. DeepSeek-V3 Technical Report (2024.12) — the MTP paradigm
- EAGLE GitHub — official implementation