Consider how human beings read familiar phrases. When you see: 'Actions speak louder than...' your brain does not decode every letter of the word 'words' from scratch; your subconscious anticipates the word before your eyes even reach it. Speculative decoding brings this exact predictive capability to artificial intelligence systems.
The Single-Token Autoregressive Tax
Standard language models are strictly autoregressive: to generate token $N+1$, the model must inspect tokens $1$ through $N$. When generating a 500-word response, an inference server must execute the full forward pass of a massive 70-billion-parameter neural network 500 consecutive times in lockstep.
Because each forward pass processes only a single token vector, the GPU's massive parallel arithmetic units are starved for work. Modern GPUs are designed to process thousands of mathematical calculations in parallel; forcing an H100 GPU to multiply a single vector against its weights is like using a 100-passenger bus to transport a single commuter across town.
[Standard Sequential Decoding: 5 Slow Forward Passes]
Step 1: Load 70B Weights ──► Emit 'The' (50ms)
Step 2: Load 70B Weights ──► Emit 'quick' (50ms)
Step 3: Load 70B Weights ──► Emit 'brown' (50ms)
Step 4: Load 70B Weights ──► Emit 'fox' (50ms)
Step 5: Load 70B Weights ──► Emit 'jumps' (50ms)
Total Latency: 250ms (GPU compute cores idling 90% of the time!)
[Speculative Decoding: Draft Fast, Verify in Parallel]
Step 1: Small 1B Draft Model guesses: ['The', 'quick', 'brown', 'fox', 'jumps'] (15ms)
Step 2: Large 70B Target Model checks all 5 tokens in ONE single forward pass (55ms)
Verification: ['The' ✓, 'quick' ✓, 'brown' ✓, 'fox' ✓, 'jumps' ✓]
Total Latency: 70ms (3.5x Faster with Zero Loss in Accuracy!)
The Mechanics of Speculative Decoding
Speculative decoding pairs a large, highly capable target model (e.g. 70B) with a tiny, ultra-fast draft model (e.g. 1B) or speculative draft heads:
- The Drafting Phase: The small draft model runs unhindered, rapidly predicting $K$ consecutive future tokens (for example, 5 tokens). Because the draft model is tiny, this takes mere milliseconds.
- The Parallel Verification Phase: The large 70B model accepts all 5 candidate tokens simultaneously in a single parallel forward pass. Because modern GPUs process batches of tokens at nearly the same latency as a single token, verifying 5 tokens takes virtually the same time as generating one!
- Rejection Sampling: The target model compares the draft model's probability distribution against its own. It accepts valid tokens in order and modifies the first divergence using mathematically guaranteed rejection sampling.
The Magic of Mathematical Equivalence
What makes speculative decoding so profound is that it involves zero degradation in output quality. Through exact rejection sampling, the output probability distribution matches the original 70B model down to the exact mathematical epsilon.
In common conversational English, code generation, and structured outputs, the draft acceptance rate often exceeds 75% to 85%. This translates directly into a 2x to 3x real-world speedup for users streaming answers in real time.
Engineering Takeaway
Speculative decoding exposes a foundational law of modern hardware: latency is dominated by memory trips, while parallelism is effectively free. When you can trade unused parallel compute to eliminate sequential memory cycles, you win every single time.