Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
Speculative decoding is a lossless inference acceleration technique that uses a small draft model to propose multiple tokens, which the large model verifies in a single forward pass, achieving 2-3x speedup without altering output quality. The key insight is that GPU memory bandwidth, not compute, is the bottleneck, so idle compute can be spent on checking multiple tokens at once. In production with batching, gains typically range from 1.2x to 2x, and the technique works best when the GPU is memory-bound—common with long contexts, mixture-of-experts models, and quantized weights.
Key points
Speculative decoding uses a draft model to generate multiple token guesses, then the large model verifies them in one forward pass; rejection sampling ensures the output distribution is identical to the large model alone.
The technique was independently published by Google (November 2022) and DeepMind (February 2023), and is now deployed in Google products like AI Overviews.
Variants include Medusa (extra prediction heads), Eagle (feature prediction from internal layers), multi-token prediction (DeepSeek V3), and Engram (prompt-based repetition for code editing).
Eagle 3 (March 2025) reported up to 6.5x speedup at batch size 1, but Berkeley's production evaluation found gains of 1.2-2x at realistic batch sizes due to compute saturation.
A Red Hat study on the MoE model GPTO OS (117B parameters) showed 27% more output tokens per second at 200 concurrent requests, because MoE models are memory-bound.
The effectiveness depends on acceptance rate, draft cost, and speculation length; acceptance rate is domain-specific—higher on code and math, lower on open-ended writing.
Berkeley's oracle experiment showed that knowing when to stop guessing could unlock over half the remaining speedup, suggesting a scheduler is more valuable than a better draft model.
Practical advice: use Eagle 3 if available, use pre-trained prediction heads, enable Engram for code edits, and measure acceptance length on your own workload.
Tools mentioned
Techniques
- Speculative decoding
- Rejection sampling
- Multi-token prediction
- Tree-based verification
- Feature prediction
- Engram (prompt-based speculation)
- Continuous batching
- KV cache
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
A language model finishes one word to start the next one. It reads its entire brain again. Every weight, every parameter off the memory chips across the bus into the processor once per
word. A 70 billion parameter model at 16 bits is 140 GB of weights and all of it moves to produce a single token. An H200, one of the fastest cards you can rent, moves about 4.8 tab a second. do
the division. That is 29 milliseconds per token. 34 a second. You're ceiling. Not because the math is hard, because the wires are narrow. The silicon can run hundreds of trillions of operations
a second. For those 29 milliseconds, it runs almost none. Which is the question this whole video answers. If the compute is sitting idle, can you spend it to buy the time back? You can. People have
since the end of 2022. The trick is called speculative decoding. Same model, same text, sampled from the same distribution as before, two to three times faster. By the end of this, you
will know how it works, what it costs, and the one number that decides whether it is worth turning on for you. Start with the thing it is replacing because most explanations skip it. Auto
reggressive decoding is the loop every transformer runs to produce text. You hand it a prompt, it processes the whole prompt in one shot, and that part is fast because a thousand tokens go
through the matrix multiplies together. Engineers call that prefill. Then it produces the first word of the answer. And to produce the second word, it needs the first one because the first one is
part of its input. Now that dependency is the entire problem. You cannot compute token 2 before token one exists. So the loop is serial by construction. The paper that started this field opens
with the flattest possible statement of it. Decoding K tokens takes K serial runs of the model. Now engineers did fix part of this. The KV cache stores the attention keys and values for every
token you have already processed. So you are not rerunning the prompt every step. It is a huge win and it is also the source of a misunderstanding. The cache saves you from redoing attention. It
does nothing about the weights. The weights are the bulk. Every layer, every projection matrix, every expert has to arrive at the compute units before it can multiply anything. And it arrives at
whatever speed the memory bus allows. Google research put the imbalance in one sentence. Modern accelerators can perform hundreds of operations for every bite read from memory. Transformer
decoding performs a few. Picture it physically. You own a factory floor with 200 machines and a single delivery van. The van is the bottleneck and buying more machines does not make the van
faster. The industry does have a standard answer and you already know it if you have run a serving stack. Batch the requests. Load the weights once. run 64 users through them together and
suddenly one delivery feeds 64 jobs. Throughput goes up enormously, which is great for your cloud bill and does nothing for the person waiting. Batching improves tokens per second across all
users. It does not improve tokens per second for one user. And that gap is exactly where the last two years of product design landed. Coding agents, long reasoning traces, a model writing
4,000 tokens while one developer watches the cursor. Your agent is not being batched with 63 strangers to hide the latency. It is one stream and the ceiling is the one we just calculated.
So back to the idle compute. Here is the observation the whole technique rests on. And it is a little bit funny once you see it. Running the big model on one token cost you a full weight read.
Running the big model on five tokens cost you the same full weight read. The five positions go through the same matrices in parallel. Five times the arithmetic on hardware that had
arithmetic to spare. moving exactly the same bytes. Checking five words is nearly free. Producing five words is five times the cost. So if you could get five words from somewhere cheap, the big
model could check them all in one go. That somewhere cheap is a second much smaller model, the draft model. Same tokenizer, same vocabulary, a fraction of the parameters. The draft runs its
own serial loop, five quick steps, and hands over a guess. The cat sat on the mat. And then the big model does one forward pass over all five positions at once. And for each position, it reports
what it would have produced there. You walk the list left to right. The cat sat on. Four matches. The fifth is wrong. Stop there. You keep the four accepted words. And here is the part people miss.
That same forward pass already computed the correct fifth word because it computed a prediction at every position. So you take that one two five words out of one big model pass and the draft cost
you a rounding error. Worst case the very first guess is wrong. You keep one token which is exactly what plain decoding would have given you plus the draft overhead. You are barely behind.
Best case the whole guess lands and you jump several words in one step. On average you land somewhere in between and that average has a name we will come back to. Now if you were paying
attention you should be suspicious. A small model just decided part of your output. Is the text still what the big model would have written? Yes, provably. Bit forbit at the distribution level.
And this is the thing that separates speculative decoding from every other speed trick. Quantization changes your outputs. Pruning changes your outputs. Distillation changes your model.
Speculative decoding changes how long you wait, not what you get. The mechanism is a modified rejection sampling scheme. For each drafted token, the draft gives a probability, call it
Q, and the big model gives its own probability, call it P. If the big model likes that token at least as much as the draft did, you accept it outright. If it likes it less, you accept it with
probability P over Q. And otherwise, you reject. On a rejection, you do not just fall back to the big model. You sample from the difference between the two distributions normalized, which corrects
for the bias the draft introduced. Work through the algebra and the output distribution comes out identical to sampling from the big model directly. The original paper says it plainly,
identical outputs. That guarantee is why this shipped everywhere instead of staying a research curiosity. You are not trading quality for speed. There is no quality dial in this system. The
history is short and slightly unusual. Two teams landed on it independently about 2 months apart at the end of 2022 and the start of 2023. Yaniv Leviathan and colleagues at Google published first
in November showing two to three times acceleration on a large encoder decoder model with identical outputs. Then a deep mind team published speculative sampling in February benchmarking
chinchilla 70 billion parameters at two to two and a half times faster in a distributed setup. Google says the technique is now deployed across its own products including AI overviews inside
Google search. When two labs invent the same thing in the same quarter, the idea was ready. Here is the arithmetic that governs how much you actually get. Your speed up is a function of three things
and only three. How often the draft is right, how much the draft cost relative to the big model, and how many tokens you let it guess before you check. The first one is the acceptance rate, and it
is the number that matters most. Everything else in this field is a strategy for pushing it up. There is attention baked into it. A bigger draft model guesses better, which raises
acceptance, and it also costs more per guess, which eats the winnings. The Berkeley measurement puts real numbers on that trade. Against a 70 billion parameter target, a good draft takes
about 12 1.5% of the target forward pass. Speculation is cheap. Against an 8 billion parameter target, that ratio rises to 37 12%. The tax roughly tripled and the winnings did not, which is a
rule you can carry away right now. The bigger the model you are serving, the better speculative decoding works because the draft is a smaller slice of it. Except running a second model is
annoying. Two checkpoints, two sets of weights in memory, and you have to find a small model that shares a tokenazer with your big one. So, the field spent 3 years deleting the draft model. And that
is what every acronym you have seen in a serving config is about. The first move was Medusa in early 2024. Instead of a separate model, bolt extra prediction heads onto the one you already have.
Head one predicts the next token, which it already did. Head two predicts the token after that. Head three, the one after that. One backbone, several guesses, no second checkpoint. Medusa
reported over 2.2 times faster with the backbone frozen and 2.3 to 3.6 times when they fine-tuned the whole thing together. It also introduced a trick that stuck. Rather than one linear
guess, propose a tree of candidate continuations and verify all the branches in a single pass with a masked attention pattern. Then came Eagle and Eagle is the one you will actually meet
in production. So it is worth understanding what it changed. Instead of predicting the next token from scratch, the draft head predicts the big model's own internal feature vector for
the next position and reuses it. It is a smarter question to ask. The features carry more information than a token ID. So, a tiny head that reads them guesses far better than a tiny model guessing
alone. Eagle 3, published in March 2025, dropped feature prediction for direct token prediction and fuse features from several layers of the target instead of only the top one. Their headline is a
speed up ratio up to 6.5 times, about 1.4 times better than the previous version. Hold on to that number. We are going to take it apart later. The cost of all this is smaller than people
expect. Berkeley measured eagle style heads adding under 10% to both static weights and per token cache. A separate draft model is the expensive option pairing a small model with an 8 billion
target raised per token memory by 1.77 times. And this is still moving. In May 2026, the Eagle VLLM and Torch spec teams shipped Eagle 3.1 with Nvidia tuned for long context reporting up to
twice the accepted length of Eagle 3. A few months earlier, the same project shipped Parallel Eagle, which generates all of its draft tokens in one forward pass instead of one at a time on a
coding benchmark that took the average accepted run from 3.03 tokens to 3.94, 30% more per step for free. There is a third family, and it may be the most elegant. Skip the draft entirely and
train the model to predict two tokens at once during pre-training. Deepseek shipped exactly that in their December 2024 V3 report. Multi-token prediction, one extra head, and a measurement most
labs would not publish. The second predicted token was accepted between 85 and 90% of the time across generation topics, which took decoding to 1.8 times the tokens per second. 85% acceptance is
enormous. It means the model's own guess about its own next next word is right nearly nine times in 10 and the whole verification dance pays off nearly every time. That approach spread. Open models
increasingly ship prediction heads inside the weights themselves. And in June 2026, Deepseek open- source DSpark, a framework that runs per user generation 60 to 85% faster than their
own single head baseline. Now the last family and it is the one that rarely makes the marketing engram speculation. No draft model, no extra heads, no training whatsoever. The idea is almost
stupid. Look at what the model just wrote. Look at the prompt, find a matching phrase, and guess that the continuation repeats. On most workloads, that is a bad guess, and it
underperforms everything else. On one workload, it beats every trained method on the market. Code editing. When you paste a file and ask for a change, most of the output is the input. Berkeley
found that once about 60% of the output phrases already appear in the prompt, engram wins across every batch size they tried by as much as 100%. Its behavior is completely different from the trained
methods, too. On that same code editing set, against the 70 billion model, the Eagle draft accepted a steady 2.7 to 7.4 tokens per step. Engram swung from 1.1 to 15. Steady versus bursty, which is
the right shape. If you serve code edits and document rewrites, take the bursty one. It costs you nothing and it is already in your engine. So, we have a lossless trick with publish speedups
from two times to 6 and a half, which raises the obvious question, why is it not switched on by default everywhere? At the end of 2025, a group at Berkeley Sky Computing Lab with ion stoka among
the authors decided to check and their writeup is titled speculative decoding performance or illusion. Their complaint about the prior literature is specific and reading it back hard to argue with.
Earlier evaluations use research prototypes and they tested at a batch size of one. So they ran five variants across four models and six workloads on production VLLLM with continuous
batching and everything else a real deployment has at batch sizes from 1 to 128 and up to 512 when they profiled it. At a batch size of 1, Eagle reached up to 1.96 times on the 70 billion model.
Respectable and already well under the 6 1/2 you see in a paper title. at a batch size of 128 on a smaller 8 billion model doing grade school maths it was 1.21 two one times different model and the shape
held on every curve they plotted they all trend down still a win just not the same conversation and their explanation for the decay is the same physics we opened with running backwards at batch
one the GPU had idle compute so speculation was free at batch 128 the machine is already compute saturated now every drafted token you verify is competing for arithmetic that somebody
is using they profiled where the time goes and the answer is blunt Verification by the big model takes between 42 and 95% of execution time. The rejection sampling maths, the
elegant part, the part with the proof, costs under 1.7%. The expensive thing is not the algorithm at all. It is the forward pass which reframes what a rejected token actually is. Think of it
as a slice of your biggest model burned on a word you threw away. And you can push that too far. In a separate test on a different serving engine, they tried wide speculation trees. 21 candidates
verified at once. At batch one, the tree beat the simple chain 1.65 times up to 1.85. At batch 64, the same configuration dropped below one time speed up on every
workload. Speculative decoding configured badly made the server slower than doing nothing. The decay also scales with model size in the direction you would not guess. On conversational
traffic, going from batch 1 to 32 cost the 8 billion model 4.3% of its speed up and the 70 billion model 14%. So the practical lever exists and it is one line. Serving engines let you switch
speculation off above a batch size threshold. So a server that speculates for your chat traffic can stop the moment a batch job floods it, which would be a tidy ending. Great for your
chat app. Switch it off for your batch jobs. Done. Except four months later, another team measured the opposite. In April 2026, engineers publishing through Red Hat benchmarked speculative decoding
on OpenAI's largest open model, GPTO OS, a mixture of experts with 117 billion parameters in total, running on an H200. They took it to 200 concurrent requests, not 64, 200. And the gains did not
collapse. 27% more output tokens per second at peak on conversational traffic. 20% on a software engineering benchmark which came with a 19% cut in cost per million tokens. Their
conclusion is that the improvements stay consistent even at high concurrency which is a polite way of contradicting the best known description of this technique. So which measurement is
wrong? Neither. And the reason they disagree is the most useful thing in this video. Batch size was never the real variable. Batch size was a proxy for the real variable which is whether
your machine is waiting on memory or waiting on arithmetic. Speculative decoding spends compute to buy back memory time. It pays whenever you have spare compute. It stops paying the
moment you do not. Batch size is one thing that consumes compute. It is not the only thing. And three changes in how models are built have moved that line a long way. First, context length. A team
showed back in 2024 that as sequences get long, the KV cache itself becomes the dominant thing being read. The cache grows with batch size and with sequence length together. So a 100,000 token
conversation at batch 64 is not computebound. It is drowning in cash reads. On an 8 billion model, they reported up to two and a half times speed up at batch sizes from 32 to 256.
Exactly the regime conventional wisdom called hopeless. Second mixture of experts that 117 billion parameter model does not use 117 billion parameters per token. It activates about 5 billion of
them per token. So the arithmetic per token collapsed while the memory footprint stayed large. That is the definition of memory bound and it is why the high concurrency result came from an
expert model rather than a dense one. Third 4-bit weights. Quantizing to four bits cuts the bytes you move to a quarter of what 16 bit needs and leaves the operation count exactly where it
was. Same trade, same direction. Put those three together. A modern serving stack, sparse and quantized and running long conversations, sits in the memory bound regime far longer than a dense
16-bit model from 2023 ever did. Which is the rule worth writing down. Ask what your GPU is waiting for. If the answer is memory, speculate. If the answer is arithmetic, do not. So, concretely, what
do you turn on? In VLLM, it is one config block. You name a method, you name a draft, and you say how many tokens to guess. If a draft exists for your model, use Eagle 3. Berkeley called
it the best all-round choice and drafts for it now ship alongside open models on the major serving stacks. If your model was prepained with prediction heads, use those. They are already inside the
weights you downloaded. Check whether your engine is actually using them because a head nobody enables is a head nobody benefits from. If your traffic is code edits, document rewrites, or
retrieval where the answer quotes the source, turn on engram first. It needs no training at all and on that shape of work, it beats the methods that were trained. Then measure the right thing,
not the vendor speed up number which was measured on their workload. Measure your accepted length, the average tokens you get per big model pass. Because acceptance is domain specific in a way
the marketing tends to skip. One production draft averages 3.16 accepted tokens on code and 3.12 on maths and drops toward two on open-ended writing. Same model, same hardware, same config.
Your chat product and your coding product will not get the same speed up. And if you only benchmark one, you will be surprised by the other. So here's where I come down. Turn it on. For
anything a human or an agent is waiting on, speculative decoding is the highest return change available in inference, and it is the only one that cost you no quality at all. The receipts are on the
table. Two independent labs prove the output distribution is unchanged. Berkeley found it helping at every batch size they tested. Red Hat measured a 19% cut in cost per million tokens on a
coding workload with the gains holding to 200 concurrent requests. The concession is real and I would rather say it than have you find it. The 6 and 1/2 times headline is a batch of one
number. In production, plan for something between 1.2 and two and configure conservatively because a wide tree at batch 64 is a way to pay for a slower server. And if you are running
maximum throughput batch jobs on a dense model where nothing is waiting and every unit of arithmetic is already spoken for, leave it off. That is the workload this does not serve. What I keep turning
over is the last finding in that Berkeley paper. They built an oracle that knows in advance exactly how many tokens will be accepted and letting it pick between two speculation methods at
every position hit 4.9 times on code editing against about 2.2 for the best single method oracle. more than half the available speed up is still sitting there and a bigger draft model is not
what unlocks it. What unlocks it is knowing when to stop guessing. So what would you rather have? A draft that guesses better or a system that knows when its guess is not worth checking?
One of those is a research program. The other might just be a scheduler.