Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
Tencent's AngelSpec is a training framework for speculative decoding that lets you compare six drafting architectures under one config flag, but the repo has been largely ignored (219 stars) while one of its methods, Dflash2, went viral (70k downloads in a week). The paper's core finding is that no single drafter performs best across all workloads: multi-token prediction (MTP) wins on conversational data, block diffusion wins on code and math. AngelSpec's value is not in being the fastest—Dfly, its own method, achieves 1.98–2.40x speedup—but in being the only place you can train all six against your own model and discover which one your traffic needs.
Key points
AngelSpec is a training framework for six speculative decoding drafters, not a ready-to-use inference accelerator.
The repo has 219 stars and 19 forks, while Dflash2 alone had 70,000 downloads in under a week.
The paper shows that no single drafting architecture performs best across real-world workloads.
Dfly, Tencent's method, combines D-Flash's shared projection, D-Flare's per-layer refinement, and an autoregressive head.
Dcut treats verification as a shared compute pool across batch requests, yielding 15.7% more throughput on live traffic.
Training data choice (code/math vs conversational) had nearly as much impact on acceptance length as architecture changes.
Tools mentioned
Techniques
- speculative decoding
- multi-token prediction (MTP)
- block diffusion
- autoregressive drafting
- Dcut batch-aware verification scheduling
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
A file called Dflash2 went up in the middle of August and within a day it sat near the top of Hacker News. 99 points on the front page, 18 comments, about 70,000 downloads of a single model file
in under a week. What it sells is speed. Up to 4.6 times more words per second from the same model with the answer coming out identical. So, the excitement
made sense, but there's a second half to this story and almost none of that traffic reached it because a framework that trains Dflash style drafters and five other kinds beside it has been
sitting in the open for a month. Tencent put it there. Its name is AngelSpec and the repo has 219 stars, 19 forks, and one person watching it. One watcher. And the seven model files that repo
ships have been downloaded 1,773 times across 30 days. Dflash2 did 70,000 in five. And if you search Hacker News for AngelSpec, you get nothing. Not one submission ever. So, this video does two
things. It explains what's actually inside that repo because the ideas in there are worth your time. And it asks why the tool that lets you compare six methods draws a fraction of the
attention of one of the six. We should start with why generating text is slow at all because every idea after this one leans on that. When a model writes an answer, it writes one token at a time. A
token is roughly a word piece, sometimes a whole short word. To produce each one, the chip has to pull the model's entire set of weights out of memory, run a single pass, and emit that one token.
Then it does the whole thing again for the next one. For a big model that's hundreds of gigabytes hauled across the bus per word. The arithmetic sitting on top of it is almost nothing by
comparison. So, the chip spends most of its life waiting on memory rather than computing. Roofline analysis puts this kind of decoding at roughly one floating point operation per byte moved, which is
deep in memory bound territory. You're renting a very expensive machine and leaving most of it idle. The public list price for a B200 ran from $3.49 to $14.24 an hour across clouds in April
2026 and cost per token is just that hourly rate divided by the words you actually got out. Now the trick. If you're paying to haul all those weights out of memory anyway, you may as well
check more than one token per trip. That's speculative decoding and the idea has been around since 2023. A small cheap draft model guesses the next several tokens ahead of time. Then the
big model looks at every one of those guesses in a single forward pass and says which prefix it would have written itself. Everything it agrees with is free. The first disagreement gets
corrected and the loop starts over. What makes this more than a shortcut is the accept or reject rule. It's built so the tokens that survive follow the same distribution the big model would have
produced alone. Not approximately, exactly. So you're not trading quality for speed. You're spending memory trips that were being wasted anyway, which raises the obvious question, if it's
free speed, why is there still an argument about how to do it? Because how you build the guesser matters enormously and the best answer changes with what your users are typing. There are two
families. The first is multi-token prediction, usually written MTPU, bolts a small extra head onto the model itself and it predicts a few tokens ahead. It's light, it's stable and Tencent's own
Honeycomb 3 ships with one built in, a 3.8 billion parameter layer inside a 295 billion parameter model whose entire job is guessing ahead. Block diffusion is the second family.
Here the drafter produces a whole block of tokens in a single pass rather than one after another, working the way an image model does, not painting pixel by pixel, but resolving a whole picture out
of noise at once. Same move applied to text. That means longer guesses and the cost of drafting gets spread across all of them. So block diffusion ought to win
outright and for a while people wrote as though it had. The finding at the center of the AngelSpec paper sits in the first sentence of its abstract. No single drafting structure performs
best across real-world workloads. Think about why. If somebody's chatting, what comes next is wide open with dozens of plausible words and high entropy, so a long parallel guess mostly gets thrown
away. But if the model is writing code or working through maths, the continuation is far more predictable. A closing bracket, the next line of a loop, there a long block guess lands. So
Tencent did something most labs skip. Rather than training one drafter on a blended mixture and calling it general, they trained two specialists, MTP on conversational data, block diffusion on
code and mathematics. Two experts instead of one compromise. And that leaves a question a benchmark table can't answer for you. If the right drafter depends on your own traffic, how
would you ever find out which one you need? That's what the repo is for. Six drafting architectures, one training pipeline, and you move between them with a config flag. Those six names in a row
sound like noise. Look at their dates and they turn into a lineage. Eagle 3 came first in March 2025. It's auto-regressive, so it still guesses one token at a time, but it fuses features
from several layers of the big model to guess better. It became the default across most serving stacks at a claimed 3 to 6.5 times faster. Then in February this
year, a lab at UC San Diego published D-Flash. Same idea, whole block at once, and about 2.5 times the speed up Eagle 3 could reach. That paper is where the release that
trended last week comes from. The lab is run by Ge Yan Liu, who teaches at UC San Diego and works on efficiency at Nvidia. D-Flash went to ICML this year, and its repo now carries 5,894
stars against AngelSpec's 219. In June, a paper called D-Flare pointed at D-Flash's weak spot. Every draft layer was being handed the same fused summary of the big model, so no layer could
specialize. D Flash let each one learn its own blend instead. D Spark arrived a month later from Peking University in DeepSeek with a different complaint. Guessing a whole block in parallel means
no token inside that block knows what got chosen before it. So acceptance falls away as the block grows, and their fix was to put a little of that sequence back, which brings us
to D Fly, Tencent's own from the Angel Spec paper. Look at what it's built from. D Flash's shared projection, D Flash's per layer refinement, and an autoregressive head bolted on the end.
It's the previous three welded together. And D Spark and D Fly were published three weeks apart by two separate organizations having reached the same conclusion on their own. Pure parallel
drafting throws away too much, so put some of the sequence back. So the repo isn't six options on a menu. It's 17 months of an arms race, and it's the only place all of it sits behind one
flag. So in short, where we are. Speed comes from guessing ahead and checking in bulk. The guessers keep leapfrogging each other, and Angel Spec is the workbench holding all of them. Now the
numbers, and what matters isn't a speed multiplier at all, but accepted length. On average, how many tokens survive each round of verification? Every accepted token past the first is a memory trip
you didn't have to pay for. So accepted length is the raw material the speed multiplier gets made from. On Honey 13, the paper's table puts MTP at 3.00,
D Flash at 3.69, and D Fly at 4.79. That gap is the 30% the abstract leads with, meaning 30% more accepted tokens per round than the drafter currently getting all the attention.
On Twin 3, it's tighter. D Fly reaches 5.41 against D Spark's 5.32, which is close enough to call a tie. It's an unglamorous table, and it's the most useful page in the paper because
it's the only place someone has run all of these against one model on one set of benchmarks. Dfly runs 1.98 to 2.40 times faster than ordinary decoding end-to-end, tested from four concurrent
users up to 64. You may have seen 2.86 times quoted instead. That figure comes from Alpha Signals write-up describing the peak on code and math specifically, and it isn't in the paper. If you're
sizing hardware, use the paper's range. There's one more idea in there and it has an ugly name, Dcut. Every system up to now has treated one request at a time. Draft it, verify it, move to the
next. But a real serving cluster has dozens of requests running together and verification is the expensive half. Some of those drafts are almost certainly right, while others are a coin
flip. So Dcut treats verification as one shared pool of compute across the whole batch. It scores each request by how confident the drafter was, prices the work against a profiled cost model of
the actual hardware, and then spends the budget where it pays off. So it checks hard where the drafter was confident and barely at all where it wasn't. On live Honey 1 traffic at 64 concurrent
requests, that's 15.7% more throughput than Dfly on its own with mean accepted length slipping only from 2.50 to 2.46. That trade is the difference between a paper and something that survives
contact with production. And it's the piece of this work I'd expect other labs to copy first. One more number because it undercuts the whole architecture race. If you hold the architecture fixed
and change only how it was trained, Montana Piatesi's mean acceptance rate goes from 52.8% to 66.4% and Dfly's own ablation says the same
thing from the other side. Their backbone change alone took accepted length from 3.77 to 4.40, then to 4.60 with an autoregressive head, and then to 4.75 purely by switching the training
data to code and maths. So the data change was worth almost as much as an architecture change, which is the two specialist argument again, proved on their own ablation table. So, why hasn't
it caught on? Three reasons, and the quality of the work isn't one of them. The first is that this is a training framework, not a download. By itself, it makes nothing you already run any
faster. You point it at a target model, generate hidden states, train a drafter against them, then serve the result. Training reaches sequences of 128K, split across GPUs with U less C sequence
parallelism. If you're not already running your own model on your own hardware, none of that applies to you. Dflash 2 is a file you pull. Second reason, the state of the repository
itself. Three commits total, and the last one landed on the 31st of July. On the 30th of July, a day after release, one account filed 10 issues in about 10 seconds, clearly automated, and the
reports are specific. A background refill task that fails silently on any exception. A cleanup routine that can delete the best checkpoint while the pointer file still names it. And a loss
path that flattens a tensor and breaks the Dspark head downstream, which means one of the six advertised architectures has an open correctness bug in its training code. Seven minutes later, that
same account closed most of its own reports to keep the tracker on the higher severity ones. Three are still open. So is the pull request carrying the fixes, more than 3 weeks on with no
comment and no merge. Third reason, small but telling. GitHub prints the license as other, which usually reads as proprietary, but the file itself says Apache 2.0 with a Tencent preamble that
detector can't parse. So, here's where I land. For most people watching this, Angel spec is not the thing to use. If you want faster local generation this week, take Dflash 2. There are ready
backends for MLX on Apple hardware and for any OpenAI compatible server. The license is MIT. Over 130,000 people have pulled it already, and it works. That isn't a close call, but for the smaller
group serving one model on their own hardware at real concurrency, Angel spec is worth more than the multiplier that beat it. Not because Dfly is the fastest thing in existence,
because it's the only place you can train all six against your own model and find out which one your traffic wants. Their own paper's answer was two specialist, not one generalist. You
can't reach that answer with a repo that ships a single drafter. What I'd argue against isn't Tencent, and it isn't Zijian Liu's lab or Inco AI, but the habit of picking tools by
the largest number in the headline. Look at the three numbers this story produced. 4.6 times, 2.86 times, 1.98 to 2.40 times. Each one is accurate, and each was measured on a different model,
on different hardware, under different load. Only the biggest one traveled. That's what headline worship costs, and it isn't abstract. 74 downloads of the one file for every download of the whole
toolkit, and the toolkit had five times long to collect them. Go and read it anyway, not to deploy, but for the table where six methods meet the same model, which is a thing few labs publish.
And if you do serve at scale, Dfly and Dcut are two ideas worth stealing whether or not you run the code, which leaves the part I can't settle. When the tool that lets you compare six
methods draws a fraction of the attention of one of the six, is that a community being sensible about what's usable today, or a field learning to optimize for whichever number fits in a
post?