1,000 Tokens/Sec on One RTX 3090 (Here's the Config)

summarized

TLDR

A configuration of nine software changes on an RTX 3090 achieves 1,000 tokens/sec for 64 concurrent users by addressing memory bandwidth, quantization, and kernel utilization. The real story is that inference speed is primarily a software configuration problem, not a hardware one, and many benchmarks may be misleading because they don't check output quality.

Key points

The configuration achieves 133 tokens/sec for a single user and roughly 1,000 tokens/sec for 64 concurrent users on an RTX 3090.

Quantization serves two purposes: 4-bit weights to fit the model, and 8-bit integer activations to move computation onto integer tensor cores, quadrupling the rate.

The 8-bit activation path initially produced garbage output because the kernel read scaling numbers as unsigned; two patches fixed it, highlighting that speed benchmarks can look perfect while answers are broken.

Three-quarters of the model layers use gated DeltaNet (linear attention) with a fixed-size recurrent state; the state was stored in 32-bit float (150 MB per request), causing queuing; halving to 16 bits allowed all 64 requests to run without perceptible quality loss.

Speculative decoding uses a draft head with 4-bit weights, calibrated on the model's hidden states, and a smaller vocabulary covering 97.5% of outputs to increase acceptance rate.

The standard attention kernel for verification used only 24 of the 82 compute units; a replacement kernel cut that layer's time by more than half.

The fastest mode is for a single user; adding users dramatically reduces per-user speed, and the fastest trick is prompt-copying where the model drafts from the prompt itself, achieving 381 tokens/sec on long documents.

The tuned stack loses about 1 point on instruction following benchmarks and a couple percent perplexity, but the trade-offs are documented; an RTX 4090 with twice the power was only 1.9% faster, confirming bandwidth is the bottleneck.

Techniques

  • quantization
  • speculative decoding
  • draft head
  • prompt-copying
  • attention kernel optimization
  • state precision reduction
Transcript (captions)

0:00 This graphics card is almost 6 years old, an RTX 3090. It carries 24 GB of memory, which is the only number that matters for running models at home. Put a current model on it, the new 2 3.8

0:12 open weights out this month. Then take a serving setup a Danish engineer published 9 days ago. Turn off the trick that makes it fast and ask it a question. 46 tokens a second. Turn the

0:23 trick back on with the rest of his stack and the same card gives you 133. Same weights, same server. Same 6-year-old silicon. And if you stop chatting and serve 64 people at once, that one card

0:36 holds around 1,000 tokens a second. He bought nothing. He swapped nothing. He actually turned the card down to 250 W on a part rated for 350. What he changed was nine things in

0:47 software. So, where does a near tripling come from if it isn't coming from the hardware? His name is Mads Hendrickson. He's a founder at a seven-person AI shop in Copenhagen called syv.ai.

0:58 And the repo is 9 days old. It already has hundreds of stars and a tracker full of people running it on cards he has never touched. Here's the map. I'm walking five of his nine changes and

1:09 each one teaches a different reason a model is slow. The first is the one most guides get backwards, quantization. You've probably heard that word in exactly one sense, shrinking the model

1:19 so it fits. At full precision, this thing wants more than twice what your card has. Squeeze the weights down to 4 bits and it fits with room left over. That's the fitting trick and it's real.

1:30 But fitting isn't feeding and here's the arithmetic the write-up skip. To make one token, the card has to read every weight it owns out of memory, 15 GB of them, moved at about 900 GB a second.

1:42 Divide one by the other and you get 61 passes a second. That's the ceiling. His no speculation number sits 3/4 of the way to it and no amount of shrinking gets you past it. Now, look at that

1:52 ceiling again because it's hiding something. It only applies to one user. Put 64 people on the card and it reads those weights once and answers all of them in the same pass. Memory stops

2:03 being the thing you wait for. You start waiting on the multiply instead, which is where the second meaning of quantization lives. The weights are already four bits. Quantize the numbers

2:13 flowing through them, the activations, down to eight-bit integers and the multiply moves off the floating point units and onto the integer tensor cores. Same silicon, four times the rate, one

2:24 word, two completely different bottlenecks. The first makes the model fit, the second makes the machine feed. And on his ladder, feeding is the single biggest jump in the file, more than 200

2:34 extra tokens a second from one setting. Except when you first switched it on, the model spoke fluent nonsense. The engine already had an eight-bit path. It benchmarked beautifully and it served

2:44 garbage because the kernel read its scaling numbers as unsigned and about half of them were negative. Two patches fixed it and one of them just folds the sign into the weights as they load.

2:54 Gotcher number one in his own docs says it better than I can. A benchmark cannot tell you the output is garbage. He had an hour of gorgeous throughput numbers before a quality check caught it. So, if

3:05 a speed number can look perfect while the answers are broken, how many of the ones you've read this year did anybody check? So, that's the compute half. The memory half hides somewhere stranger.

3:14 Three quarters of this model isn't attention at all. Most of its layers are a linear attention design called gated DeltaNet. And instead of a cache that grows with

3:23 your conversation, each one carries a fixed-size chunk of state. Fixed-size is the good news. The bad news is the precision it asks for. The config wants that state in full

3:33 32-bit float, about 150 megabytes per request, allocated the moment a request arrives and rewritten on every single step. So, when he asked the server for 64 requests at once, the log told him

3:45 the truth. 37 running, 27 waiting. That queue had nothing to do with the KV cache. It was the recurrent state, a number in a config file. One flag halves it to 16 bits, all 64 run. Throughput

3:58 climbs by more than a third, and perplexity doesn't move to two decimal places. Two changes down, both of them about serving a crowd. So, go back to one person at a keyboard, where that

4:08 ceiling is real again. Past a ceiling like that, the only move left is to make one read pay for more than one token. Speculative decoding does exactly that. A cheap little model guesses the next

4:19 two tokens, and the big model checks all of them in a single pass, keeps what it agrees with, throws the rest away. Because of how that check works, the text you get is the same distribution as

4:28 the slow way. The 20 team ships a draft head for this, and out of the box it barely paid. Every guess it made had to run the model's full output layer, a quarter of a million rows wide. So, each

4:39 extra draft cost 3 milliseconds of dead overhead. So, he shrank the drafter to 4 bits, calibrated on the model's own hidden states. Then he did the clever part. He built it a smaller vocabulary,

4:51 counted over the model's own writing. That sounds like housekeeping. It's worth 10% on its own, and the reason is sharp. A draft head can only propose words that are on its list. So, a

5:01 missing word is a guaranteed rejection, and a rejection ends the chain. His list covers 97 and a half percent of what the model actually says. The generic web text list he started with covered 92.

5:13 Then there's the one I'd frame and hang on a wall. Checking a block of guesses means several queries per request, and the standard attention kernel only spreads its work across the chip when

5:23 there's exactly one. So, the verify step ran on 24 of this card's 82 compute units. The rest sat idle mid-benchmark on the exact step the benchmark was timing. A replacement kernel cut that

5:34 layer's time by more than half. The last piece I didn't see coming. A lab called Inco published a better drafter 6 days ago, one that proposes seven tokens in a single shot instead of chaining four.

5:46 And their post opens with the line, "This whole repository is really about inference is the bottleneck of the agent era." But the free trick isn't the drafter at

5:54 all. Think about what a long context assistant actually does all day, quoting your document back to you, reproducing text that's sitting verbatim in the prompt it was handed. So, when the model

6:04 starts copying, draft from the prompt. Those guesses cost nothing to make because they're already written down. On a long document you pasted in yourself, that takes him to 381 tokens a second.

6:15 Same answers, character for character. So, that's the recipe. Now, what it costs because a stack this aggressive isn't free. On the instruction following benchmark, the

6:25 quantized stack gives up about one point against the untouched model. Grade school maths comes back in the mid-90s across every configuration he ships. The 8-bit activations cost a couple of

6:35 percent of perplexity, and he prints the whole table so you can pick a gentler row. The real limits are shape, not quality. His fastest mode is a one-person mode. Add a second user, and

6:46 each of you loses about a third of that speed. And at four users, two-thirds of it is gone. And all of this is patches against one specific version of the engine. Upgrade, and you reapply them by

6:56 hand. Here's the receipt that made up my mind, and it came from a stranger. Somebody with a 4090 at the office ran the identical stack with no power cap at all against his 250 W limit. It finished

7:08 1.9% faster. A generation of newer silicon and nearly twice the power bought 2% because this workload is bandwidth, and the newer card has 8% more of it. So, here's where I land. If

7:20 you already own a 24 GB card, the correct next purchase is not a second card. It's a weekend. Speed on this workload isn't something you buy, it's something you configure. And the

7:30 industry sells it as silicon because silicon is the thing with a price tag on it. There's a real other side to this. If your users each bring their own long document, this collapses to 15 tokens a

7:40 second, and you want plane batching or a second card. For everybody else, it's the best value in local inference right now and I'd take it at twice the price. One last thing.

7:51 This repo went up on Hacker News two days after it appeared and the post got a single upvote. Today his tracker carries independent reproductions from four different machines on cards he does

8:00 not own. If one engineer with one old card found nine of these, what else is sitting unmeasured in the defaults of everything you run?

Frontier News · by Hyperjump Technology