What Would It Cost to Run GPT-6 Astra Locally? (I Did the Math)

summarized

TLDR

Running OpenAI's GPT-6 Astra locally is infeasible due to fundamental bandwidth constraints, not just cost. The model is estimated at 7 trillion parameters (3.9 TB), requiring dozens of high-end machines that still deliver only 1-6 tokens per second, slower than a human can type. The electricity alone to keep the hardware running costs six times more than renting the same token volume from OpenAI, and hardware payback takes roughly 60 years. The real bottleneck is bytes-per-token memory traffic and inter-machine cable speeds—physics that no credit card can fix.

Key points

GPT-6 Astra is estimated at 5-10 trillion parameters, with ~7 trillion as the working figure.

The model requires reading 156 GB of memory traffic per generated token.

A single DGX Spark can produce only about 1.3 tokens per second running Astra.

Distributing the model across machines via tensor parallelism is bottlenecked by cables 120x slower than internal memory bandwidth.

Electricity alone for a local inferencing cluster costs over 6x the price of renting the same tokens from OpenAI.

Tools mentioned

Techniques

  • Compound compute growth estimation
  • Chinchilla scaling law
  • Sparse activation (MoE-style)
  • Recurrent depth reasoning
  • Tensor parallelism
  • 4-bit weight quantization with mixed precision attention
Transcript (captions)

0:00 Open AI put this page up yesterday. It's the official spec sheet for their newest model GPT-6 Astra. $10 a million tokens going in, 50 coming out, a million tokens of context. Every number you'd

0:14 want except one. Nowhere on that page or on any other page does it say how big the model actually is. Open AI hasn't published a parameter count since GPT-3 and that was 6 years ago now. So here's

0:27 what we're doing today. We're going to weigh it anyway. Then we're going to go shopping. By the end you'll know what Astra probably weighs, how many machines that is, and why buying more machines

0:38 won't help. You'll also get the one number that ends the whole argument and it isn't the price of the hardware. First, the weighing. To weigh a thing that its owner won't put on a scale,

0:49 what you need is a ruler. There's exactly one usable ruler on the internet right now and the good news is you can download it. Kimi K3 from Moonshot AI, the weights went public in July. 2.8

1:01 trillion parameters. And because it's public, its file listing tells you what Open AI won't. 96 shards, 1.5 terabytes of model. Divide one by the other and you get the number that does all the

1:13 work today. Just over half a byte per parameter. 4.5 bits each, give or take. That is the real exchange rate between a parameter and a byte on your disk. Kimi's own config file says why. Think

1:27 of it as writing every number in shorthand. The weights are stored 4-bits wide, packed in groups of 32 with the attention layers kept sharper because that's where precision actually matters.

1:38 That shorthand is why the download is a terabyte and a half instead of six. So we have a ruler, which means we need one more thing, a guess at Astra's parameter count. And I

1:49 want to give you a few of them. Different methods so that when they agree, the agreement actually means something. The first guess is the textbook answer

1:58 and I'm showing it to you mostly because it's wrong in an interesting way. Epoch AI tracks training compute across the whole frontier, and it's been growing four to five times a year, every

2:08 year. GPT-4's training run has actually been measured, and the figure is on screen. Astra came 3 and 1/2 years later. Compound that growth across the gap, and Astra got somewhere around 200

2:21 times the compute. Then you apply the classic scaling result called Chinchilla. It says the ideal parameter count grows with the square root of compute. Square

2:31 root of that is about 14. 14 times GPT-4 gives you 24 trillion parameters, and that answer is wrong, which is the useful part. The labs abandoned Chinchilla years ago. They overtrain

2:44 instead. Same parameters, far more tokens, because the cost that actually hurts is serving the model to millions of people, not building it once. So parameters grow much slower than that

2:55 square root. Hold on to that one, because it comes back at the very end. The second guess is the ruler again. Historically, the closed frontier runs two to three times the size of the open

3:05 one. Apply that to Chime and Astra land somewhere between five and eight trillion. The third guess is the price tag, and it's my favorite, because this is OpenAI telling on itself. Astra costs

3:18 $50 per million output tokens. The model it replaces, Soul, costs 20. That's two and a half times more, and serving cost per token tracks one thing above all others, how much of the model has to

3:31 switch on for each token. So Astra is lighting up roughly two and a half times as much silicon per token as the model before it. Two of the three methods land in the same place, and the third rules

3:42 out the ceiling. So here's the estimate. About seven trillion parameters in total, of which maybe 280 billion actually fire for any single token. That gap is what people mean by a sparse

3:54 model. Think of a workshop with every tool you can imagine hanging on the walls. Any single job only pulls down a handful of them. But you still need the walls and

4:04 the whole building and the rent. The walls are what you're paying for here. To be clear, that is an estimate. OpenAI has published nothing and anyone quoting you an exact figure is guessing, too.

4:15 The working band is 5 to 10 trillion and I'll keep showing you the working so you can move the number yourself. Now, here's the part I didn't expect. 280 billion active is almost exactly what

4:27 analysts estimated for GPT-4. 3 years apart. 200 times the training compute and the amount of model that wakes up for a single token barely moved. It got wider, not busier.

4:40 Which sounds like trivia, but it's the single reason your desk has any hope in this video at all. Remember it. Right, the shopping. 7 trillion parameters at just over half a byte each. Astro weighs

4:52 about 3.9 terabytes. That's the download before you've stored one token of your own conversation. Two and a half Kimmys stacked on top of each other. Apple's new Mac Studio tops out at 128

5:06 GB on the M5 Max chip. Divide the weights by that and you need 31 of them. 31 machines to hold a single copy of a single model. Not 31 users. One. Nvidia's DGX Spark gives you the same

5:20 memory and unlike Apple, it publishes a price. Just under $4,000 a box. So, the stack comes to $124,000 sitting on your desk doing nothing at all yet. Step up to the bigger M5 Ultra

5:35 and each box holds twice as much, so you only need 16 of them. Call it $150,000. Apple sells a version with four times that memory from late October. Eight of those would do it. They haven't

5:48 published a price yet, so I'm not going to invent one for you. Or you could go the gamer route with the fastest card money can buy. You would need 122 of them. A quarter of a million dollars of

6:00 graphics cards pulling 70 kilowatts. Your house has opinions about that. For contrast, the card Open AI actually runs this on moves nearly 5 terabytes a second. 28 of those hold Astra. Three

6:13 and a half servers in a room built for exactly this on a floor built for the room. And every count I just gave you is optimistic because you never get all of the memory.

6:24 Leave a slice for the operating system and the driver and your spark count climbs from 31 to 36. So we have a price and it's a big one, but it isn't a scary one. Plenty of companies spend that on

6:36 office chairs. If money were the only wall here, somebody would have simply paid it by now and posted the benchmark. They haven't. Here's why. And it's the part that never the trip into a

6:47 spreadsheet. To write a single token, the machine has to read every active parameter out of memory. All of them. Exactly once. 280 billion parameters at half a byte each is 156 gigabytes. Moved

7:03 per token. Not per answer. Per token. The word the cost you 156 gigabytes of memory traffic. Which means the speed limit isn't the chip at all. It's like a water pipe. A bigger pump does nothing

7:16 if the pipe stays narrow. Tokens per second is just bandwidth divided by that number. So let's divide. A DGX spark moves 273 gigabytes a second. D-rate that for what real inference engines

7:30 actually achieve and you land on about one and a third tokens a second. A token every three quarters of a second. You can type faster than that. You could probably write the answer faster than

7:41 that and you don't even know the answer. The M5 Ultra does better at a whole 1.2 terabytes a second. Call it six tokens a second, which is roughly reading speed. Now, you just bought 31 boxes. Surely

7:55 that's 31 times the bandwidth? It isn't, and this is the trap that gets people. Split a model across machines the normal way, and the token walks through them in order. Box one, then box two, all the

8:08 way to the last one, then out the other side. It's like a relay race where only one runner is allowed to move. So, the whole stack gives you exactly the speed of a single box. You bought capacity.

8:20 Speed was not on the shelf at any price. There is a way to add bandwidth. You slice every layer across every box, so all of them read at once. That's called tensor parallelism, and it works right

8:33 up until you look at the cable, because now every layer has to stop and share its results with all the others 90 odd times per token. And Thunderbolt 5 carries 10 GB a second between boxes.

8:45 The memory inside one Mac Studio moves 1,200. So, the wire between your machines is 120 times slower than the memory inside them. It's like plumbing a supercomputer with a drinking straw.

8:58 Nvidia's link is a lot better, and it's still 10 times slower than its own memory. Step back a second, because this is the shape of the whole problem. You can buy

9:07 the capacity, you cannot buy the speed. And you cannot buy your way around the cable, because the cable is a physics problem wearing a price tag. Which leaves the bill, and this is where the

9:18 argument actually ends. Each Spark pulls 240 W. The stack pulls 7 and 1/2 kW. A normal wall circuit gives you 1,800 W, so you run out of socket at box number seven. Run it flat out for a year at the

9:33 American average electricity price, and you'll spend about $12,000 on power. Not on hardware, not on maintenance, just electricity, just to keep the memory awake. So, what does that

9:45 electricity actually buy you? At a token and a third a second, running every second of every day, about 41 million tokens. Buy those same tokens from OpenAI at list price, and they cost you

9:58 around $2,000. That comparison is the number that ends this. The electricity alone is nearly six times the price of just renting it, before the hardware, before you have plugged in a single

10:10 cable. And the hardware itself, divide that bill by the list price, and you have bought about 2.5 billion tokens of Runway. At the rate this cluster writes, it

10:20 breaks even in roughly 60 years. Your grandchildren inherit a working GPT-6, and there's a final twist, a cruel one. Astra doesn't think the way older models did. OpenAI says it reasons using

10:33 something called recurrent depth. Instead of writing its thinking out as words you can read, it loops the same block of weights over and over inside itself, where you can't read it, which

10:45 means every loop is another full read of those 156 GB. Loop it four times, and a token takes 3 seconds. So, every speed number I just gave you is a ceiling, not a target. The real thing is slower, and

10:59 OpenAI won't say by how much. So, the verdict, and it isn't close. For Astra, today you rent, not because you're poor, because you're bound by physics. The wall isn't money, it's bytes per token

11:12 and GB per second, and no credit card fixes either of those. But, run that same arithmetic backwards, and you get a far more useful answer. This is the answer I'd actually act on.

11:24 Take Chinchilla 3, the ruler we started with, and squeeze it down to about a bit and a half per parameter. Now, it's under 600 GB. That's two Mac Studios. Deep Seek's big model at 4-bit fits

11:38 inside a single machine. Their fast one is smaller still. A single box on your desk for the price of a decent laptop. No cluster, no cable, no drinking straw. And that box

11:50 runs a model that was frontier class about 18 months ago. Not a shrunk-down teaching version. The real weights. One generation back, answering at reading speed while you sit there and watch.

12:02 That gap has held at roughly 18 months for 3 years running, which makes it the most reliable number in this entire video. So, here's what I keep turning over, and

12:12 I don't have the answer yet. Recurrent depth means a model can get smarter by thinking for longer, rather than by getting bigger. If that's where the frontier is heading,

12:22 does the gap stop being about how much memory you can afford, and start being about how long you're willing to sit and wait? Because waiting is the one resource your

12:31 desk has more of than a data center.

Frontier News · by Hyperjump Technology