Best Local Coding AI for Your GPU (4GB to 512GB)

summarized

TLDR

Fitting a model on your GPU is not enough; you must also account for KV cache and runtime overhead. For most developers with 8-24GB GPUs, a 9B model at 8-12GB or a 27B model at 16-24GB leaves enough headroom for real coding work. The key insight is that a smaller model with room to run usually beats a larger one that barely fits.

Key points

SparkX 2.5 (4B params) fits on 4GB GPUs but requires runtime support for its attention design.

Ornith 1.5 9B at 4-bit fits on 8GB GPUs and scores in the mid-80s on agent benchmarks.

A special build of Qwen 3.8 27B uses smart quantization to fit under 12GB on a 16GB card.

Code 3 Coder Next (80B) runs a coding agent on 48-128GB GPUs with up to 256K tokens of context.

DeepSeek v4.1 flash (552B) can run on a 128GB Mac using DwarfStar to stream lookup tables from SSD.

Tools mentioned

Techniques

  • Quantization (4-bit, 8-bit, etc.)
  • Smart quantization (uneven bit allocation per layer)
  • Mixture of experts (sparse activation)
  • KV cache management and context trimming
  • Offloading layers to system RAM
  • Streaming lookup tables from SSD (DwarfStar)
Transcript (captions)

0:00 A 4-gigabyte laptop GPU and a half-terabyte server can run the same kind of coding model. On the small card, the model is about 2 and 1/2 GB on disk. On the big one, 467.

0:12 Both of them fit. So, fitting was never the interesting question. The one you care about is which model actually helps you write code and which one just loads. That gap is the whole video. You tell me

0:23 how much memory you've got, anywhere from 4 gigs to 512, and I'll give you three things: the model to run, the run time it needs, and the exact point where it stops being useful. Here's the trap

0:34 in every which model fits your GPU chart you've seen. They list weights only. Your GPU doesn't just hold the weights. So, before a single model name, one idea that runs under all of this. When you

0:45 load a model, three things fight for your memory. First, the weights. That's the file you download, and it's the only number those charts show you. Then, the KV cache, which just means the notes the

0:56 model keeps on every token, your prompt and its own answers. It's what saves the model from rereading the whole conversation for each new word. The more context, the bigger those notes. At that

1:07 point, the runtime takes its own scratch space just to do the math. Add the three together. If they're bigger than your memory, it spills onto something slower or it falls over. Say you've got an

1:18 8-gigabyte card, a 9-billion parameter coding model. Quantized down is about 5 and 3/4 gigs of weights. First, those load. That leaves you a little over 2 gigs of room. Ask it one short question,

1:30 and that's fine. Hand it a long file plus a back-and-forth conversation, and the KV cache grows until those 2 gigs are gone. Then, you either cut the context or watch it crash. How big does

1:41 that cache get? A few thousand tokens of code and chat can run into the hundreds of megabytes, and it keeps climbing as the conversation does. On a tight card, the context is what tips you over long

1:52 before the weights do. That leaves you two dials when memory runs short. Drop to a smaller quant, so the weights take less, or trim the context, so the cache takes less. Most people only think to

2:02 swap the model. The context is the other half of the budget, and it's usually the cheaper one to give up. That's why every size I show you is weights only, and why you leave room on top. Fit the weights

2:13 with nothing to spare, and you've built a thing that loads, but can't work. So, that's the budget. Weights, then cache, then scratch. Keep it in your head. We're about to climb from 4 gigs to 512,

2:25 and at every rung the question stays the same. Does it fit with room to actually run? Start at the bottom. 4 GB, that's a huge number of older laptops and cheap desktop cards.

2:36 Your model here is SparkX 2.5, 4 billion parameters, quantized to about 2.6 gigs. It's a proper little agentic model, and it'll take a million tokens of context if you feed it. One catch the size

2:48 charts leave out. It uses its own attention design, so your runtime has to ship Spark 2.5 support, or the file simply won't load. That's the runtime half of the answer,

2:59 and it matters as much as the fit. What's a 4 gig model good for? Explaining code you didn't write, fixing one focused bug, small contained edits. It's a sharp assistant, not something

3:10 you hand a whole repo. Where it falls down is the moment you ask for more than one hop. Give a 4 billion model a task that spans five files and needs a plan, and it loses the thread halfway. That

3:20 isn't the model being bad. It's handing a scalpel a job that wanted a whole team. 6 gigs buys a little room. You could run Neo Horse 1, another 4 billion model, around 3 gigs at a higher quality

3:32 setting. Neo Horse is built on top of Qwen 3.5, and tuned for tool use in coding, and on a 10 benchmark round it beats its own base model by about six points. Same

3:43 class of work though, focused single problems. It's biggest gains land on the agentic tests, the ones about calling tools and staying on task, which is the exact weak spot of models this small.

3:54 So, between the two 4 billion options, pick by which your runtime supports cleanly and which you can keep fed with the context you need. 8 gigs is the first real step up because now a 9

4:04 billion parameter model fits. Ornith 1.5 9B is about 5 and 3/4 gigs at 4-bit. Ornith 1.5 came out in August, openly licensed from a team that trains these models to improve their own coding. At

4:18 9B, local coding starts to feel less like auto-complete and more like a junior pair. On a hard agent benchmark, the family scores in the mid-80s, which for a 9B is strong. The limit at this

4:30 size is room for context, so keep your prompts lean and it'll hold one file and a real conversation about it. Now, a decision people get wrong at 12 gigs. You've got spare memory, so do you jump

4:41 to a bigger model? Not necessarily. You can take that same 9B and run it at 8-bit instead of 4. About 9 and 3/4 gigs and get a cleaner, steadier version of the model you already trust, which

4:53 brings up the thing doing all the shrinking. Quantization just means storing each weight in fewer bits. Say you've got that same 9 billion model. At full precision, 16 bits per weight, it's

5:04 around 18 gigs. Drop to about 4 bits and it's under 6. You've made it a third of the size and for most coding you'd barely feel it. Go too low, under about 3 bits, and it starts making mistakes it

5:15 didn't make before. So, the smart move at 12 gigs is sometimes the same brain thinking more clearly, not a bigger one. And you don't have to guess where that floor sits. For most of these models,

5:25 the 4-to-5-bit range is the safe zone where the benchmark scores barely move. Below 3 bits, quality drops off a cliff, which is why the 1-bit builds are a party trick, not a daily driver. Two

5:36 things so far. 4 to 12 gigs runs a real useful assistant, and the best spend isn't a bigger model by default, it's the right number of bits. What this tier can't do is hold a change that spans

5:49 several files. For that, you need more room, 16 GB. This is where a lot of gaming cards live, and it's where it gets interesting because you can run a 27 billion parameter model on hardware

6:00 that on paper shouldn't hold it. The model is Qwen 3.8 27B in a special build. At a plain 4-bit quant, a 27B is too big for 16 gigs once you add context. So, this build quantizes it a

6:14 smarter way. Instead of squeezing the whole model equally, it keeps the sensitive layers at higher precision and crushes the ones that matter less, all under a fixed size

6:23 target. The result is under 12 gigs, and on some coding and math benchmarks that under 12 gig version matches the full precision model exactly. That's a 27B doing real work inside 16 gigs of

6:36 memory. The method has a name you'll see on the download page, and it's two ideas. One part measures how far each tensor can be squeezed before the output suffers.

6:45 The other spends your size budget where it matters, more bits for the sensitive layers, fewer for the rest. The result on a 3-bit build, it lands on the same score as the full model on a math test

6:56 and a coding test, not close, the same. That's the whole trick. Below a threshold, you lose the model, and a smart quant shoves that threshold way down.

7:05 At 24 gigs, the classic sweet spot card, you run that same model with more bits and, more to the point, real room for context, about 16 and 1/2 gigs of weights, and the rest of the card is

7:16 free for a long file and a long conversation. This is the tier where multi-file edits and running your test suite actually work because the model can hold enough of the project at once

7:26 to reason about it, and that's why the bigger model earns its memory here. A real change lives across files. Rename a function in one, update its callers in another, fix the test that breaks in a

7:37 third. The model has to see all of it at once, and a 9B just doesn't have the room to keep that whole picture in its head. 32 gigs gives you a choice, and the choice teaches you something.

7:47 Option one, max out that 27B at a high quality 6-bit setting, about 22 gigs. Option two, step up to a 35 billion parameter model, Ornets 35B, around 25 gigs. But it has a trick worth

8:01 understanding because you'll see it again higher up. It's a mixture of experts, which just means only a few of its parts run for any one token. So, even though it's 35 billion total, only

8:11 about 3 billion fire per token. It runs closer to the speed of a 3 billion model while carrying the knowledge of a 35. The part people miss, you still have to fit all 35 billion in memory. Sparsity

8:22 buys you speed, not space. You just don't pay compute for every parameter on every word. Where it disappoints is when you expect 35B quality on everything. With only 3 billion active per token,

8:34 the hardest reasoning can feel thinner than the headline size promises. Quick and capable for its speed, not a free lunch. Where are we? 16 to 32 gigs is the real sweet spot for most developers.

8:45 Multi-file changes, tests, a model that keeps up. And the winner here isn't the biggest number that fits, it's the setup that leaves room for the context's real work needs. Now we leave consumer

8:56 territory. 48 to 128 GB is workstations, rented cloud GPUs, and the top of the enthusiast world. This is the first tier where a model can run a coding agent. Not just answer a question, but work a

9:09 task across many steps, editing, testing, and recovering when it gets something wrong. Picture the loop. First, it reads the failing error. Then it edits the file. Then it runs the

9:19 tests. At that point, it reads the new error and goes again, holding the whole chain in context. That's the line between a chat box and something that finishes the task while you get a

9:29 coffee. The workhorse across most of this range is Code 3 Coder Next, an 80 billion parameter model built for exactly that. 48 gigs runs a compact 4-bit version

9:40 near 38 gigs. 64 gigs runs a fuller 4-bit, about 48 and 1/2, and that one's the official figure straight from Quen's own repository. 80 gigs runs a 6-bit near

9:51 66. It takes a quarter million tokens of context natively, which is what lets an agent keep a whole task in view. One report from someone actually running

10:01 it, and I'll flag it clearly, this is one person's experience, not a controlled test. They got useful coding out of Coder Next on a 24 gig GP plus 64 gigs of system RAM by offloading some of

10:12 the model's layers onto the slower RAM. It worked, but reliability was mixed, and that offload trick exposes the real bottleneck at this tier. It isn't the chip, it's the bus between the GPE and

10:24 system memory. When the model keeps reaching across that bus for the parts you pushed into RAM, the fast chip sits there idle, waiting on data. Fix the memory fit, and you can still be

10:34 bottlenecked by the road between the two. It's worst in one phase. When it first reads a long prompt, the model has to pull those offloaded parts across the bus before it writes a single word. So,

10:44 the GPU just waits. A good runtime hides some of that by fetching ahead during compute. A careless setup leaves the fast chip idle half the time. Then a detail at 94 and 96 gigs that breaks the

10:56 bigger number wins Instinct entirely. Those two sit 1 GB apart on the chart. The 94 is an H100 NVL, a data center card with fast HPM memory and a dedicated high-speed link between chips.

11:10 The 96 is an RTX Pro 6000, a workstation card with GDDR7 and a slower link. On one card, benchmarks have the 96 actually edging ahead. Put several together, and the 94 pulls away because

11:23 its link between chips is many times faster than the workstation cards. 1 GB apart, and speed goes to whichever one your workload needs. Memory fit told you nothing. To put rough numbers on it, and

11:34 these come from published benchmarks, not my own bench. On a single card, the workstation 96 has come in slightly faster on raw throughput, but the data center card's link between chips moves

11:44 hundreds of gigabytes a second against the workstation's low tens. Across four cards, that gap is the whole ball game. And in this range, you meet a model that forces you to count carefully. When 3.8

11:55 flash next, people call it 125 billion parameter model, and that's the mistake. It's 125 billion main model plus 51 billion parameters of lookup tables, plus a 4 billion prediction head that

12:08 lets it guess several tokens at once. Budget only for the 125, and you run out of memory partway in. The good part, those 51 billion of lookup tables can live in ordinary system RAM instead of

12:20 your GPU because the model knows where it'll need them and fetches them ahead of time. So, the job isn't fitting all of it on the card. It's putting each piece where it belongs. That 4 billion

12:31 prediction head isn't decoration, either. It lets the model draft several tokens at once and check them together, which is a real speed trick. And the main body alternates a cheap

12:40 linear attention with a sparse one, so most of it stays light to run despite the size. One more single person report, same caution as before. Around 124 tokens a

12:51 second on 296 gig cards with the lookup tables held in RAM. Quick, but it leans hard on the exact setup. So, the 48 to 128 tier runs a real coding agent. The condition is that you count every

13:04 component and respect the bus. That's two tiers behind us. Now, the frontier. 141 GB and up. Multi-GPU rigs, big unified memory max, the serious end of local AI. Up here, the headline is GLM

13:18 5.3 flash, and above it, the full GLM 5.3. Flash runs from about 120 gigs at a low quant up to 240 at a high one across the 141 to 288 range. The full model fills

13:32 the top from around 280 gigs to 467 at 512. That 467 is the number from the very start of this video, but there's a runtime catch that will burn your afternoon. This model's architecture

13:45 isn't in the standard runtime yet. You need a specific branch that the Unsloth team maintains or their desktop app. Download the weights without the right runtime and nothing runs. The reason is

13:56 simple and annoying. It's a brand new architecture and the main open source runtime hasn't merged support for it. As of a couple of weeks ago, it still wasn't in the main line.

14:05 So, today the working path is the maintainers own branch or the app they ship and that's it for now. Now the payoff this whole climb was building toward. At 256 gigs, you could load GLM

14:16 5.3 flash at a high quant about 200 gigs. Or for almost the same memory, around 208, you run MiniMax N3, a 426 billion parameter model. It sounds like the obvious bigger pick, but MiniMax is

14:31 a mixture of experts that fires only 23 billion parameters per token. So, it carries far more knowledge, yet runs light and leaves you room for context and speed. Same memory and the

14:42 sparser model is usually the more practical one to actually work with. Larger versus smarter and smarter keeps winning. Think about your actual afternoon. The GLM at 200 gigs eats most

14:52 of your 256, leaving little for a long file. The MiniMax at 208 holds more total knowledge, but fires a fraction of it, so it runs faster and leaves headroom.

15:03 On the same box, the sparser one is the model you'd rather use all day. Two distinctions up here that cost people real money. First, unified memory, which is one pool the whole computer shares.

15:14 512 gigs of that is not 512 gigs of GPU memory. A Mac with 512 unified shares that pool with the entire system at slower bandwidth. Two data center cards giving

15:25 you 512 total are fast, but the model has to be split across both, which is its own headache. Same number on the spec sheet, completely different machine. Bandwidth is the quiet

15:35 difference. A unified memory Mac is convenient, and it's real memory, but its bandwidth sits well below a stack of data center cards. So, a model can fit comfortably on the Mac and still write

15:46 tokens slower than that big memory number would ever lead you to expect. Second, and it earns the repeat. Fitting a model has never once proved it runs well. Every tier we've climbed says the

15:56 same thing a different way. There's one more model that deserves its own segment, and it's still experimental, so treat it that way. DeepSeek v4.1 flash, released September 10th, 552 billion

16:08 parameters, openly licensed, and it tests ahead of the company's own larger model. A Mac engine called DwarfStar does something clever with it. It keeps one big chunk of the model, a

16:19 lookup structure, permanently on your SSD instead of in memory, and streams it in as needed. That's how you can run a compressed version on a single 128 gig Mac, and it's why bumping the quality up

16:30 adds 142 gigs, because the main model grows while that lookup structure just stays on disk. Why can that part live on an SSD when the weights can't? Here's how it works. When it reads that table,

16:42 it does so only now and then, and predictably enough to fetch from disk in time. The core weights it touches constantly, so those stay in fast memory.

16:51 Once you split the model by how hot each piece is, a 500 billion parameter model runs on a laptop's worth of RAM. It rewrites the entire weights only chart without any asterisk, because part of

17:02 the model never enters your memory at all. That's the frontier direction. Not a bigger box, a smarter place to keep each piece. It's early and it's finicky, so don't bet a deadline on it yet.

17:13 But, it's the most interesting idea at the top of this chart, because it stops treating memory as one flat wall and starts treating it as a hierarchy. Fast memory for the hot parts, slow disk for

17:23 the cold ones. And if you land between tiers, go back to those two dials from the start. When a model almost fits, drop one quant level or trim the context and it will. You need the next GPU up

17:35 far less often than the charts want you to believe. So, back to where we started. The 2 and 1/2 gig model on the 4 gig laptop and the 467 gig model on the half terabyte server, both fit and

17:47 fitting told you almost nothing. Here's the call. For the GPU most of you actually own, somewhere between 8 and 24 gigs, the right pick is the model that leaves headroom, not the biggest one

17:57 that squeezes in. A 9B at 8 to 12, a 27B at 16 to 24. That's real coding for most people watching. The exception is real. If you're running an agent through long sustained work, the 48 to 128 gig models

18:11 earn their keep and the frontier machines exist for good reasons. But even at 512 gigs, a smaller, sparser model that leaves room to run usually beats the giant that barely loads.

18:22 Memory fit is the floor, not the finish line. And the way you settle it isn't a chart, mine included. It's five things. The model file, the runtime it needs, the context you gave it, the time it

18:33 took to reach a correct fix, and the one place it failed. Fits is a spec. Helps is a test. A chart can only ever tell you the first one and you just watched the first one lie at every single tier.

18:44 Run that on your own hardware, on your own code, and you'll have your answer in an afternoon. If this saved you a bad download, subscribe. One of these every week.

Frontier News · by Hyperjump Technology