Qwen 3.6 35B on 3GB RAM?

summarized

TLDR

Edge Zero's 2.9 GB memory claim for running Qwen 3.6 35B is technically accurate but misleading: it measures what the MLX allocator handed out, not the machine's total memory cost. The real requirement is fast storage plus enough spare RAM to cache a 19.5 GB checkpoint in the page cache. The project's key engineering innovation is a trained 'pre-router' that predicts which experts the next layer will need, overlapping I/O with computation and eliminating stalls from wrong predictions. On a 24 GB Mac with a fast drive, this makes a 35B model usable at ~14 tokens/second; on an 8 GB machine already in swap, it drops to 0.7 tokens/second.

Key points

Edge Zero runs Qwen 3.6 35B on Apple Silicon using 4-bit quantization and expert offloading to SSD.

The 2.9 GB memory figure measures MLX allocator usage, not total system memory including page cache.

A trained pre-router predicts next-layer experts one token ahead, overlapping I/O with computation.

The model uses 4 experts per token instead of 8, halving the per-token weight transfer to 283 MB.

On an 8 GB M1 in swap, throughput drops to 0.7 tokens/second versus 24 tokens/second on a well-provisioned machine.

Tools mentioned

Techniques

  • 4-bit quantization
  • mixture of experts (MoE)
  • expert offloading to SSD
  • trained pre-router for expert prediction
  • memory-mapped file I/O
  • page cache warm-up
Transcript (captions)

0:00 A 35 billion parameter model running on a Mac Mini reporting 2.9 GB of memory. That's the project's own benchmark table. The footnote printed under it is why I wanted a closer look because it

0:11 changes what that number means for your machine because the weights didn't shrink. I added up the publish checkpoint myself. 19 1/2 GB. 18 of those gigabytes are expert weights and

0:22 they stay on the SSD while the model is talking to you. So, the reading is honest. It's the peak the MLX allocator handed out, not what the machine spent. The project is edge zero. Apache 2.0

0:35 open sourced on the 8th of September. About 1,600 stars today. Two questions then. If the weights live on a drive, why isn't every token painfully slow? And what is that figure leaving out? By

0:46 the end, you'll know both. And which number to check first? Underneath it is Quen 3.6. Most of the coverage still says 3.5. The maintainer corrected that this morning. A mixture of experts model

0:58 cuts each layer's feed forward block into many small ones. This checkpoint has 256 per layer, 40 layers, and exactly four of those experts fire per token. The rest do nothing at all, and

1:10 you still have to keep them somewhere. One shared expert per layer stays in memory, and it's small, 47 megabytes in total. Everything else has to arrive when a token asks for it, which raises

1:20 the objection you're probably already making. Pulling weights off a drive for every single token should be slow enough to make the whole thing pointless. So, let's work out what one token actually

1:30 costs. Safe tensors files publish their headers separately from the weights, so you can read those over HTTP without downloading a single parameter. I did exactly that. Every layer's expert block

1:42 comes to 453 megabytes. Their own streaming dock says around 310. The ship checkpoint disagrees with the dock, so I'm going with the file you'd actually download. One expert is 1.77 megabytes.

1:55 Four per layer, 40 layers, and a single token needs 283 megabytes of weights that aren't in memory yet. Run that at 15 to 18 tokens a second. And you need four to 5 GB a second sustained, not as

2:08 one long stream either, as 160 separate little reads per word. Back in July, we watched Kim K3 stream off disk on an M1 Max. Same basic idea, none of this machinery, and it produced one token

2:21 every 14 seconds. Far bigger checkpoint, but that's the shape of the naive version. A drive is perfectly happy reading one long file. Scattered reads by the 100, each with a matrix multiply

2:32 stalled behind it, is a completely different workload. The sequential number on your drives box does not apply, and that's before the operating system gets involved. So, that's the

2:41 bill for one word on your machine, and it doesn't look payable. Either those reads aren't reaching the drive or nothing is waiting on them. It turns out to be both. And the second half is what

2:51 the title is about. Start with the stall they had to remove. Routing depends on the previous layers output. So you learn which four experts you need at the precise moment you need them. That's

3:01 fine when the weights are already in memory. On a drive, it's ruinous. Every layer of every token would open with a weight. So their answer is a small trained network per layer. And they call

3:11 it the pre-outer. 33 of them sitting on most of the stack. Each one a 512 wide hidden layer. It ships as its own weights file next to the model. And the shape of the guess is the clever bit.

3:23 The head sitting on layer n doesn't predict its own experts. It predicts what layer n plus one is going to want. And that layer uses the prediction made a whole token earlier. Take one token

3:34 start to finish. First, while the model is still finishing the word before yours, the head on layer 19 names four experts out of 256. Then those four bundles 1.77 megabytes

3:46 each are read while that arithmetic is still running. At that point, layer 20 arrives and its experts are already staged. It routes on exactly those four. Nothing is thrown away and the read you

3:57 were dreading finished a whole token ago. That overlap is the entire design and they put it at up to 59% more decode throughput. They're careful about when you get that gain, too. Slower storage,

4:08 bigger model, more experts per token, and it buys you more. Every writeup I read calls this pre-fetching, which sounds harmless enough. Their own files say something stronger in two places.

4:19 The tier documentation says routing is supplied by the trained pre-outer heads and that the mixture of experts block routes on the pre-outers logits. A comment in the source repeats it almost

4:29 word for word. At decode time, the guess never gets checked against the real router. Whatever the head predicted is what runs and nothing downstream tells you it was wrong. Zero drop by

4:39 construction they call it because the set that got loaded is a set that fires. So a wrong prediction doesn't cost you a stall. It changes which experts computed your token. That's why the thing is

4:49 trained rather than heristic, why there's a flag to switch it off, and why you can see the blast radius in the bug tracker. On the 11th, someone reported outputs collapsing into exclamation

4:59 marks and every later request in that process going the same way. It didn't happen with the pre-outer off. A layer overflowed in 16 bit. The logit went to not a number and Argmax landed on token

5:10 zero. Fix the next day. That's how the drive stops being the bottleneck. Two things in this release that I haven't seen mentioned anywhere though. The first is sitting in the config file in

5:20 plain sight and it changes every number you've heard so far. Quenzone config for the base model says eight experts per token. Edge zero pins it to four. Half the routed width decided by the runtime

5:31 rather than by the checkpoint that ships with it. That having is what takes the bill from 566 megabytes per token down to 283. It's the difference between needing 10

5:42 GB a second and needing four. It also means the quality gap I'll get to shortly isn't purely a quantization story. They changed how the model routes and then they measured it. The second

5:52 thing is the cache. There's one shared cache for decoded experts and it holds 64 of them for the whole model, not per layer, 64 in total. A single decode step touches 160 experts on this model. On

6:05 the smaller one, it's more. Either way, the cache holds under three slots per layer. Their own source comment says what that produces, and it's the sentence that reframed this whole thing

6:14 for me. The release slot count measured exactly zero hits, rebuilding 184 bundles every step. zero hits. So, whatever's making this fast, it isn't the expert cache they ship. It isn't

6:26 pinned resident experts either because this tier switches those off. What's left is a prefetched thread pool and the union of recently used expert sets. And underneath all of it, inside the

6:37 operating system sits the thing that's actually doing the work. Experts are read through memory mapped bite ranges. So, the page cache is what's really holding them. Their Redmi even defines a

6:47 warm run as exactly that, the pages being resident, which is the memory the headline figure doesn't count. 2.9 gigabytes is what the allocator asked for, not what your machine hands over.

6:58 Their own code sets the acceptance bar at 3.4 GB and 13 tokens a second, a less flattering pair than the Redmi's. We can test that directly because they publish the cold and warm split. First request

7:10 after starting the process, 113 tokens a second on the prompt. every request after it 140 about a fifth slower on the first request and then it's gone. A 24 GB Mac has plenty of room to keep a 19

7:24 gigabyte checkpoint in its file cache. So after one pass the drive barely gets touched. That only works if your machine has the room. Now take the room away. On the 12th, someone filed an issue running

7:35 the smaller tier on an 8 GB M1 that was already deep into swap. They measured about 0.7 tokens a second. The model card says 24 for that tier, so roughly 30 times slower. And the memory

7:48 footprint came out right at about 1 GB, which means the offloading itself was working exactly as designed. They ran a control 2, which is the part that makes the report worth trusting. A dense 4B

7:59 model through ordinary MLX, same machine, same session, gave 11.4 tokens a second. Nothing wrong with the laptop. The maintainer confirmed the cause in the thread. On a memory constrained

8:11 machine, decode is page faultbound rather than computebound. Their first suggestion was an environment variable that warms the page cache over the checkpoint before you start. And a

8:20 second thread adds another data point, a 16 GB MacBook Air streaming off an external USB drive about five tokens a second. So the requirement was never 3 GB of memory. It's 3 GB of allocator

8:33 plus fast storage plus enough spare RAM on your machine to cache a 20 GB file. That's the speed and where it comes from. Now the cost. The 4-bit base is frozen and a Laura adapter trained by

8:46 distillation from the full precision teacher sits on top recovering what the quantization lost. The adapters never get merged so one readonly base can serve several of them. They ran a

8:56 benchmark suite against the full precision base model with identical settings on both sides. 79.2 against 83.2 averaged call it four points. The sharpest single drop is on the AIME math

9:09 set. 86.6 against 92.7. So if you're leaning on it to reason, you'll feel that. If you're using it to write code, you mostly won't. Keep in mind what's packed into the one figure

9:20 you'll see quoted. 4bit weights, half the routed width, and a predicted router with the adapter pulling back as much as it can. Reading it as quantization loss underscells what they actually changed.

9:31 Two things worth saying plainly. They ran both sides of that table themselves, and I couldn't find anyone who's reproduced the throughput yet. It's Mac OS and Apple Silicon only. It serves one

9:42 request at a time, and the KV cache grows on top of that headline figure as your context gets longer. Here's where I land. The prediction head is real engineering doing more work than the

9:52 coverage gives it credit for because it doesn't hint at the routing, it decides it. 2.9 GB though is a true reading of the wrong thing. So, the number to check isn't 2.9. It's how much memory you've

10:04 got spare after everything else. Because that spare RAM is the cache doing the work. On a 24 gig Mac with a fast drive, this is a usable 35B. On 8 GB, it isn't. Subscribe for a tearown every week.

Frontier News · by Hyperjump Technology