Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
FreeToken lets you run DeepSeek V4 Flash on a single 3090 by offloading to system RAM, achieving around 10 tokens per second. It's a beta desktop app that makes large MoE models accessible on modest hardware, though performance depends heavily on RAM speed and capacity. The real story is that this changes the game for local AI by enabling models that previously required multiple GPUs.
Key points
- FreeToken is a desktop application that allows running large MoE models like DeepSeek V4 Flash on a single GPU (3090 or 4090) by offloading to system RAM.
- With a single 3090 and sufficient system RAM, DeepSeek V4 Flash achieved approximately 10-11 tokens per second, varying based on which experts are active.
- The presenter recommends at least 32 GB of system RAM, with 64 GB or more being preferable, and notes that DDR5 RAM can roughly double performance compared to DDR4.
- Qwen 3.8 27B BF16 failed to run in FreeToken, producing a non-descript error; the presenter could not determine the cause.
- GLM 5.2 NVFP4 was attempted but required 217 GB more system RAM than the test system had available.
- FreeToken is labeled as beta software; the presenter observed bugs in tokens-per-second estimation and limited model support.
- The presenter claims FreeToken 'changes the game' for local AI accessibility, especially for single-GPU owners, despite needing more polish and model expansion.
Tools mentioned
Techniques
- offloading model layers to system RAM
- mixture of experts (MoE) model architecture
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
We're going to be running free token today and this allows you to do things like run DeepSseek V4 flash off of 1390 and your system RAM. There is a desktop application that you can run. I'm going
to show you how you can run this also on COS. You may run into a couple of instances where it might be a little bit hard. There is a Windows distributable a Ubuntu also a common app image and the
Arch Linux package which we'll be running the Arch Linux Linux package later here. There is a huge dependency. A lot of people leave this out when they say things like run it on a desktop
system. Yes, one GPU is going to be able to perform insanely well for given what we're going to be running, but you have to have the RAM and it's going to be best when you're using those are mixture
of expert models and that's things like Deepseek V4 flash and many of the other frontal models like GLM 5.2 as well. So, this is one of those instances also where system bandwidth to your memory is
going to make a big difference. So, if you're looking at running 2400 speed versus 3200 speed on something like a Zen 3, you're going to get your full bandwidth off of the 3200 speed. You're
going to take a step down every time you go down a little bit on your RAM. But, of course, the cost of RAM is insane. So, check the links in the description below for some recommendations on RAM,
what GPU you would want to consider. A lot of stuff that we've got to check out here today. So, let's get this puppy up and running completely and loaded up on just one GPU. Luckily, this does load
substantially faster. And this should be just a few minutes to get this one loaded up. And if you keep your eye down here, you'll be able to see the actual utilization of RAM that's happening.
Okay. And when it says API server is ready, you're ready to go. I've got ours hooked up over here through Open Web UI. So, we'll be using this today for our chat. And I'm just going to toss it a
howdy really quick. And you'll start seeing some activity down here. And it'll actually give you your TG tokens per second that you're able to hit. But once it really gets churning here, we'll
be able to see kind of the actual tokens per second. And so it thought for 3 seconds and gave us that. So we'll actually toss a few back and forth here. So,
asking to tell me a little bit about the 1968 Camaro lineup. And again, this is running off of just one of the 3090s. So, you can see the other three sitting there not doing
anything at the moment. And this is what I was seeing last night. It comes in right around 10 to 11 tokens per second. This varies wildly. This is because the experts that are active and if you run
into new experts. So, seeing that we're hanging right around 10 still with this. So, yeah, there was the twodoor coupe and the convertible. Yeah, you can see the token generation is
while it's not super fast, maybe interactive speed. I don't think you're going to be running a Gentic stuff off of this, but if you're looking for a chat, this is actually at 10.5 or so
tokens per second. Incredibly decent. So, the performance on this something I would say if you're looking at a single 3090 can pack quite a punch. All right, we're going to run Free Token, the
desktop version. And yeah, so like I said, it actually is a pretty decent user interface. We'll just go with hugging face of course for our end point.
Hopefully it handles the VIN creation and everything perfectly. Fingers crossed on that. Looks like it gives us a little quick guide here. So, I downloaded two models
and I really think that we've got really a great one here, Deepseek V4 Flash, that we're going to check out also. And we've also got another one. If you go to the downloaded tabs, it'll show you
Quinn 3.8 27B BF16. So that would definitely not fit in just a single 30 uh 4090. So again, 4090 on this rig. We tested uh 3090 recently. But we'll start off with the Deepseek V4
Flash and that way we can get a good idea about what the performance difference between the Cache OS with 4090 versus Debian LXC13 with 3090 performance looks like. I'm
going to guess it's not going to be that different, but we'll find out here. And since this is running directly on the desktop, I can pop that over and we should be able to see the free token
desktop start to kick up here eventually. One thing you do want to keep in mind is if you have a bunch of other things like OBS running, uh, like for right now, I am actually not using
the Nvidia Inve might evict it. I'm not sure if it would fully stop it or not, but it could impact the amount of what's stored in there also, which would have a negative
impact on your performance. So, something to keep in mind if you're a desktop user and you're like doing this and starting up games and stuff, it might also have some impacts there. I
have this virtual machine now set to 192 because I changed my mind and I wanted to be able to run the DeepSeek V4 headto-head kind of a little bit uh just to see if there was a difference between
the 4090 and the 3090. So, I've got some bad news about 128 GB. I don't think that's going to be enough. But 156 probably would be 168 probably more common for you to run into. Again, it's
the same memory. It's just DDR4 2400 speed. So, of course, not the fastest DDR4 by any stretch of the imagination. All right, so it's now running. I guess we can just go over to chat.
It is up. And let's uh turn on thinking. We'll just turn it on max. Why not? If you're going to use Deep Seek V4 Flash Max, sure. See what those tokens per second look
like. Hopefully, it gives us that information. Oh, yeah. It will up here. So, that was 1.8 tokens per second that it said on that, but I think we need to run a few more to see if we can get a
little bit better information about that because uh there's a couple things that I would say it's a little bit buggy. Uh there are obviously not every model out there available for this. So there
probably it does appear and it is labeled as a beta. There probably are some things that may not be fixed up yet completely and tokens per second estimation may not be one of them. We'll
see if it can do this. Not for the quality of the story or anything just to see how many tokens per second it can generate. And recall we were at about 10.5 tokens
per second when we were doing this uh looking at this on the server side. So I I feel like this just kind of looks a little bit lower than that. We'll see what it comes out and says as
far as tokens per second. I I do like it in the Llama C++ interface that they've come up with in their UI. It gives a really very real time updating of the tokens per second for both prompt
processing and also for TG. So it would be nice if they did the same here. And it looks like it's using about 20.380 uh gigabytes of memory right now. So
that came in at 8.8 tokens per second. So the desktop client may be slower. Definitely does appear to be a very easy way to get up and running though. I can't tell you what the performance
impact of that would be on Windows, but I would expect it would probably be a little bit lower there also. So, let's go ahead and stop that. And we're going to as soon as that evicts out of
RAM here. Looks like it's getting kicked out. Kick up the Quinn 3.8 27B. Now, this one's a dense. So, I thought it was interesting that this one was a potential. This one I actually expect to
be even lower than 10 tokens per second. So I expect the performance to be not good because of the fact that oh what are we looking at here? The engine for Quinn 3.8 27B BF 27B BFF16 kept erroring
or exiting unexpectedly. Raw error below. You can restart it or shut it down. So that's a very non-escript error. I couldn't tell you. I guess I could copy the server logs and
open an issue. Might do that, might not. uh but that did not function. So that's one thing to keep in mind if you are up and running the Quinn 3.8 27B didn't run. But let me go ahead and just be
thorough about it and we will make sure that it's not. And we'll start up the free token desktop web application again and try to run it there. that way in case it was
parked or something like that and just there was for some reason not the ability to run it. Maybe it could have been that. We'll find out.
Looks like it closed out again. So, I think that's a pretty good review of what you could expect. And of course, a lot of people are going to be looking at this for the big things. If you've got
512 GB of RAM in your system, for instance, GLM 5.2 to NVFP4 may be a potential for you and for sure it will tell you insufficient RAM if you do not have enough utiliz utilizable RAM and
VRAM together. So you can see here I would need 217 GB more system RAM. So on the Epic that I still have over there, we can throw enough in there to be able to actually run that. So I might give
that a check out. Let me know in the comments below if you would find that interesting. And yeah, I think we're looking at what you could expect for, you know, a pretty innovative new piece
of software. And not a bad uh interface. Honestly, I do like this little uh this little uh thing they've got up here, which kind of compares your costs and stuff like that versus what you would
have paid if you went through the API for that specific model. And I think this is actually something you should definitely give a run, especially if you're a single GPU owner. And you've
got at least 32 GB of system memory. Probably 64 would be a much better place to be. 96 128 most of these are going to run. And the performance on things on DDR5 versus DDR4, I would expect that
you get probably a pretty easy doubling there. So I think that should be a really decent expectation if you happen to have the DDR5s in your system. pretty cool. I mean, yeah, like one GPU and
system memory running offload of uh DeepSeek V4 Flash, I can tell you would not be that fast for sure with just 1390 in any offload scenario that I can think of. Uh pretty cool. Like really pretty
cool. So, huge hat tip to the team. Huge hat tip to all of our members. Thank you for joining and everybody who likes, subscribes, shares. This is something that definitely makes local AI a lot
more approachable. This changes the game even, I think, is not too big of a statement. Does the interface need a lot more polish? Yeah. Does the amount of models that it's running need to
possibly be expanded? Sure. Did we hit a bug on Quinn 3.8 27B BFF16? An excellent model. My daily driver actually. Uh yeah. So, there are some issues outstanding. It is beta software, so
hopefully those can get resolved. So, there you have it. I think you should check it out. And if you're looking for more information on how to get up and running with software for your local AI
gear and setup, you can check out the playlist that I've got here. And if you're looking for more build guides and rig guides in general, you can check out the playlist that we've got here.