Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
Qwen 3.8 Flash Next running in a Hermes agent on a quad-3090 setup achieves 55-60 tokens per second with a Q4 quantized model, making local agentic AI both fast and practically useful for game generation and iterative coding tasks. The key is a custom vLLM configuration that trades peak speed for stability—disabling async scheduling loses ~10 tok/s but prevents crashes, and the setup requires 128 GB system RAM and 96 GB VRAM.
Key points
Qwen 3.8 Flash Next runs at 55-60 tok/s on four 3090s with a W4A16 (INT4) quantized model.
Disabling async scheduling in vLLM sacrifices ~10 tok/s but prevents crashes on this hardware.
The setup requires 128 GB system RAM and 96 GB VRAM (four 3090s) to load the model.
Vision support is included in the quantized model used (from Viney AI).
The agent generated a functional web-based abduction game in two prompt iterations.
Tools mentioned
Techniques
- Q4 quantization (INT4 W4A16) for vision models
- KV cache optimization with 131072 context length
- Prefix caching and tool calling parser (Qwen 3XML)
- Disabling async scheduling for stability on older GPU architectures (SM86)
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
Allan, grab the lard. We got ourselves a situation. Coin 3.8 flash next needed to be able to run for me in a Hermes agent kind of setup to be able to really produce things that I was interested in.
And I was able to get there. And I'm going to show you all the steps and I've got a recipe that you can copy built upon the backs of great other people also. But there are some twists to this.
So, make sure you pay attention. And we're going to also check out one of the outputs that I was able to generate with this in a couple of sessions. And this is a new game that we're working on.
Basically about abducting things from farms and aliens and stuff. So I think this is pretty fun. We're basically running everything two times as fast as we were before. So I'm going to take you
really quick here looking through the run block that we've got. And we're going to be using VLM for this today. But basically there are some things you definitely want to make sure you have.
131072 is our max model length. Our max number sequences is set to two. Our matchbox tokens is set to 2048. GPU memory utilization set to 0.96 quite aggressive here. DT type auto for our KV
cache. CUDAGraph mode full is set. And we've also got enable prefix caching. Enable auto tool choice. Our tool calling parser is Quinn 3XML. And importantly, no async scheduling is set
on this. Now, typically I don't run down what I call the cherry-picking route for setting up and running and displaying a bunch of information to you, the end user. And why do I not do that? Because
cherry-picking is something that is not a long lived thing. It is good today. It'll be mainline tomorrow, but in this instance, there may be some of these cherry pickicss that may not make it in
there. So, I wanted to be sure for especially the quad 3090 owners out there to give you a little bit of an idea and give you a playbook that you can use. You can follow along with that
on digital spaceport where I've got the entire post up here. We'll quickly jump over here. So I've got your requirements here. Now since this does have that very large engram and we are running VLM
basically safe tensors in an in4 variant which essentially that's a Q4 but it is going to work in things like VLM. Uh 128 GB of system RAM is pretty recommended here. You will see that we're going to
push right up under 100 gigabytes. That is kind of non-negotiable. So 128 gigabytes really is where you probably want to be. 96 gigabytes of course of VRAMm for the GPUs because we've got our
Quad3090s running on our rack here. Build video on this almost completed and I think it's one of our best produced videos yet. So, make sure you hit like, subscribe, and ring that bell. Huge hat
tip also to the people who helped get this out the door. And this was all spawned basically by a thread that I saw on X from Lokar and Lokar and he's a quad 390 owner now. He's an 83090 owner,
but I'm I'm super jealous over that. But he was able to hit 198 tokens per second. We're not going to go that route. You can go that route if you want to quant your KV cache. There is extra
steps to that. There is a lot of extra steps to that. And basically, I went through his repo. And his repo was, let's say, clawed it up. If you've read through Claude repos before in the past,
you probably know what I mean. Like, wow, it's got a lot of everything under the sun. It uses clawisms everywhere. And I really wanted one specific thing. The ability to run Quinn 3.8. 8 flash
next very fast and with acceptable performance for agentic uses. So, this is able to happen here. But I would also say this was built a lot of the work here that he pulled was built upon,
let's go back here, uh, Alexi Fetvs and that's Super Alashsh. Both of them, if you're on X, very good people to follow if you are a 3090 owner because they both have 3090s and they publish quite a
bit of stuff around them. But basically, I wanted to pick this apart and make it easy for you to use. Of course, you can point your agent to this and have it probably decompose it fairly decently,
but I wanted to give you a really quick oneliner that you can basically have your build script up and running if you have a base system up and running. If you note up here, I've got CUDA and NVCC
13.0 or better yet 13.3 plus your drivers installed on whatever system you want to run. If you've been following along with things, of course, Proxmox and Lexe is what we're going to be using
here today on our Prox 2 host. And you can see that we've got our VLM Quinn 3.8 8 running here. And our CPU usage is about 40% with 16 CPUs. And our memory usage is at about 103 GB right now of
192 GB that I've got aside. Like I said, this is going to 128 GB be your kind of hard limit that you're going to want to hit. Now, if you look over here, you can see that as we're loading up, our GPU is
parked right at about 22 each one of them gigabytes right now. And the host MIM is there. It is right there, 99.101 MIB. But definitely you need a tremendous amount of system resources.
For a average machine though, this is actually totally possible. Did you hear that? For an average machine, this is totally possible to do. So, we're going to fire up our Hermes agent over here
and it should be pulling it up here hopefully very soon. It looks like we have not fully loaded yet. On the right hand side, you should see that application ready in VLLM and we'll see
that in the upper right hand side. Lower right hand side here we've got our NVT top. So this provides us visualization kind of nice little bouncy charts of what each one of the four 3090 GPUs over
here is doing. This is a fantastic model. So a lot of people have been like 35B A3B 3.8 where is it? And also a lot of people have been like is there something faster than the 27B? I've been
like that. And this may well be the answer to that. And why? because it is a larger model first off, but it also has a very very uh just the performance of it does not degrade. It is a really good
model as far as the quality of outputs even at an in4 or a Q4. We saw that already. And it is capable, as you're about to see, of really decent agentic performance. So when you think about
like what can I actually accomplish with something, you want to be able to hook up your agent nowadays and not just be stuck in a chat interface. That is kind of the slow way old way of doing things.
Still not there 100% and it does take some time to get this fully loaded in. It is very interesting the way the new engrams load into your system RAM. So we're going to have it recall the
project that we had been working on. This is a little kind of 16-bit shooter game all web- based. And th this one's a little bit more complicated than something it remembered like some of the
earlier arcade games or possibly, you know, a flappy bird flippy bit kind of thing. So, I wanted to check out the progress it had made, see what it looked like. And yep, it did find it in the
abductum game folder and hopefully it can get that launched here and we can check out the progress that it had made on it. So, this is now two round iteration. The prompt for this is shared
on digitalspaceport.com and you can play along with this on arcade.digital digitalspaceport.com as well. All right, so you're basically out there trying to abduct as much as you can and avoiding
the things that will shoot you down. You can see the uh beam isn't quite lined up with the spacecraft, but the cow basically getting beamed up. That's that's what this is all about.
Oh yeah, and there's farmers that are mad and they'll chase after you, too. And there's crows that will attack you. And this all costs health damage. Oh,
okay. So, there you go. If you come up to a scarecrow. Hey, guess what? While I'm at the scarecrow, duck some things there. Oh my gosh. Yeah, this might be a little
bit too hard the way things are going right now. Oh, there we go. There's a cat. It was supposed to make a moon noise whenever it's getting abducted, but it
didn't. Oh my gosh. Okay, so we leveled up there. Oh my gosh. Yeah, the crows have murder eyes also. Oh, it's quite a hard game. Uh, so I
might need to have it kind of gradually ramp up, of course, as it's going through the levels, but pretty fun. Got this in one mega prompt and then one secondary prompt. You can find those,
like I mentioned, at digitalpaceport.com if you want to make a abduction game yourself. And this is just really cool because of a couple things. So, first off, I got to come back over here, and
what we're going to do is we're going to actually have it review and look through what it's doing. And I'm going to give it a couple of instructions like you'll kind of be able to see the tokens
per second that we're generating over here. And if you remember in Llama C++ we were getting not very good speeds like very much not good speeds. So the prompt processing that hit about 399
tokens up to about 516 tokens it looks like over there. And the TG is really impressive on this prune. And this is why I definitely have got to say this video needed to be made because there
will be faster and faster coming to you for this and that is going to definitely improve your experience if you wanted to do something like aentic work with this newer model. Of course, Quinn 4 is going
to be coming out very soon and you should also be checking out and hit that ring and like button so that you can get notified when we have more content on Deepseek V4 flash with vision which is
now out. However, the supporting bits to get that one up and running are not in yet all the way down the stack. So, probably gonna have to wait a little bit for that. Maybe check out the 3090 back
ports for VLM. That's most likely where uh we'll be heading for that. All right. So, I told it the progress is good, but we need to have a level of difficulty start off maybe 25% less hard and each
level gets slightly harder. It's kind of starting off a little bit too hard. This is much better, but the difficulty is pretty hard right off the bat. Also, review the loading screen animation. and
the beam and the UFO of the cow and the cow are offset from the craft. Don't forget to use your visual processing as well as your play rate to troubleshoot this. So you can see we're hitting right
about 59.4 TG on this and that's pretty good. However, you can actually go a little bit faster. There is one this is another reason why I usually don't do kind of the talking about these kind of
hatches because there's always things that will break. So, we have no async scheduling and that's taking about 10 tokens per second off right now. But with the patch set that's out there,
this right here, if it's not present, is going to cause problems and it'll cause you to crash. Does go faster. You might be able to get a pretty good run out of it before it crashes, but it definitely
will eventually crash out on you. So, definitely this will all create your entire run block and everything. That'll be servey- flash next.sh. It'll chod it for you. So then you can just basically
go in run the script with the commands that you've got here and be up and running with VLM provided you have 128 gigabytes of system RAM and at least 96 GB of VRAM. Now this is of course also
generated towards SM86. I think you can definitely upgrade this and get it to work with SM89 if you wanted to. You might need to make a couple of changes. SM86 that's the amper
generation. Definitely make sure your compression threshold is set. Just you can probably tweak this later up higher. I've got mine set at 0.5. I think I can probably push that to 065 pretty safely.
And I've got my contacts link set to 1310728K. And of course, you can copy along the prompt here and get that up and running. And a huge hat to to everybody who's
been working along on getting faster and faster performance for 3090s. These GPUs still have a lot of life left in them. And certainly the VLM back ports and the focus that they have on 3090s as well.
probably gives us a lot of longevity going forward even at high performance levels. Props to Viney AI who put out the actual quant that we're using the W4 A16 and this one does have vision
support in it which is an important thing. Again, vision support one of the reasons that definitely we needed to see DeepSeek V4 have that in their flash lineup and now it does. I I definitely
got to say though, this is this is exactly what we need to have happen to keep the longevity in the ecosystem for 3090s going and it doesn't look like it's stopping. As a matter of fact, it
looks like it's keeping pace. And in the case of 43090s, being able to fit the full 262 context window with a KV cache of 337K tokens and hitting 61 tokens a second at
260, that's the decay is insanely insanely impressive. Going faster, name of the game, and having better quality results also. That is why the vision is important. The INT4 definitely a great
option if you're a quad 3090 runner versus the Llama alternative out there right now. Hopefully we can see better and better performance out of llama. C++ for running things like better and
better prompt processing at deeper and deeper context so you can have a much better agent experience also. So that's pretty much it. Uh I think it's a fun little game and yeah, we're chugging
along here. It's recording 55 tokens a second that it's got right here. Making its patches, updating things. Okay, let's start abducting. I got four in one there.
Still not getting the cow noises. Avoid the birds. The birds are bad. The birds are fast. Also, the fast ones you cannot get away from mainly. Oh man. Come on. Come on.
All right, let's get to abducting chicken. can kind of split the difference and get two for one sometimes. The heart is not working. This should be
something that Yeah, the heart's not working. I ain't got no money, but I got a shotgun and a short temper. Get in your cow. Get in your cow. Allan, grab the lard. We got ourselves a
situation. It's an insane amount of people walking around out in a field in the middle of the night. This is just an abnormal of people out in a field walking around at night. What the hell
is that? Why can't I abduct it? Somebody get the sheriff. Somebody's slow. Is that Oh, went up a level. All right. I'm a man of the soil and you're flying
in a tin tin of trouble. Oh crap. Well, so that's the game. You can find that at arcade.digitalpaceport.com and it's up until whatever I create
next. So huge hat tip to my wife on this one also because she came up with this prompt and this idea and we've been working together on this and it's pretty fun. It's actually uh quite hard also.
So getting the balance of difficulty on it has been uh not easy to do, but you can check that out again at arcade.digitalspaceport.com. Of course, you can also grab your own
prompt, fire up your own llm and create your own abduction game. I I mean just impressive. Just impressive the scope of where we were a year ago versus where we are now. Now we are in a position where
we have actually seriously useful local AI. like not just you've got to become a magician of prompting to get something out of it. Literally throw it something pretty easy in a Hermes agent and the
the agent helps you align and helps the brain of the LLM align and get substantially better outcomes. We are at a really great point. We are at a really great point and I hope to read what you
guys have to say in the description below. Huge hat tip to all of our channel members, everybody who likes, subscribes, shares this out there. Huge hat tip to you guys also. all the
Patreons, all the buy me a coffees, everybody who supports this channel, thank you very much for everything you do. You guys are the reasons that I am able to be here doing this for you and
not have to be sponsor driven as far as the content we produce, which leads you to a different kind of outcome for a channel versus what we're doing here. Kind of fun, kind of exploratory, and
having a good time. Everybody, if you're interested in checking out more, you can always check along at with our very massive mega AI playlist here. And that has everything you could possibly ever
want under the sun for local AI running. And if you're interested in HomeLab, you can check out our Home Lab setup guides that we have over here.