Qwen 3.8 Flash Next + HERMES AGENT = AWESOME LOCAL AI AGENTS!

summarized

TLDR

Qwen 3.8 Flash Next running in a Hermes agent on a quad-3090 setup achieves 55-60 tokens per second with a Q4 quantized model, making local agentic AI both fast and practically useful for game generation and iterative coding tasks. The key is a custom vLLM configuration that trades peak speed for stability—disabling async scheduling loses ~10 tok/s but prevents crashes, and the setup requires 128 GB system RAM and 96 GB VRAM.

Key points

Qwen 3.8 Flash Next runs at 55-60 tok/s on four 3090s with a W4A16 (INT4) quantized model.

Disabling async scheduling in vLLM sacrifices ~10 tok/s but prevents crashes on this hardware.

The setup requires 128 GB system RAM and 96 GB VRAM (four 3090s) to load the model.

Vision support is included in the quantized model used (from Viney AI).

The agent generated a functional web-based abduction game in two prompt iterations.

Tools mentioned

Techniques

  • Q4 quantization (INT4 W4A16) for vision models
  • KV cache optimization with 131072 context length
  • Prefix caching and tool calling parser (Qwen 3XML)
  • Disabling async scheduling for stability on older GPU architectures (SM86)
Transcript (captions)

0:00 Allan, grab the lard. We got ourselves a situation. Coin 3.8 flash next needed to be able to run for me in a Hermes agent kind of setup to be able to really produce things that I was interested in.

0:11 And I was able to get there. And I'm going to show you all the steps and I've got a recipe that you can copy built upon the backs of great other people also. But there are some twists to this.

0:20 So, make sure you pay attention. And we're going to also check out one of the outputs that I was able to generate with this in a couple of sessions. And this is a new game that we're working on.

0:30 Basically about abducting things from farms and aliens and stuff. So I think this is pretty fun. We're basically running everything two times as fast as we were before. So I'm going to take you

0:40 really quick here looking through the run block that we've got. And we're going to be using VLM for this today. But basically there are some things you definitely want to make sure you have.

0:49 131072 is our max model length. Our max number sequences is set to two. Our matchbox tokens is set to 2048. GPU memory utilization set to 0.96 quite aggressive here. DT type auto for our KV

1:02 cache. CUDAGraph mode full is set. And we've also got enable prefix caching. Enable auto tool choice. Our tool calling parser is Quinn 3XML. And importantly, no async scheduling is set

1:14 on this. Now, typically I don't run down what I call the cherry-picking route for setting up and running and displaying a bunch of information to you, the end user. And why do I not do that? Because

1:27 cherry-picking is something that is not a long lived thing. It is good today. It'll be mainline tomorrow, but in this instance, there may be some of these cherry pickicss that may not make it in

1:37 there. So, I wanted to be sure for especially the quad 3090 owners out there to give you a little bit of an idea and give you a playbook that you can use. You can follow along with that

1:46 on digital spaceport where I've got the entire post up here. We'll quickly jump over here. So I've got your requirements here. Now since this does have that very large engram and we are running VLM

2:00 basically safe tensors in an in4 variant which essentially that's a Q4 but it is going to work in things like VLM. Uh 128 GB of system RAM is pretty recommended here. You will see that we're going to

2:15 push right up under 100 gigabytes. That is kind of non-negotiable. So 128 gigabytes really is where you probably want to be. 96 gigabytes of course of VRAMm for the GPUs because we've got our

2:27 Quad3090s running on our rack here. Build video on this almost completed and I think it's one of our best produced videos yet. So, make sure you hit like, subscribe, and ring that bell. Huge hat

2:36 tip also to the people who helped get this out the door. And this was all spawned basically by a thread that I saw on X from Lokar and Lokar and he's a quad 390 owner now. He's an 83090 owner,

2:49 but I'm I'm super jealous over that. But he was able to hit 198 tokens per second. We're not going to go that route. You can go that route if you want to quant your KV cache. There is extra

2:59 steps to that. There is a lot of extra steps to that. And basically, I went through his repo. And his repo was, let's say, clawed it up. If you've read through Claude repos before in the past,

3:09 you probably know what I mean. Like, wow, it's got a lot of everything under the sun. It uses clawisms everywhere. And I really wanted one specific thing. The ability to run Quinn 3.8. 8 flash

3:21 next very fast and with acceptable performance for agentic uses. So, this is able to happen here. But I would also say this was built a lot of the work here that he pulled was built upon,

3:34 let's go back here, uh, Alexi Fetvs and that's Super Alashsh. Both of them, if you're on X, very good people to follow if you are a 3090 owner because they both have 3090s and they publish quite a

3:45 bit of stuff around them. But basically, I wanted to pick this apart and make it easy for you to use. Of course, you can point your agent to this and have it probably decompose it fairly decently,

3:55 but I wanted to give you a really quick oneliner that you can basically have your build script up and running if you have a base system up and running. If you note up here, I've got CUDA and NVCC

4:05 13.0 or better yet 13.3 plus your drivers installed on whatever system you want to run. If you've been following along with things, of course, Proxmox and Lexe is what we're going to be using

4:15 here today on our Prox 2 host. And you can see that we've got our VLM Quinn 3.8 8 running here. And our CPU usage is about 40% with 16 CPUs. And our memory usage is at about 103 GB right now of

4:28 192 GB that I've got aside. Like I said, this is going to 128 GB be your kind of hard limit that you're going to want to hit. Now, if you look over here, you can see that as we're loading up, our GPU is

4:39 parked right at about 22 each one of them gigabytes right now. And the host MIM is there. It is right there, 99.101 MIB. But definitely you need a tremendous amount of system resources.

4:53 For a average machine though, this is actually totally possible. Did you hear that? For an average machine, this is totally possible to do. So, we're going to fire up our Hermes agent over here

5:05 and it should be pulling it up here hopefully very soon. It looks like we have not fully loaded yet. On the right hand side, you should see that application ready in VLLM and we'll see

5:15 that in the upper right hand side. Lower right hand side here we've got our NVT top. So this provides us visualization kind of nice little bouncy charts of what each one of the four 3090 GPUs over

5:26 here is doing. This is a fantastic model. So a lot of people have been like 35B A3B 3.8 where is it? And also a lot of people have been like is there something faster than the 27B? I've been

5:37 like that. And this may well be the answer to that. And why? because it is a larger model first off, but it also has a very very uh just the performance of it does not degrade. It is a really good

5:53 model as far as the quality of outputs even at an in4 or a Q4. We saw that already. And it is capable, as you're about to see, of really decent agentic performance. So when you think about

6:05 like what can I actually accomplish with something, you want to be able to hook up your agent nowadays and not just be stuck in a chat interface. That is kind of the slow way old way of doing things.

6:15 Still not there 100% and it does take some time to get this fully loaded in. It is very interesting the way the new engrams load into your system RAM. So we're going to have it recall the

6:26 project that we had been working on. This is a little kind of 16-bit shooter game all web- based. And th this one's a little bit more complicated than something it remembered like some of the

6:36 earlier arcade games or possibly, you know, a flappy bird flippy bit kind of thing. So, I wanted to check out the progress it had made, see what it looked like. And yep, it did find it in the

6:46 abductum game folder and hopefully it can get that launched here and we can check out the progress that it had made on it. So, this is now two round iteration. The prompt for this is shared

6:57 on digitalspaceport.com and you can play along with this on arcade.digital digitalspaceport.com as well. All right, so you're basically out there trying to abduct as much as you can and avoiding

7:11 the things that will shoot you down. You can see the uh beam isn't quite lined up with the spacecraft, but the cow basically getting beamed up. That's that's what this is all about.

7:20 Oh yeah, and there's farmers that are mad and they'll chase after you, too. And there's crows that will attack you. And this all costs health damage. Oh,

7:32 okay. So, there you go. If you come up to a scarecrow. Hey, guess what? While I'm at the scarecrow, duck some things there. Oh my gosh. Yeah, this might be a little

7:50 bit too hard the way things are going right now. Oh, there we go. There's a cat. It was supposed to make a moon noise whenever it's getting abducted, but it

8:11 didn't. Oh my gosh. Okay, so we leveled up there. Oh my gosh. Yeah, the crows have murder eyes also. Oh, it's quite a hard game. Uh, so I

8:37 might need to have it kind of gradually ramp up, of course, as it's going through the levels, but pretty fun. Got this in one mega prompt and then one secondary prompt. You can find those,

8:46 like I mentioned, at digitalpaceport.com if you want to make a abduction game yourself. And this is just really cool because of a couple things. So, first off, I got to come back over here, and

8:57 what we're going to do is we're going to actually have it review and look through what it's doing. And I'm going to give it a couple of instructions like you'll kind of be able to see the tokens

9:20 per second that we're generating over here. And if you remember in Llama C++ we were getting not very good speeds like very much not good speeds. So the prompt processing that hit about 399

9:32 tokens up to about 516 tokens it looks like over there. And the TG is really impressive on this prune. And this is why I definitely have got to say this video needed to be made because there

9:46 will be faster and faster coming to you for this and that is going to definitely improve your experience if you wanted to do something like aentic work with this newer model. Of course, Quinn 4 is going

9:59 to be coming out very soon and you should also be checking out and hit that ring and like button so that you can get notified when we have more content on Deepseek V4 flash with vision which is

10:09 now out. However, the supporting bits to get that one up and running are not in yet all the way down the stack. So, probably gonna have to wait a little bit for that. Maybe check out the 3090 back

10:18 ports for VLM. That's most likely where uh we'll be heading for that. All right. So, I told it the progress is good, but we need to have a level of difficulty start off maybe 25% less hard and each

10:28 level gets slightly harder. It's kind of starting off a little bit too hard. This is much better, but the difficulty is pretty hard right off the bat. Also, review the loading screen animation. and

10:36 the beam and the UFO of the cow and the cow are offset from the craft. Don't forget to use your visual processing as well as your play rate to troubleshoot this. So you can see we're hitting right

10:45 about 59.4 TG on this and that's pretty good. However, you can actually go a little bit faster. There is one this is another reason why I usually don't do kind of the talking about these kind of

10:58 hatches because there's always things that will break. So, we have no async scheduling and that's taking about 10 tokens per second off right now. But with the patch set that's out there,

11:09 this right here, if it's not present, is going to cause problems and it'll cause you to crash. Does go faster. You might be able to get a pretty good run out of it before it crashes, but it definitely

11:19 will eventually crash out on you. So, definitely this will all create your entire run block and everything. That'll be servey- flash next.sh. It'll chod it for you. So then you can just basically

11:31 go in run the script with the commands that you've got here and be up and running with VLM provided you have 128 gigabytes of system RAM and at least 96 GB of VRAM. Now this is of course also

11:44 generated towards SM86. I think you can definitely upgrade this and get it to work with SM89 if you wanted to. You might need to make a couple of changes. SM86 that's the amper

11:55 generation. Definitely make sure your compression threshold is set. Just you can probably tweak this later up higher. I've got mine set at 0.5. I think I can probably push that to 065 pretty safely.

12:06 And I've got my contacts link set to 1310728K. And of course, you can copy along the prompt here and get that up and running. And a huge hat to to everybody who's

12:17 been working along on getting faster and faster performance for 3090s. These GPUs still have a lot of life left in them. And certainly the VLM back ports and the focus that they have on 3090s as well.

12:28 probably gives us a lot of longevity going forward even at high performance levels. Props to Viney AI who put out the actual quant that we're using the W4 A16 and this one does have vision

12:41 support in it which is an important thing. Again, vision support one of the reasons that definitely we needed to see DeepSeek V4 have that in their flash lineup and now it does. I I definitely

12:53 got to say though, this is this is exactly what we need to have happen to keep the longevity in the ecosystem for 3090s going and it doesn't look like it's stopping. As a matter of fact, it

13:03 looks like it's keeping pace. And in the case of 43090s, being able to fit the full 262 context window with a KV cache of 337K tokens and hitting 61 tokens a second at

13:17 260, that's the decay is insanely insanely impressive. Going faster, name of the game, and having better quality results also. That is why the vision is important. The INT4 definitely a great

13:29 option if you're a quad 3090 runner versus the Llama alternative out there right now. Hopefully we can see better and better performance out of llama. C++ for running things like better and

13:39 better prompt processing at deeper and deeper context so you can have a much better agent experience also. So that's pretty much it. Uh I think it's a fun little game and yeah, we're chugging

13:49 along here. It's recording 55 tokens a second that it's got right here. Making its patches, updating things. Okay, let's start abducting. I got four in one there.

14:06 Still not getting the cow noises. Avoid the birds. The birds are bad. The birds are fast. Also, the fast ones you cannot get away from mainly. Oh man. Come on. Come on.

14:23 All right, let's get to abducting chicken. can kind of split the difference and get two for one sometimes. The heart is not working. This should be

14:41 something that Yeah, the heart's not working. I ain't got no money, but I got a shotgun and a short temper. Get in your cow. Get in your cow. Allan, grab the lard. We got ourselves a

14:57 situation. It's an insane amount of people walking around out in a field in the middle of the night. This is just an abnormal of people out in a field walking around at night. What the hell

15:09 is that? Why can't I abduct it? Somebody get the sheriff. Somebody's slow. Is that Oh, went up a level. All right. I'm a man of the soil and you're flying

15:28 in a tin tin of trouble. Oh crap. Well, so that's the game. You can find that at arcade.digitalpaceport.com and it's up until whatever I create

15:44 next. So huge hat tip to my wife on this one also because she came up with this prompt and this idea and we've been working together on this and it's pretty fun. It's actually uh quite hard also.

15:53 So getting the balance of difficulty on it has been uh not easy to do, but you can check that out again at arcade.digitalspaceport.com. Of course, you can also grab your own

16:05 prompt, fire up your own llm and create your own abduction game. I I mean just impressive. Just impressive the scope of where we were a year ago versus where we are now. Now we are in a position where

16:16 we have actually seriously useful local AI. like not just you've got to become a magician of prompting to get something out of it. Literally throw it something pretty easy in a Hermes agent and the

16:29 the agent helps you align and helps the brain of the LLM align and get substantially better outcomes. We are at a really great point. We are at a really great point and I hope to read what you

16:41 guys have to say in the description below. Huge hat tip to all of our channel members, everybody who likes, subscribes, shares this out there. Huge hat tip to you guys also. all the

16:48 Patreons, all the buy me a coffees, everybody who supports this channel, thank you very much for everything you do. You guys are the reasons that I am able to be here doing this for you and

16:58 not have to be sponsor driven as far as the content we produce, which leads you to a different kind of outcome for a channel versus what we're doing here. Kind of fun, kind of exploratory, and

17:07 having a good time. Everybody, if you're interested in checking out more, you can always check along at with our very massive mega AI playlist here. And that has everything you could possibly ever

17:18 want under the sun for local AI running. And if you're interested in HomeLab, you can check out our Home Lab setup guides that we have over here.

Frontier News · by Hyperjump Technology