Qwen 3.8 27B and Hermes Agent built a vLLM Monitoring App for Local AI

summarized

TLDR

Qwen 3.8 27B at FP16, despite being slower, outperforms Qwen 3.8 Flash Next at int4 for agentic code tasks because multiple slower agents working in parallel produce higher quality output. The presenter demonstrates a vLLM monitoring app built with Hermes agent, showing that with proper settings (max sequences 20, prefix caching, chunked prefill), a swarm of sub-agents can achieve 220+ decode tokens per second even on a dense 27B model. The real takeaway is that in agentic workflows model speed matters less than model quality, and Qwen 3.8 27B is the current sweet spot for local agentic AI.

Key points

Qwen 3.8 27B is run at FP16 on a quad RTX 3090 rig with specific vLLM settings.

Qwen 3.8 Flash Next at int4 fails in agentic code tasks but excels in deep creative chat.

The vLLM monitoring app tracks prefix hit cache rate (~94.8%) and token throughput across agent swarms.

The presenter uses Hermes agent to delegate tasks to multiple sub-agents (up to 24) that work in parallel.

Optimal max number of sequences is 20 for quad 3090s to balance workload and output quality.

Tools mentioned

Techniques

  • Agentic swarms with parallel sub-agent delegation
  • vLLM configuration: chunked prefill, prefix caching, max model length, max num sequences
  • Tool call parsing using Qwen coder (not XML)
  • Precise quantization: FP16, int4, KV cache at FP8 E4M3
  • Mamba cache mode set to line
Transcript (captions)

0:00 newest Quins, latest GLMs. There has been a lot of model drops, including that latest Deep Seek V4 Flash with Vision now built into it. And people probably are experiencing a little bit

0:10 of model fatigue. I know I am myself. So, I wanted to do a video today that's really quick and just show exactly what I'm running, exactly why I am back to running what I am running, and give you

0:20 kind of a demonstration of some of the quality of what I've been able to get out of what I'm running, and a comparison specifically between the Quinn 3.827B, 827B which I've put

0:29 through a fair amount of usage. We'll take a look at that here. And the Quinn 3.8 flash next which I was having to run at an int4 versus Quinn 3.827B which I'm running at an FP16.

0:42 So real quick, let's take a look at the VLM run block that I've got here. And I'll run through a little bit of this for you so you can see what the settings are that typically are going to impact

0:50 you. So GPU memory utilization I've got set to 0.95. I've got my max model length set to 18244. The max num sequences I have set to 20. And I'm going to show you today kind of a flex

1:01 of why it has not mattered to me. Whereas a lot of people say things like it's too slow to run the FP16 or this is you know too resource intensive to run a dense where that kind of falls apart

1:13 that logic when you're using a lot of agents and sub aents. The max number of batch tokens I've got at 16K. The KB cache I am at FP8 with the specific E4 M3 for our Amper 3090s. This is the Quad

1:28 3090 rig that you see here. Also, the 4090 is running my Cashio OS desktop which is what you're seeing here. And of course, I'm doing all this on this one machine recording this everything

1:38 playing video games all the same all all of that at the same time. One machine. We've got our enable chunk prefill. We've got enable prefix caching. critically important to have enabled

1:46 prefix caching by the way and chunked prefill if you want that fast reload. Our mm processor so multimedia is shim and we've got that set to 384 megabytes which is pretty heavy but I do a lot of

1:58 direct the sub agent to take a screenshot of what it's seeing and evaluate with the other agents so that you can iteratively improve. So if you're doing a lot of screenshotting 384

2:07 is probably where you would want to be. 256 can be a safe other alternative for you. Reasoning parser Quinn 3 enable auto tool choice Quinn 3 coder not XML is what I have for the tool call parser

2:18 take note of that if you're on the 3.8 8 branch the default chat template quarks. Okay, so we've got thinking true reasoning effort X high. Yes, it is a lot of extra tokens but the quality of

2:31 those tokens the outputs is substantially higher making it actually very worthwhile versus medium and preserve thinking I've got enabled is true also that one's just a good one on

2:42 any of the settings to enable. Mamba cache mode is set to line and I have disabled custom all reduce to get rid of the annoying message that you'll get in your warnings. So we'll go ahead and

2:51 start this up and I'll show you some of the stuff that we've been looking at and particularly where things started to fall apart at the quant that I was running in 4

3:01 basically. Uh that's very similar to quant 4 if you want to think of it in those terms as far as the bits per weight. But that was where things started to fall apart for me in agentic

3:11 use specifically when I was running Hermes agent in large processes for creating code creating projects. So I've spent time over the past couple of days here just playing with it and like I

3:25 can't produce as many videos as a result of that unless people want to sit around and watch me play with it which I I doubt that's entertaining. Let me know in the comments below if you actually

3:32 want to watch that. But it's very very you have to be hands-on with it for a lot of time to get an actual feel for whether the model is going to fit your workloads and what you want to do with

3:45 it. And while I would say I am for my Hermes agentic uses back to the Quinn 3.827B, if I wanted to have a really high quality chat that had some deep insights

3:55 and high creativity, I would actually switch back over to the Quinn 3.8 Flash Next for that. even at the IM4 for direct streamline communications that are back and forth that go really deep.

4:08 It is an insanely creative model. That's exactly why I would do it. If I was doing nothing but code work, I'd probably be looking at something different than the Quinn lineup.

4:17 Actually, I would probably be looking at let's just have Deepseek V4 Flash and that's probably going to be good enough for us. So, GLM 5.3 really good. A little too slow. Although, there is some

4:28 improvements out with Llama C++ that have sped things up. I read an unsloth tweet that you can expect up to 3x faster performance. So, I do need to revisit that. GLM 5.3 was excellent code

4:41 quality. Like, wo, code quality was just through the roof. But beneath that, for most people that are running moderate to higherend machines, you're going to be in the Deep Seek V4 Flash. That is a

4:53 very good coding model right there. So, we've got our setup up and running here. Let's get our Hermes model. And I just thought I would take you through the process today. It's been a while. So if

5:04 you needed to get connected to your VLM instance, you can follow along with what I'm going to show you here. Again, all of this running same machine. So we've got our VLM Quinn instance here. We've

5:13 got our Hermes Buds06. Budo 5 is retired at the moment, but it'll be coming back. And my desktop, which is the Cache OS 772. It's not a 7702. I just haven't renamed the container. It's the 3945WX

5:25 Thread Ripper instead of the Epic. Epic's still back there. We're going to be using that in some upcoming videos also. But yeah, different IP addresses for these. So if you go to summary, you

5:34 can see your IP addresses. So we've got 211 for the VLM instance and our Hermes BUZO is 58. Those are LXC containers. I was reminded the other day, people are still trying to do things like run GPU

5:47 servers in VMs. Don't do that. You do incursion performance hits. Run them in an LXC container. Performance hit is not existent. So custom endpoint here and we're going to do http col/192.168.1.211

6:04 and that is running on port 9876 and that is v1 for the endpoint. Our API key is nerdtastic and you can select one for auto discovery and it will detect the model

6:18 and we will just click yes on that and it will auto detect the context length also for this. This is the display that you will see if you're selecting models in the drop down from Hermes models. So

6:28 we'll just put Quinn 38 27B. That will be pretty specific for me. And I'll also put VLM. Okay. And so now we can just get our Hermes up and running.

6:42 And one of the first things that we'll do is we'll take a look at some of the statistics for usage. So you can get a little bit of an idea of how many tokens I've been using. Insights. This is it.

6:51 Yes. So insights will show you kind of what you've been up to. And so actually I'm going to go to scroll up here. And a lot of these are just kind of throwaway

7:02 chats that I've had. So a lot of active time though. So 5.9 days of active time. Average messages a session 30.6. Total number of uh tokens input. So 208 million output about 4.8 million. 213

7:20 million total. But if you see here, the Quinn 3.827B, I would say about 60,000 of that was last night, me refining something that didn't quite work out. And then you can see down here, Flash

7:32 Next W4 A16, which I put together the video and the guide on how you can get that up and running. Ran about 68 million tokens through that. Sub agents have been quite a bit of that uh

7:43 activity, quite quite a bit of that activity, about 92 million tokens there. And really when you're looking at what kind of information you can get, really really awesome. And so notable sessions,

7:54 longest session was uh two days. And yeah, that was that was an epic session right there. And I'm going to show you what we created during that session and where the difference between 3.8 27B was

8:04 and flash next was flash next could not get it across the finish line. 27B came in, finished everything, got it across the finish line. So, uh, let's go ahead and

8:17 check the VLM dashboard project and startup. But definitely 27B kind of this golden model and we've we've really gotten lucky that the Quinn team has released something this awesome. I do

8:32 believe the architecture that we saw though come out with the 3.8 a flash next is revolutionary as far as being able to offload to system memory and the the capabilities that'll bring to Quinn

8:44 4. Insane and I've got to show you. So AJ went on a little bit of a fishing expedition that leads me to believe we are definitely going to get Quinn 427B. Him just calling out the props that is

8:55 what very welld deserved. 3.5 was good. 3.6 was a huge step up. 3.8 is another step up. 3.8 ridiculously good. 3.6 was like revolutionary though. And everyone is better than the previous all in the

9:07 same year. But surely 2026 can't end here, right? Quindev. When are we going to get the next drop? We are all ready for Quinn427B. And that is at its me AJ AJ KV on X.

9:19 Really good follow. He also has a lot of stuff for smaller GPU people. So if you're a smaller GPU person, definitely somebody you would want to check out. And they say soon. Oh man. So I'm

9:28 excited. I think we're going to get that Quinn 427B. So definitely I would give that guy a follow. He's always pushing the right buttons and getting some responses. So, okay, let's see if it got

9:37 it up and running here. Our VLM monitoring app, which was developed with me and Quinn 3.827BFP16, a small local LLM. There's this uh narrative out there that you can't use

9:48 small LLMs for anything that's important or even slightly meaningful. And that's just totally not true. You can absolutely do very highlevel coding projects that are incredibly good, but

9:59 you do have to have that patience if you're running it locally because your ability to run things fast not super awesome sometimes. But this is going to I think what I'm about to show you

10:09 illustrate a little bit of why you don't have to have the fastest model if you are using a bunch of agents to do things. So we're going to have it spin up an agent swarm and do some research

10:20 for us. So having a box, let's call this the box, any box, but this box in particular for this uh mental model, having a box full of, let's say, geniuses, and you've got maybe 15, maybe

10:35 20 geniuses in the box, and you've given them a task, and they do work slower, but the quality of their output is higher. that box of geniuses is going to do better than a box full of

10:48 mid-managers or senior managers in many instances or senior engineers. Sometimes it just comes down to who's in the box. And this is where when you're working with agents, you really want to be

11:00 optimizing your tuning. Like you saw, we had maximum sequences set at 20 in this. So, we're going to see some agents spin up and we'll be able to watch the progress of this also happening over

11:12 here as we see our streams start to grow. And so, as soon as you see this preparing delegate task, so it's it's decided five is a a good number here. So, it's delegating five researchers out

11:23 there and it's going to send them out. I had this up to like 24 the other day. So, you'll see number of requests now is queued up there at four. And if we come over here, this will really start to

11:34 illustrate why having a tremendous amount of throughput in any one single speed agent. Grantite, while it can be slow for the chat that you have back and forth with one agent, if you're giving

11:44 tasks out agentically and walking away and not being super duper hands-on, well, things kind of work out better that way for you because you don't have to be this micromanager of what has been

11:57 in the past, I will say, a bleep show of going off the rails with your LLM. And this is why I actually like using Quinn 3.827B, 827B, a slower model even on the quad 390s at FP16 precision because it

12:13 does such a good job. And so we can see that our running requests are now at six here. Our prefill tokens are ticking around the 1900 range and our decode tokens are going to surge up there after

12:25 the prefill gets up there. So ingesting, taking in the information, and then the decode is going to kick in here is pretty freaking awesome. Our prefix hit cache rate is about 94.8. So pretty darn

12:37 good. And the completed request 352. This is going to be actually a pretty good way if you're running VLM to hook it up, monitor it, see what kind of work your workers are doing. This is really

12:49 important if you've got a bunch of agents. So you can be like, "Hey, by the way, agent that's managing other agents, go to the dashboard, screenshot it. If you see utilization below blank, kick up

13:01 more sub aents for the project and make sure that they're on task and working. And the decodes now in the 70s peaked out there it looked like and in the 60s. That's with four running sub aents at

13:11 the moment. Our decodes over a 100 now. That's Quinn 3.827B. This is the power of agentic swarms. This is the power of using agents. You have a bunch operating at the same time

13:24 working together on a project. It's effectively like growing the number of smart people in the box. So, we should see some pretty good numbers here. I've gotten this up to 350 on the decode. So,

13:35 we'll see if we can get it kind of close to that. I think that we'll probably end end up in the like 250s to 260s range most likely here. Oh, yeah. It's climbing.

13:48 133 on that one. And it looks like we hit 188 tokens a second on our swarm there. About 20 minutes into this pass six to eight, you're going to see that decay quite rapidly, but definitely you

14:01 still have quite a few working on the same task. So I would say 16 is a very safe number to be at for your maximum sequences. 20 gives you a little bit of wiggle room in case you accidentally go

14:12 over or there's a compaction going on or something like that. So if you do have quad 3090s and you're running say four static profiles for different departments that you would have in

14:21 Hermes agent, then each one of those can spin up a couple of maybe three, four sub agents without a big risk. If you go too heavy on that, of course it can blow up. If you're in the 20s, I've seen it

14:32 as high as 24, it still will work and you can still be at about 350, but the time that you will spend waiting for the outputs is significant. The outputs are actually really good, though. So maybe

14:43 it's worth that. Maybe it's not. It depends on the task like I said that you're doing. If you're doing say code, yeah, this is probably not a bad way to go having a bunch of sub aent

14:52 delegation. If you have a whole department of different things that are kind of multiaceted, say like design. So, it looks like we hit 220 there. 222. Yep. That's a pretty good amount of

15:05 decode tokens when you think about it. all that work going on in the back end, even including summarizations and compactions. That's a lot of workload. That's a lot

15:15 of workload. And this is where having a bunch of agents, even on a slow model like the dense 27B still just makes it completely viable to use for getting in results.

15:26 And so this is why in my opinion running a gentic work really does offset the fact that you could be running a slower model. If you're looking at something like text gen and prefill and one single

15:37 stream of conversation, of course you want to be very careful to pick a fast model in that instance. But if you're using a gentic work and you have a lot of workers working, especially because

15:47 of the detached nature that you really I mean it's feet up, it's not lean in, it really can be quite effective. And usually I don't go this heavy on any research. So this is an awful lot of

15:58 agents to be tasking for this. Usually about four agents would be totally capable of not only researching but producing a report or a summary for me that is incredibly high quality. A great

16:08 thing to run overnight and typical run times on those are about 30 to 35 minutes. I really do want to tip everybody that is a channel member. Thank you very much for all you do.

16:17 Also, huge shout out to everybody who likes, subscribes, all the buy me a coffees, all the Patreons, everybody who is just sharing this out there. Thank you guys for all you do for the channel.

16:26 And if you're looking to learn more about running your own local AI, you can find out more in our local AI mega playlist here. And if you're looking to learn more about getting up and running

16:35 with your own home lab equipment, check out our playlist that we've got here.

Frontier News · by Hyperjump Technology