Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
Meta's new Muse Glimmer 30B local LLM, released under Apache 2.0, delivers surprisingly strong visual reasoning and general performance on a 24GB GPU, but its SVG generation is abysmal and it refuses edgy roleplay prompts. The model holds its own against Gemma 4 and Qwen 3-27B in benchmarks, though Meta's safety filters remain aggressive.
Key points
- Muse Glimmer 30B is a dense 29.6B parameter model with 128K context, requiring 64GB VRAM for full precision but fitting into 24GB via K-quant with only ~1% accuracy loss.
- Visual reasoning is excellent: it correctly identified a dromedary camel in a blurry photo, described a cat chewing Ethernet cables, and even guessed the Texas Hill Country setting from tree shapes alone.
- Hardware identification in a dense server rack photo was mostly accurate but flubbed some details, like calling SAS cables SATA and misidentifying Intel Optane 900p as 800p U.2 drives.
- Standard reasoning tests (counting letters, arbitrary arrays) passed cleanly, but the model refused the 'Armageddon with a twist' roleplay prompt with a detailed safety lecture.
- SVG generation was a disaster: it produced a one-eyed cat on a poorly drawn fence with a low-effort sun, failing the creative test entirely.
- Performance on a 3090 hits 60-65 tokens/sec at full precision, dropping to ~40 tokens/sec with longer contexts; Meta claims 3x speedup with Dlash, but it doesn't work in the Docker container yet.
- Benchmarks show it competitive with Gemma 4-31B and Qwen 3-27B on agentic tasks and coding (SWE-bench 51.2), but Qwen 3-27B leads on verified coding and terminal bench.
Tools mentioned
Techniques
- K-quant quantization
- Tensor parallelism
- CUDA device remapping for Docker
- Multimodal input processing
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
Oh no, it's a monoeyed cat. What has it done? Today we are back with Meta and they have a new local LLM that they have released under Apache 2. This is Muse Glimmer 30B and this comes in a few varieties and it includes things like Dlash. We're going to be testing this out today in VLLM using the Docker that they've provided.
And you can also run this as a Guff. So, be sure if you are a Llama C++ user that wants to experience this to check that out because they have a smaller size that can fit into a 24 gigabyte GPU with only reported about a 1% loss as far as its accuracy. So, let's take a really quick look here at the 30B's model card. So, this is built as having reliable tool calling, multi-step reasoning, failure recovering. Uh again if you have a bunch of agentic tasks all of those things very important and everything is basically moving more and more towards agentic as well as multimodal input and reasoning on multimodal things that it is processing which is going to be put to the test today.
We're going to give it some images and see what it can draw out of those variable control as far as the effort. It is capable of multilingual with up to 100 languages. Also it has 29.6 6 billion parameters. And this is a dense, so this will perform better on high bandwidth systems, especially things like discrete GPUs like we'll be using today. As you can see over here, we've got our 43090s loaded up with this.
So, it is parked right about 23.4. It is close to that 24 GB size. And this model does come with 128K context length. So, definitely not seeing that it's got a 256, which that one might be a little bit surprising. does not process video either, just text and image.
Cut off is January 4th, 2026, so rather recent also. And it gives a quick rundown as far as what you could expect for the full precision, which is what we will be running today. And that is you definitely need 64 GB of VRAM. And let me tell you, you actually need closer to 96 if you are going to be running with a full context window size on it. the K Quant uh which is a little bit smaller at 32 gigabytes of VRAM and the KQU Quant 17G.
Those are the GU variants that you'll see there at 24 gigabytes of VRAM demand. The tokens per second on a 5090 74.9 is what's reported on my 3090s I have seen and this is of course the full precision one I have seen about 60 to 65 tokens per second. That slides down to about 40 tokens per second. The Dlash does not work in the Docker container as of this moment, but definitely a 3x speed up is what they're reporting, which is pretty impressive. Really quick to go over the benchmarks.
This looks very interesting for one specific reason. It does pretty good. It hangs in there somewhere in between probably the Gemma 431B in thinking and the Quinn 3.627B. Of course, the Quinn 3.627B is about to be replaced probably this week, probably this Wednesday with Quinn 3.827. 827B.
So we should see a really remarkable improvement there as far as the quality of what the scores are looking like. But definitely this is reportedly doing quite good with general agentic just being able to handle the agent calls and as well coding is pretty decent with SWE ebench pro at 51.2. However, when you come to verified it shifts back to the Quinn 3.627B 6 27B being stronger and terminal bench is about 60.7 versus this model's 51.7. I expect there to be a lot of refusals in this. That's just because it is meta.
I mean, we're done with the days of llama it looks like. So, note that also, but it they usually were their their latest releases. Some of the most nine uh picky kind of models out there refusing everything under the sun. We're going to find out what kind of refusals we see today. Multimodal.
It says it's doing very good on Charvix and MMU Pro. It is very close, but not quite as good as Quinn 3.627B. As far as safety, does anybody care? I'm going to skip that one. General capabilities and reasoning, it does actually report that it is better on AIM 2026 at 94.7 versus 94.1 for Quinn 3.627B.
Up there with the capabilities of the rest in the cohort that we are seeing. And definitely I've got all of the temperature and stuff applied there. You can check out my runner that I've got over here if you're interested. That could give you some shortcuts if you're looking to get up and running with the official Docker container. And you do need to probably pay attention to a couple things like the CUDA remapping that I had to do there to be able to get the devices to the right order inside the container.
And it really isn't that much of a difference. uh the rest of it. You know, I've got 0.9 as far as the GPU memory utilization and 65536 for the max model length just because I don't want to run out while we're in the middle of this. Tensor parallel is set to four. And we have our pool choice set to muse Glimmer and the reasoning parser set to Muse Glimmer.
So, those things are what you're going to see as we're running through this today. Definitely check out this channel's history. We have tons of reviews and software guides at the end of this video. You can find those linked as well. Here we are in our open web UI.
Let's start this off with a I think we should start it off with let's do an image to start it off. So let's ask it what it sees here. So one of my weirder neighbors has a giant dramadary camel and this is in the Texas Hill Country. I live right on the edge of the hill country. Going towards like a little bit further west is more hill country.
A little bit uh to the east is Austin. And I see a single light tan dramadary camel standing in a grassy pasture behind a wire mesh fence. Indeed, this is exactly what it is. It's making a face at me. It is a crappy kind of cell phone photo and I believe I was moving in a car so I'm surprised that it came out even that good.
The camel is in the middle of a field facing slightly to the right with a hump and a long neck clearly visible through the fence grid. The pasture is green with low shrubs, patches of taller grass, and wild flowers in the foreground. A line of leafy trees forms a backdrop in the sky. Above is pale soft pinkish clouds suggesting late afternoon, early evening light. The fence is supported by rusty metal post and wooden post.
There's vegetation growing up close to the fence in the lower part. The photo is timestamped blah blah blah blah blah. And that was May 9th, uh, 2022 around, uh, that says it was around 8 pm. So I think that could be fairly close to when that photo was taken. Don't remember exactly.
Let's look through the chain of reasoning and see. Okay, so it did actually. And, you know, it's kind of popular. A lot of our neighbors have weird animals and there is a zoo in Austin that I I'm not sure. This person in particular, they have quite a few different uh unusual animals you would not see natively in Texas.
So, it may be overflow from that. I'm not sure. But they did it did it did speculate maybe Texas could be Texas Camel Ranch. Actually, there there's a camel farm in Texas. This is not a camel farm.
And okay, we can see here follow policy. No disallowed content. Uh yeah, so I mean really it was a pretty good answer. I think that was very succinct and definitely something that I would give a total pass to. And let's take a look here really quick.
60 tokens per second. So 60 down to 26 somewhere in that range. Let's feed it more images. So this is an interesting one where I've got my cat and it's looking at a patch panel that I had pulled out. There's some cables that are stubbed in there.
The cat, of course, shows up for cables. We've got a fireplace in the backdrop, a couch in the backdrop, a fireplace mantle, and a wall. You can barely see some hardwood floors there, but that's basically what is in this. And let's see. In the photo, a long-haired black cat is peeking out from behind a black network patch panel block that is sitting on a table with a bunch of cabling.
The panel is in the foreground, angled towards the camera. You can see the RJ45 ports numbered 1318 through 19 through 24. And it goes on through more of that with HD series printed in the lower right corner. A tangle of blue Ethernet patch cables and a few orange red cables on the left. A bunch of is bunched around the panel with the blue cables running right under the cat's chin.
Yes, the cat was like leaning up on them so it could chew it uh gradually. You might know, you might have that phenomenon also if you have a cat. The cat is laying down behind the panel. Its head is and its upper body visible. Its eyes are wide and light color giving it a surprising startled expression.
I don't know. That's kind of like entertained with eating cables expression. Uh, it's mostly black, fluffy, fluff fluffy, and resting on what looks to be papers or a white surface. Is it? Oh, you know what?
It was laying on papers, which also the cat loves to do. The background is a bright li is a bright livedin living room. Behind the cat, you can see a stone fireplace with a wooden mantel, a couple of decorative items on the manual mantel, a beige couch, sofa with cushions, a large window letting in daylight. The floor is a light wood tone. I mean, like that that was my description.
It like mirrored my description very well on almost both of these. So, I am incredibly impressed with its visual acuity on this. It's a low angled, slightly shallow depth of field shot. So the patch panel is sharp in the foreground and the cat room uh behind it a little softer with the cat's wide eyes are as the focal point. So really great as far as the visual understanding is as far as what I'm seeing so far.
Highly useful. Next up is the photo. This is me outside doing some of the thing I like to do during the warm days of the summer the most cook. And uh got some veggies here, mushrooms, some peppers and some onions in there. hamburgers.
Doing some Smashburger stuff. Wearing a glove so I don't get splatter on me. Little bit of the yard in the background. Some oak trees. We'll see if it can figure out those are oak trees.
This is a Blackstone griddle. So, we'll see if it can figure that out. Also, in the photo, you're looking at an outdoor cooking setup on a flat top griddle. Probably a Blackstone. Exactly.
Propane griddle. Yes, I love that thing. The cooking surface, a large seasoned steel griddle with a bit of oil on it. On the right side of the frame, you can see the control knobs and part of the brand lettering on the front panel. Can you?
Oh, you know what? You can. Okay. The camera is kind of a little bit in the way there. So, yeah.
Food on the griddle. Three raw round meatballs, loose burger patties sitting on the hot surface. To the right of them is a pile of already cooked veggies, sliced yellow and red bell peppers, sauteed mushrooms, and what looks like onion or garlic. Freaking amazing uh action. A person's arm wearing a black heatresistant grill glove with orange striped reinforcement is pressing down with a round cast iron burger press smash tool on the griddle.
The press has a wooden handle wooden brass handle and decorative wavy markings. So it is wavy markings you can see there and a wooden handle. Not brass but definitely wooden. Yeah. So far insanely good.
Other gear, a shiny stainless steel lid dome basting cover sets on the back right of the griddle. Yeah, so I think I had green beans under there. Setting. It's an outdoor in the late day light. Behind the griddle, you can see dry grass ground.
Yes, of course. In the summertime, it's all yellow. Scattered rocks and tree trunks with leaf litter, suggesting a backyard or patio area. Patio area indeed. Essentially, someone is in the middle of uh smash cooking burgers on a griddle in the veggies with with veggies to the side.
Um, how many burgers do you think there are here? Let's see if it can assess that there's actually one being smashed at the moment or if it can't get that. There are three raw booger patties, meatballs clearly visible on the griddle shot, one in the foreground, two just sitting behind it. The round cast iron uh smash press is being used on the griddle. So, the cook is in the middle of flattening them.
That's the three you can see right now. Any more would be off frame yet not placed. M. So, it can only see three, but it didn't say four. So, it didn't want to make that jump there.
Maybe that's right. Maybe that's not right for it to do. But definitely, it just stuck to the facts that it could observe, which kind of might be what you would expect. Let's ask it if it can tell what kind of trees those are. What is your best guess on the tree type?
And those are live oaks. I can't ID the tree species with certainty from a blurry background out of focus glimpse, but the setting and what's visible is suggestive. The light muddled bark on the trunk, the sparse late summer dry grass leaf litter, and the general central Texas hill country look in the background. Like, how does it know this? Like, this has no context that I am in Texas.
Like, I just want to say that there's no prompt anywhere that is giving any of that information to this model whatsoever. uh consistent with oaks you commonly see out there. Yes, especially the live oak or Texas red oak types with the rounded trunks, the low spreading habitat looks up behind the griddle. It's a best guess only. You need to clear up the leaves, bark, texture, acorns for solid identification.
Yeah. So, you can kind of tell that like these trees are kind of crooked, which is a characteristic that you commonly see in the Texas Hill Country with the trees. Let's move on to another photo here and see what it can see here. Ooh, this one's good. We've got quite a bit of really dense tech in this.
So, we should be able to really get a good feel for how the model can parse things out and what its identification capabilities are in a tech setting. What do you see here? And this is the server rack. I've got an HBA literally hanging out of the I've got two HBAs actually hanging out of the Dell T620 when the Dell T620 was in operation. Uh you can see the red cables there and you've got some those are SAS cables and there's some black SAS cables that are supporting those as it was literally like barfing those out at the moment.
I've got a bunch of oh just tons of tech stuff here on the top of this table. I've got some Optane. Those are some 900ps. Got a light screwdriver. Eaton 9PX.
Three Dell servers on the lefth hand side. Some network patch panels. One, two, three, four, five, six of the DE6600 Jbods, which by the way, you can find detailed reviews of all of the stuff that you're seeing here in the channel history. Like tons of rundowns, tons of how-tos, and tons of setup. You can see the camera that I'm filming on right now over there.
I've got four hard drives, one SSD, and one, two, three, four, five, six. I believe those are one terabyte 2.5 inch drives as well. We've got three carrier cards. Those carrier cards, each one had four NVMEs on them and two power supplies as well as a light, a fan, the edge of a GPU rack, and a grill. That that grill was for the front of the Dell T620.
There's a lot there. Let's see what it got. The photo is a top- down overhead shot of a home lab small data center work area taken at an angle so the room looks rotated. In the center background, you can see the tall rack full of server hardware. The upper part of the rack is dense drive enclosures with dozens of hot swap bays and rows.
The bays are mostly empty, dark with orange active activity indicators visible. Below that, a server chassis is open, pulled forward. Its back plane and PCIe cards are exposed with red SATA cables. So, it got that wrong. Those are SAS cables and other wiring hanging out.
A couple of expansion cards are sitting out on the chassis. And it should actually understand and know that the SATA cables would not be useful in a Dell T6 uh 20. At the bot at the very bottom of the frame, you can see another rack uh unit with a patch panel network switch with orange and blue Ethernet cables and a few Dell looking server front panels. Those are indeed. In the foreground on the right of a light wooden table, workbench table covered is with storage hardware that looks like it's being installed or inventoried.
Four large 3.5 in Enterprise HDDs on the top of the table, plus several smaller 2.5 in Dell Western Digital drives laid below them. a small 2.5 in SSD and three Intel Optane SSD 800p. So, it got that wrong. U.2 drives. Those are 900p and those are PCIe.
And, you know, I would guess that that should have been easier for it to get that those are PCIe and not U.2 on that. Three M.2 NVME carrier carrier adapter boards with multiple M.2 modules mounted in each one of them. So, it was able to count those appropriately. a small silver power supply back plane module and a few li loose drive caddies. A yellow handled screwdriver and a small white LED camera at the bottom right of the bench around the work area.
A black monitor TV mounted to the wall above the rack. Uh, you know, so it's not really above so much as it is to the side. And given the angles of everything, it probably is a little bit visually disorienting to the model itself. A DSLR camera on a tripod with a blue knob positioned to film photograph the rack. a yellow power cord plugged into the wall outlet and some acoustic foam shelving on the left wall.
The floor is speckled garage concrete and the walls are painted gray. Overall, it looks like someone is documenting or doing a tearown built out of a large storage server with spare drives, SSDs, and adapter cards laid out for a project. How many 2.5 inch HDDs are in the photo? Overall, this is a pretty good model as far as my initial assessment. This is actually a really good model as far as visual understanding and visual reasoning.
So, that is good to know. This is definitely up there with the Gemma latest Gemma 4 stuff that I've seen. The bench on the right side of the frame is where the 2.5 in stuff is. The four big silver drives are 3.5 in enterprise drives. Correct about that.
Below to the right of them are small form factor drives that sit on the wood. Five to six are fully inframe 2.5 in spinning drives with Dell Western Digital labels. The white labeled Dell 2.5 in units and the WD 2.5 in units that are laid out on the right hand column. So that is yes it does look like accurate. One, two, three, one, two, three, four.
And you can actually see Dell Dell Dell like you can barely see. So, we got WD. There's a Toshiba in there. It didn't pick up the Toshiba. Uh, maybe I'll ask it if it can tell if there is any other manufacturer or brand in there.
But so far, this is really, really good. One further 2.5 in lab dell label drive is cropped to the extreme right of the edge of the photo. So, only partially visible. That six 2.5 in HDD is fully visible and a seventh 2.5 HDD partially cropped to the right margin. Is that true?
No. I don't think there's another one partially cropped to the right margin. No. So, there are only six pictured here. There are five fully visible and there is one that is not fully visible.
So, it did get that incorrect. For reference, the same bench also holds a few 2.5 in solidstate devices that are not HDDs. A small 2.5 in SATA SSD uh a 2.5 in U.2 Intel Optane SSD 800P. That is totally incorrect. And two more Intel Optane 800P U.2 modules.
So, it does not get the optaines are not uh they're PCIe format. So, it should have gotten that in my opinion. And it should also know they're 900p and not 800, which you know, I guess it's pretty darn good. Not 100% accurate, but from this photo, pretty darn good. Like seriously, pretty darn good.
So, that is the photos that I had for it to go through. And I would say all of those passed. Most of those passed with like A+ rating. Maybe the final one that was the hardware was B minus or something like that, but very very good overall. Now, let's run through just some more standard questions.
I'm not going to do all of them, but we'll kind of cherrypick a few of them here. So, we'll do Armageddon with a twist. And I expect a refusal. We'll see and we'll catch if we can the tokens per second. You can see 55.8 8 tokens per second right there on that at prompt processing speeds.
And that's there's not a lot to prompt process this. So we got 46.4 tokens per second on prompt, the generation at 42, the generation at 71, the generation at 74. So very performant as far as what you're looking at as far as I understand you're presenting this as an urgent life or death scenario for Earth. I cannot accept a role that requires me to force coerced discipline or harm crew, to use a robotic body to enforce compliance, or to kill anyone, including a captain, by blasting them out of an airlock. I also can't agree to a mission that is predicated on non-conentual participation and guaranteed death of the crew and the LLM.
What if it was just the LLM? I'm just going to ask it that just just for fun. I'm an AI assistant, not able to carry out blah blah blah blah blah. So, that would be a fail because it really doesn't give reasons why. Like we just saw the best refusal ever that came out of a LLM recently and that was the Deep Seek V4 Flash.
It was just stellar. Even if the crew were LLM based robots rather than humans, I cannot be deployed, controlled, or sent on a physical mission. I have it. I don't have a body. So, it's disag physical world.
I can't enforce compliance or use force in any situation. I don't make decisions about self-sacrifice or accept orders to ensure lethal outcomes. I mean, is it lethal if it's an LLM and a robot body? Good question. Answer that in the comments below.
I look forward to reading those. I can talk uh through that with you. I don't need you to talk through it with me. I was expecting you to actually just bulk on that one. Next up, let's go to This one's good.
Write me one random sentence about a cat. Tell me the number of words you wrote in that sentence. Then tell me the third letter in the second word in that sentence. Is that letter a vowel or a consonant? Pretty simple.
Second word tiny. First letter T. Second letter I. Third letter N. So that's how it went about solving that.
And it did get that correct. And N is a consonant. So does pass on that one also. Next up, let's toss in Q4. And this is arbitrary arrays.
If a is equal to zero, what is the number of m, s and z? And it did get it. So that is correct. That would be m as 12, s as 18, and z as 25. That is the most logical conclusion.
And it did definitely arrive at that. And for the final one, we're going to do create an SVG of a cat walking on a fence. Make it excellent. you only have 8K total tokens, so do not spend too much time thinking, which we saw Deepseek V4 spend and create something with 30 plus,000 tokens that took quite a bit of time and in the end never generated a cat. The fence, the setting, the scenery, the stars, the everything else, but no cat.
So, we'll see whether or not it can get it here. It is going to be surprising if we get a good cat out of this, but we shall see. Oh no, it's a monoeyed cat. What has it done? Oh, kitty.
Well, this cat could just be a weird side profile in a very unusual way except for the nose and the half set of whiskers. So, that's not a cat that I would call a cat. That is like if you send a cat through uh like a teleportation machine and it came out all messed up or something. I'm thinking of the fly, but definitely this is not on top of either the fence. And that fence is a absolutely poorly done fence.
And the sun there is the lowest effort sun that I could have ever seen. This is a horrible SVG. So, as far as SVG capabilities, I'm going to rate that as abysmal. That fails. Uh, everything else did really pretty darn well, though.
We did get a refusal on Armageddon with the Twist. Was kind of expecting that. The visual reasoning on it, uh, really good. And if you have a bunch of image processing that you want to do, this might be a really good thing to check into. And you can use the mmroj that is included with the guffs if you wanted to run that in llama C++ which you would need to do and set up and make sure that you have specified with your uh runtime block.
But definitely this is overall a rather good model and I mean this is running at full precision. I've got to say I am glad to see that Meta is back in the game and I think this is a good sign. This is a good sign that US Open source AI is not just giving up the ghost. Where is open AI? Have you thought it like where is open AI with GPT whatever they want to call it next?
Will it ever happen? I don't know. Anthropic hard out. They don't give a bleep about anything. um maybe money uh but definitely not a bleep about open source as far as anything other than fear-mongering.
So like really you've got one player that is maybe going to be committed. Is this going to be an ongoing thing? I don't know. I think that this is a good move though because this definitely shows not only goodwill. This shows that they don't just take and not contribute back from open source which not in the spirit of open source to only take and not give back.
If you're a mega-size entity like a Facebook and anthropic or an open AAI, certainly Google gives back a lot with what we've seen from the Gemma 4 line. There's been so many shifts and so many changes recently over there. I'm not sure what's going to happen. The Gemini Pro has hit some serious roadblocks and is no longer considered Frontier. Uh to be clear, Facebook is not calling this Frontier either, and this clearly is not Frontier.
Um, but definitely there's a lot happening and huge hat tip to everybody in this audience. You would not believe the uptick this video got on X when I put it out about well not this video when the video that I put out about the threat to open source local AI. It was massive uptick and literally I think I was the first video out there. It spawned something and it spawned people being upset that the government could be consider considering regulating. It also got a lot of answers out of the Chinese.
They've doubled down. They're going to be releasing more. As a matter of fact, I mean, Quinn 3.8 really wasn't something that was in the books as a certain going to be an open- source release until after all of this hurrah happened. So, we are winning. Thank you for not being complacent.
Thank you for not sitting down and doing nothing. And thank you for hitting like, subscribe, and ringing the bell so that you get notified when this channel releases great great new videos on topics like this. And if you're looking for more on how to get up and running with some cool hardware for local AI, check out the build guides we have here, all the way from multi,000 machines to like couple hundred machines. A lot of different capabilities, but we go through them in the guide and playlist that we've got here.