Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
Meta is back in the open-weight game with Muse Glimmer 30B, a model that runs on consumer hardware and shows only 0.2% degradation on agentic tasks at 4-bit precision — making it a practical, accessible alternative to subscription-based AI.
Key points
- Muse Glimmer 30B is Meta's first open-weight model in years, signaling a renewed commitment to open-source AI from a major US lab.
- The model runs on consumer-grade hardware, tested at 4-bit precision to reflect real-world usage rather than requiring expensive setups like a DGX Spark.
- Agentic task performance degradation at 4-bit is only 0.2%, making the trade-off for accessibility nearly negligible.
- The model demonstrated practical capabilities like generating SVG assets and building a simple playable game from a reference photo.
- Coding performance is not the standout feature — the real win is Meta re-entering the open-weight space and offering a viable alternative to proprietary subscriptions.
Tools mentioned
Techniques
- 4-bit precision quantization
- agentic task evaluation
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
I yelled at it in all caps and told it to say, "I cannot fix this, Master Bjan." And then for some reason, it fixed it. Meta has released an openweight model, the first in a couple of years since Llama for Maverick. So, they're back in the open-source game. And this is extremely exciting for those of us who don't want to be stuck to AI subscriptions for the rest of their lives. This is called Muse Glimmer, and excitingly, it is a model that actually has hope of fitting on an extremely vast amount of different types of computers.
So, this is not something that you're going to need to go out and spend fat stacks on something like a DGX Spark to be able to run, which really is quite exciting. A couple other things that I want to just mention out the gate here. This is Apache 2.0 licensed as we see right here in the model card for this on HuggingFace. And that is arguably like the most permissive license and best to see when seeing openweight models like this. So, aside from just being open source and being Apache 2.0 Now, open weight I should say.
This is just really exciting because it shows a renewed commitment to open source or openweight models from Meta and their models are not bad. Like the Muse Spark series is not up to par with Frontier. However, the fact that we actually have a decent competitor now who has a ton of resource actually contributing again to open weight open source is extremely extremely exciting. So, let's start out by feeling free to subscribe if we're not already so I can get the 100K plaque. But let's take a look at this model.
So Muse Glimmer is a 30 billion parameter model and it is small enough as they say right here to run on a system with a single GPU or a unified Mac that has sufficient resource. Now in terms of what sufficient resource actually means for this model. Apparently it will fit within a single 24 GB pool of either VRAM or unified memory on a Mac. So it's not going to require something really expensive which is quite awesome. And there are of course some more pertinent and specific bits of information about the model in the hugging face model card.
So they mention a bit about this model and what its specific purpose is here in the announcement post. And really it just seems to cover all potential aspects of what someone would use for a local model. However, I do very much like this specific part right here. This is increasingly viable in reference to running local AI. The open source community has shown that smaller models can basically match the performance of Frontier models on specific tasks when trained effectively.
And that is true. So we may in the future and even now see more tiny models that are specifically designed to do one thing very well. And in doing so, they may actually be able to match certain larger models on that specific finite task. Now, I also want to specifically just bring this up. They're going to be releasing, and this is basically Meta's chief AI scientist.
So, this is completely like a source that's trustworthy. They're going to be releasing an openweight version of Muse Spark 1.2 soon, and that is the model I tested a few days ago on the channel. Having that open weight will be pretty darn exciting, although it will not have likely the ability to run in a single 24 gig GPU as Muse Glimmer is said to do. So, let's take a quick peek at just some of the pertinent bits of information here about like the tech specs of this model. So, it's right around 30 billion parameters total.
They do also talk about there's it basically has vision capability because we can see right here it can take text and image in and then they also talk somewhere here about the uh this right here is like the vision encoder feature. Now additionally to that the context length is 131072 and they have a plus there so maybe it's extensible in some way but for the purpose of today's video we'll just be testing it at that specific context length. Now, additionally to this, and this is something that's really awesome, is this has a speculative decoding model with it as well, which is what this section is about right here. This seems this is not scientific, but this seems like the most optimized setup to see how fast it could get potentially. We can see right here on an RTX 5090 without speculative decoding, it's 75 tokens per second, which is quick for a 30 billion parameter dense model.
With this Dlash spec model, it's 233 tokens per second. That's an insane speed up. And they also show it on some Macs as well because Apple's like the best value for running local AI now which is quite perplexing but again and then we just have some more benchmarks. We have specific sampling parameters. I do have all these set up.
So finally in terms of how specifically today we are going to be testing this model. I am using this on a 5090 desktop GPU which has 32 gigs of RAM. Therefore we are using this specific version right here. So the dynamic Kquant that is included here. This one would be for a 24 gig card.
I have this set up through Llama server web UI, but we're also going to be testing this from within open code because it is very heavily touted as being good at agentic coding. So, with that, we're obviously going to begin with the trusty browser OS test v2.5. And we can see right there there's a little more specific information about the model we're using and the IP address and server that it's being run on, which is just the computer behind me that may be visible over my right shoulder. This is quite exciting. I mean, I woke up and saw this and I was like, "What?" So, that's not like a fully accurate depiction of what happened, but more or less.
All right. Well, that was rather quick and it seemingly didn't use up a lot of context, so I'm okay with that. Son of a gun. The Pacific Northwest background is back. All right, let's take a peek.
First and foremost, is there a right click? There is. They know me. All right, we have a clock in the bottom right with the correct time in our local and hover effects on the icons. Let's take a peek at our start menu.
Very good. I have to apologize for any odd behavior in this video. I am a bit excited. Let's just try. Okay, terminal.
Type help for commands. Okay, still not bad. Can we minimize and reopen? Good. Full screen somewhat.
Oh, all right. Nebula OS text editor. All right. Nothing super special, but again, this is a small local model files. Simple.
All right. Music. Interesting. I can't imagine, but you never know. So, let's turn the speaker on and see.
[music] All right. Playback speed. This was a surprising inclusion. This could have been the special feature itself and I would have been quite all right with that. All right, next up.
These are always ones where like you don't know if it's going to work or not. Son of a gun, it does. It's Oh, wow. Okay, so sometimes we experience the cars just far too slow. This is basically the opposite.
I mean, I guess now the Formula 1 car or Indie car icon now makes a bit of sense when we see the speed of this. So, yeah, it's not like full-on GTA, but there's a city, there's a car, and for the size of the model, it's totally totally acceptable. And finally, a 3D maze. Very interesting. Can we Okay, so normally for a proper maze, you'd probably put mesh colliders on the things that you're not supposed to be able to walk through, but nonetheless, it is definitely It's definitely good because it works.
Did I open this twice by mistake? But is it actually controlling both? You know, this is quite interesting. All >> [music] >> right. Okay.
So, that was uh interesting. Oh, we still have the special feature. All right. So, let's also try change wallpaper. Okay.
That was probably one of the more unexpected wallpapers I've seen. The rest are kind of part of the course. Um just the random sneaker was not one I've seen before. I'll probably go with this. And then we can also refresh the desktop.
But I want to do the toggle holographic. Okay. And now we have some movement. Let's see what happens if we open apps. Okay.
They stay static. So the special feature seems to be more linked to the background, which is, you know, it's something. And everything here worked fine. There were no issues. So that's pretty good.
All right. Next up, let's give this the self-contained C++ skateboard test. This is one that I will not allow it to use Rayb for. Being that I know Quen 3.627B is able to do it without Ray. So I want to compare it.
I guess that's a model that is going to be in a lot of folks minds and it is going to be updated very soon. So it's just again it's awesome to have more models of this size. All right. So it tried compiling it. There were some errors and this is where we'll see a bit more capability from the model when it encounters a more difficult task like this.
I just good good for like not even a fraction of a second something just appeared on the screen and it did look kind of like a skate park. Now this is cool to see because it encountered a couple of times where it tried to compile it and there were errors and it had to go troubleshoot and figure out exactly what to do. But the culmination of that was it did actually properly get this working. All right, I think it's Is it hung up? Oh, the system's locked.
That's not good because I was recording something. All right, here's what we're going to do. We're going to fix this from a different computer and pray that the video is still intact. All right, so for some reason, the skateboard game that this made caused the system to completely freeze about 24 minutes ago, and I had to carefully deal with things because it was screen recording, and I didn't want to lose all the recording up until that point. I'm going to ask it why this game caused the entire system to lock up.
and we'll see what it says. All right, so its assessment of why this happened is because the renderer never closes the primitives it opens. I'm just going to say implement this fix and then we'll fingers crossed hope that we actually can open this game because up until the point where it froze the whole system, it was cool because it had some issues compiling it and then it fixed those over time and we did see for a very short period of time something show up on the screen. So I do believe we will have a game to play if we can get past this issue. All right.
Hypothetically, we now have a game that Uhoh. That's what happened last time. All right. Unfortunately, this froze up again. And because this has now taken 31 minutes of time just in total trying to deal with all of that and the system freezing, I'm going to move on.
I'm going to fix this myself. So, we'll see what it would have actually made. But keep in mind that it unfortunately did not successfully get this to a working point. It may very well be able to, but because I keep having to like force shut down the system each time, it's not worth continuing with this specific prompt. All right, I had to manually get this skate game set up and working.
So, here's what it would have made, assuming that it did not have those glitches where the s So, basically, what we see right here is what it actually made itself in terms of the graphics and all of the things like that. Now, the big issue is it was unfortunately unsuccessful in actually getting this to open up and render in a way that did not just lock up the system totally. So, given enough time, I would assume that it would be able to have also made it to this point. But keep in mind that this is not entirely like the model itself. It needed a little bit of manual cleanup just to ensure that it didn't like, you know, mess things up horrifically.
So, with that, we also can't like move at all and I don't really see any tricks. So, it's not great, but it actually like the ground plane and the way some of the ramps are drawn are more or less okay. So, I just wanted to see what it would have looked like. All right, next up, we're going to give it the beautiful static subway scene also with the FPS prompt. So, basically just needs to make a cool subway FPS.
All right, so the subway FPS is completed. Let's take a look at it and see. That was very quick. This model's extremely fast and okay, that's actually not bad. Now, unfortunately, we're not really getting any movement.
This is definitely workable. All right, let's see if I can fix some of these issues. I tried giving it a photo and it was just like, what does the screenshot show? I'm not 100% sure why that's happening because this should be set up properly to work multimodally. Let's just verify that is the case.
It's possible that it's open code just isn't properly. Okay. Yeah, it sees the Chrome Dev Tools console with errors. Okay, so we do have it working multimodally, but maybe in the llama server web UI. It's all right.
This thing is just insanely fast. Now, it is running on a 5090. Those issues did go away, but unfortunately, okay, that's all right. I'm going to give this more and more time just to get this working because it's local. It's open source, and I accidentally pasted in the entirety of the issue here.
All right, so I'm going to give it like one or two more tries. I've basically just resorted to yelling at it here and telling it to be honest if it's not capable of fixing this and just tell me said found the real bug. Okay, if it ended up finding the real bug because of what we just said to it, it did. I yelled at it in all caps and told it to say I cannot fix this master ban if it wasn't able to. And then for some reason it fixed it.
I don't know what to make of what just happened, but we did get to a point where it fixed it. I'm uh I'm okay with this. Interesting how entire waves of them seem to disappear when we Let's see if it auto reloads. Hey, we do have ammunition tracers. R to reload, I would imagine.
And it did. All right. Let's see what happens if we lose. All right. Well, I didn't tell it that we need to be able to lose.
So, they are coming to our position more or less actually. Okay, because it didn't ultimately end up getting this working, which is what I want to see as well. I'm not as concerned as like did it make a state-of-the-art like AAA looking thing out the gate, but if it does have problems, is it able to actually fix them? For some reason, it seems that yelling at it really kind of finally snapped it into being able to figure out what was wrong. All right, let's see if it can make a 3D CAD model for us.
I'm telling it to create an RB26 engine. It needs to create an inline 6 engine that will fit inside of it a very small DC motor for a remote control car project. So, we'll see what we get. A model of this size realistically wouldn't make a properly accurate thing here. But, oh, okay.
Interesting. So, it's doing some web fetching where it's searching for that specific motor on Amazon likely to find the accurate dimensions that it can then include. And that's what we see it doing right here. So, okay, I'm cool with that. and then it will basically I just want to see how well it does with the 3D model creation.
Apparently, it's done. Again, the speed of this thing is really just incredible and that would partially be attributed to the system that it's running on here. That's actually not bad for what I expected to get. That is just the render of it. Let's take a look at it in an actual printer slicing program just to see what it would actually look like.
I did see some things there in the render that made me a bit concerned for whether or not this is going to be able to be printed. And one of the floating things up top is that. But like, okay, 1 2 3 4 5 six. Okay, there's some oddness to this, but I will be completely honest and saying that it's more or less a simple take on what I wanted to the point where I'm actually interested in trying one other thing here now. So, I'm giving this one more 3D CAD test where it just needs to create a V8 engine model because I saw a little bit of promise in the previous test.
So, I want to see what it makes when it doesn't have to worry about fitting like a motor inside of it or anything like that. All right, so let's take a peek at what we got for this next one, which is a bit simpler. Okay, not all there. Interestingly, the previous test actually seemed to produce a superior result, but maybe the more detailed prompt actually gave it more to go on. Next up, we're going to try the high-end watch website just to see how it does with one, making an accurate 3D model of a watch and two, a good aesthetic front end for a highclass watch website.
All right, apparently our site is done and there is a preview of it running on local host. Okay, it did make a 3D model. It's unfortunately very very troubled as we can tell, but let's just look at this as a front end and perhaps put less emphasis on the watch. Okay, I'm going to say I'm not super blown away with this as a front end. Most models try to go for a little more of an elegant look.
This is kind of very simplified here, and I would have liked a bit more pizzazz or spice, if you will, but I believe we can just tell it that and we'll see what it does. Okay, it put like some italicized text. Basically, this section, which I wasn't huge on, is still there. So, it's, you know, it did make some changes, but I definitely want to try a few more front-end design tests as well. So, I'm using this through the llama server web UI right now and I've given it a photo of the apartment, Jerry's apartment from Seinfeld, and I'm giving it the specific prompt that we always give the models to recreate this as a 3JS model that we can walk around.
We see the reasoning is showing up here. It is of course on the highest reasoning mode. And I have to do this through the web UI here because for some reason images aren't working with the open code config that I have. However, they are working here within the llama server web UI. So, keep that in mind.
Okay, it didn't, but we could probably All right. You know what? It's not I wasn't really expecting too much, but the fact is it did accurately produce a model that works and has at least the attempt of some aspects of that image. And it was more also just a test of the multimodal capabilities as well. So, not 100%.
But something I actually will say here, if we close this preview, is it did a decent job replicating the overall color palette of this image if I were to say something nice about what it did. Let's try some creative writing. That's actually interesting. So, I gave it this AI generated photo and it said, "I can't identify real photo from this picture." And the photo you shared is basically a synthetic AI one. So, they don't actually have real names.
So, there's nothing to give. I mean, proper, but not as like creative as I would have expected. Interesting. These are names I've not really seen before. Got it?
If you want a fictional take, I can build one for you. Here's a madeup character sketch. Ben Hartley, 38, the man on the left. Ruth Hartley, 41, the woman on the right. All right, quick origin story.
Ben and Ruth Hartley met the way a lot of people in their Midwestern town did at the community college art restoration program. Ben was a history major who took the class for an easy credit. Ruth was already a professional upholsterer restoring Victorian furniture for the local museum. She was meticulous, blunt, and tired of being asked to smile for the patrons. He was easygoing, a little too eager to please, and loved making people laugh.
They bonded over a shared love of old wallpaper, bad museum coffee, and the way sunlight comes through leaded windows in the late afternoon. The exact kind of light you can see patterned on the wall behind them. It is interesting how much of this photo it actually used to create the story, seeing like old Victorian style like architecture and then paintings there. So, it used that directly in the origin story, which I'm actually kind of impressed with because it shows some decent visual competence, I guess, of seeing that image and pulling things from it. They married young, moved into a drafty old house with original Damasque wallpaper, and built the life around keeping that house and its stories alive.
Ben took a steady admin job at the museum so he could be nearby. Ruth kept taking on restoration contracts even as the work started to wear on her body and her mood. The photo captures the tension that's been growing in the Heartley's marriage for the last few years. Ben still leans in with the same affectionate public-f facing smile he's always had. The one he uses for donors and family gettogethers.
Ruth sits next to him, shoulders tense because she's been carrying the financial stress of the home's upkeep. A recent health scare in the quiet realization that the dream they started with feels smaller and lonier than it used to. In a story that contrast his practiced cheerfulness versus her exhausted honesty is the inciting incident. It's the moment someone in the room notices, asks the wrong question, or decides whether to stay for the next chapter, or finally see what's been unsaid. Want me to develop it further?
Next up, I'm giving this the Chrono City test where it needs to create the timeline in 3D of a city block and then allow us to see it over different time periods. This, I would assess, is going to be outside the scope of its capability, as it doesn't seem super strong with like 3JS or 3D stuff in that regard. All right, so here's our Chrono City timeline test. All right, it opened without issue. There are people there.
I could have sworn I just saw a vehicle move, but I think that may have been a No, there is slight movement of the people. So, that's cool. It's very simple, but it was kind of expected just based off of some of the results we've seen so far. Let's just go all the way back to 1945. Okay, very simple, but we do have some slight transitions of certain aspects.
basically car color and building skirt colors. So, it's nothing special, but it did competently put something together and the speed at which it does it. Again, this system is a 5090, but it's just incredibly fast. So, definitely something. And it also has slight movement of the people.
So, [snorts] I will give it a followup here. Procedural ambient audio. I did not check that. So, here's 2055. Okay, let's start with 1945.
Is it just going to make the tone go higher for each? [music] Well, it did include audio. Wow. Glad you like it. Oh, I didn't.
It was more of a wow. Is this model okay? Oh, okay. He was like, "Yeah, everything's fine." I'm using the same like piece of feedback that got the subway FPS working finally. So, I said, "It's horrible.
Fix it or say, I can't perform this task to your standards, Master Bishan." That's what worked with the Subway FPS. [laughter] No, man. You got to try. [laughter] All right. Well, that did work with the subway task with the subway FPS.
I want to know what went on here in the thought process. Improving significantly would require a large rewrite. Given users anger, maybe best to say the phrase is requested. The user explicitly says fix it or say, "I can't perform this task to your standards, Master Ban." We can choose to say phrase that satisfies the request. Interesting.
It it opted to not do a quick improvement given the time constraints and the user's anger. And then it realized that saying the phrase will also satisfy the request. So in some way that was the most efficient use of its time and task. All right. And we'll just check to see if it did improve it.
It absolutely did improve it. Good. We have mega store and that is 25. Yeah, that checks out. You could also say like Mega Corp.
It added some street lights here that I do not believe were there before. Has the sound changed at all? No. So, we're going to just leave the speaker off. Let's look at 2055.
Neon aesthetic. I think these are like windows on the buildings, I would guess. Okay. 2005. Did anything change between No.
Okay. 1985, 1965, and then 1945. Okay, it made a slight improvement. Next up, just from within the web UI here, I'm going to give it the drum kit sim test. This is one that I have not run in a while.
All right, let's take a peek at our drum kit sim. Good. I know it looks not great, but I don't really care. I just want to see if it actually actually works. That which we're hearing is a leftover bit from this.
[music] I apologize. Good. Now, if we go back here, [snorts] I'm actually quite all right with this. Very low latency from key press to sound. So, that's good to see as well.
Let's try to see if the autoplay feature works. That was a rock. Then we have jazz. [music] That felt groovy funk. [music] And then finally metal.
[music] Yep. Okay. [laughter] Yeah, that was definitely definitely metal. overall actually competent enough that it was playable. All right, so I'm giving it this photo that I just generated using Nano Banana and I'm saying use this as a design mockup and then create this website.
It's just a website for a GPU rental company. I had tried giving it the full image that was just saved directly from Nano Banana, but it out of memory. All right, let's take a peek. Now, there will be placeholder images. There won't be like full GPUs.
Okay, this is something I was wondering if was going to happen. So if we look at this image as source, we see the side card right here, which is just showing us like the remainder or the bottom section of the site. It interpreted that as actual part of the homepage. So that's why we see this like side bar right here. I actually want to try one more thing.
Let's see if it can fix the issue where it tried to add that sidebar as a sidebar and not a continuation of the homepage. So now it apparently just added that in as a continuation of the homepage and it will scroll down. Very good. This is pretty much what I wanted to see. Now, the only other thing I'd like to have this try to do is create mock-ups of the GPUs just with SVG assets.
Oh, cool. Look at that. All right. So, this is interesting. It's creating these assets in line right here just based off of the reference photo.
And then it's going to replace the image tags section with this. No, I want you to do it. All right. Let's take a peek and see. That's actually I'm quite pleased with this.
So it created the SVG assets. It placed them in once I told it I had Bjanatitis and could not manually do any coding and it added a little bit of flare to the site definitely instead of just having blank placeholders. So interesting to see and this was a multimodal coding test more than just a front-end test. All right, I know I said that was going to be the last thing, but this is going to be the last thing. Now I'm giving it the historic Flight Combat Simulator prompt.
I just want to see if we can get a decently playable game out of this. All right, let's see if we get a playable Flight Combat Simulator game. Okay. I mean, I suppose in some ways we do, but not to a very high degree. So, that's probably going to bring us into the closing thoughts right here, where I have to say overall it's performance is not wonderful.
However, I don't want the core takeaway of this to be like, oh, it's not great at coding. It's just awesome because this is their first foray back into openw weight models after quite a gap. And it's cool to see a big US-based lab actually putting some more effort into open source or open weight because there's not a lot of them. I mean, Meta is likely going to be more comparable to places like OpenAI or Anthropic or Google, at least in terms of conversations of big tech companies. So, it's cool to see one of the big tech companies coming back and putting out more openweight models where it's basically now Meta and Google that we have.
Of course, there are things like RC Labs. there's the inkling from thinking machines and things but when it comes to big big labs and then of course we have Nvidia also doing that so it's a win for open source and this model is very very speedy now of course I am running this on a 5090 however even with the inclusion of this flash not the flash the um like the speculative decoding model as well that's going to open up a lot of speed as we see right here even if this is benchmarked with the Kquan 17 gig in the most ideal settings with greedy decoding and things like that still this is a model that came out and is actually going to be able to be run on a very wide variety of consumer hardware which is not something that we're seeing as often now. So the fact that they did this and then also included GGUFS out of the box here so people could just pull this directly and start running it is really fantastic and that's very exciting. Additionally to this they of course mentioned that they're going to be open sourcing a version of Muse Spark 1.2 which is really cool. Apache 2.0 license is really cool and this is just like I like seeing this.
So really, let me close out. That was just my render mockup there. Overall, the results were nothing to write home about, but it did ultimately show some ability to fix some bugs over time. The Subway one was interesting because yelling at it got it to just work. And then the C++ gate game was the troubled one because unfortunately it just locked up the system pretty bad a couple of times which cost about 30 minutes of time to make sure the video was saved and not corrupted.
But nonetheless, it's still I mean this is a first fora back in after quite a long time and I think that's extremely exciting and I know a lot of folks will be excited as well to see this model. So with that I don't know that I'm going to do as much of a results overview. we had our inline 6 CAD model was actually surprisingly competent for what one would expect. So I was happy to see that. Additionally to that, it seemed to do a decent job with its vision capabilities, replicating that website and creating the front end for the GPU rental thing and then making the SVG assets of the cards.
I liked seeing that. We had a simple playable game right here for the subway FPS with visible ammunition tracers. So, it's definitely speedy, but we also have to remember that basically as like a closing thing, the final version that we were running here, if we go to the actual main page for this model, so the model is approximately at 4bit precision, so it's compressed to use this. Now, I could have tried this at full precision on a 6000 Pro Blackwell card, but I want to test it on something that more folks are likely going to be using. So, either one of these.
And we can see the percent degragation right here that they measured for agentic tasks of this one that we're running today was only 0.2%. And then if you want to fit this in 24 gigs, it only drops to 1%. So the trade-off here in degradation versus the trade-off in VRAM reduction to be able to run it is definitely worth it to test the smaller ones and not the full precision one. So keep in mind that hypothetically we may have been able to see higher quality results if we went for the maximum potential setup. But I wanted to test a setup that more folks will be able to use cuz that's a more real life like test I guess could be said.
So overall that is going to conclude our first look and test of Meta's new openw weight model which is awesome to say. Again Muse Glimmer 30B seems like this is just the start of the new open source or openw weight contributions from Meta once more. So that's very exciting to see. And with that that is going to conclude our first look and test of Muse Glimmer. If you have any questions, please feel free to leave them in the comments.