Meta Muse Spark 1.3 Is HERE – Is THIS a Real Opus Competitor?

summarized

TLDR

Meta's Muse Spark 1.3 claims benchmark scores on par with Opus 5 and 56 Soul, but practical tests reveal significantly lower quality in tasks like 3D game and website generation. The model is cheaper ($1.25/$4.25 per million tokens) and decent for its price, but the benchmarks appear saturated and misleading for real-world performance.

Key points

Muse Spark 1.3's claimed coding benchmarks (75.4 on steepswe) are not reflected in the presenter's tests, which showed results far below Opus 5 or 56 Soul.

The model costs $1.25 per million input tokens and $4.25 per million output tokens, roughly 5-10x cheaper than competing models from OpenAI and Anthropic.

It accepts text, image, video, audio (partially), and PDF inputs with a 1 million token context window.

The model performed well on specific tasks like the skateboard game and the watch website after a visual fix, but overall quality did not match top-tier models.

Meta plans to release open weights for a Muse Spark model, though the specific variant is not yet announced.

Tools mentioned

Transcript (captions)

0:00 Excuse me while I go vomit. I don't think that the Meta has released Muse Spark 1.3 which comes after Muse Spark 1.2 which came after Musepark 1.1. Now something pretty

0:14 interesting about this is the benchmark performance that is being claimed for this model puts it basically in state-of-the-art territory at least in terms of some of these coding tasks as

0:24 well as the long context tasks here. These are like freak scores for a model in an iterative.1 improvement. Basically, one would assess that it would be fair to call this like muse

0:34 spark 2.0. Basically, assuming these scores do actually stack up. This deep sore right here is basically putting this on par equal to or greater than GPT56 soul on max or opus 5. So, I'm

0:48 going to be testing this and I'm expecting things that are going to be on par with this. If they're not, we're basically going to call it out. So with that, I'm really excited to get started

0:58 with testing this. I'm going to be using it. There are a couple of different ways in which you can use this and the pricing varies depending on how. So we'll talk a little bit about that and

1:07 just a couple of pertinent things about this model. But then we'll just jump and do some testing. So before we get into it, do feel free to subscribe and let's introduce ourselves to Musepark 1.3. So

1:17 for a bit of pertinent information about Musepark 1.3, we can see right here the modalities it accepts are text, image, video, audio, and PDF. The audio portion has a little side note next to it

1:28 because they mention it's not fully supported. It does only output text as would be expected and it has a little over a million token context window. Additionally to that, they do have some

1:38 other models here that they've been also releasing recently, things like voice transcription and images. So, they've been very active in terms of their release. And then, of course, Muse

1:47 Glimmer, which was their openw weight model release. And on that topic, I do believe it was mentioned in an X announcement post that they're also going to soon be releasing the open

1:56 weights for a Muse Spark model. I don't specifically know which one, but that's exciting as well because it's a larger, more performant model than the Openweight Muse image. And then we have

2:06 things like OpenAI and Anthropic where we had GPTO OSS, which was very good for the time, but then basically nothing since and then nothing from Anthropic ever. So, it's nice to see Meta and

2:16 Google open waiting some models as well. And even XAI has open-weed older versions of Grock, but like they do exist. So with that, our pricing is something to be spoken about because

2:27 there are a couple of different ways in which you can actually use this and the pricing varies pretty significantly. There is the contributor tier where it's the same model. However, your inputs are

2:36 going to be used as data to train the model further and improve it. The pricing for that is significantly significantly cheaper at 20 cents per million out and 10 cents per million in.

2:47 However, this is rate limited by tokens as well if you don't want to do that, which we're not going to be doing today. I'd like to just use it at its full price because it gives a better judgment

2:56 of how much this actually costs to run and hopefully it's less likely to get benchmaxed on some of our tests, but who knows? So, it's $1.25 per million in and $4.25 per million out. Now, again, if

3:08 these benchmark scores do actually hold up right here in coding, these two models are like between 20 and $30 per million out. So even at a not discounted rate, this could be five times cheaper

3:20 for equal to or greater than performance, which would be basically shocking to see. So that's what we're hoping to find out. And with that, let's get started. I have already initiated a

3:30 simple test with this just from the Muse code. They're basically coding agent right here for this model. I asked it who I was. It web searched, which was correct, and then I said no web search,

3:40 and then it unfortunately had no knowledge. And we'll get started with the browser OS test v2.7. I don't specifically know what to expect for speed here. I suppose if it's on open

3:49 router, we can kind of get a feel for how many tokens this would be because Meta would be the only current provider for it. So very good it is on here and we can see token speed is 85 per second.

3:59 So that's very very solid. It should be relatively smooth to actually work with this. All right, our browser OS titled Nebulon OS Nerd is completed. So let's take a peek at this. It basically said

4:11 it verified it headlessly, but it wants us to just do like a quick visual smoke test to ensure that everything is good. Okay. Oh no, you All right. Sometimes this happens.

4:24 Excuse me while I go vomit. There's no right click. There is a clock in the bottom right. Oh, okay. starfield. That's probably we'll go with that one. Let's check our

4:43 start menu. Okay, it's just very devoid of style. However, that doesn't mean I meant to say hideous here, but I typed like hide OS. Perhaps a Freudian slip. H. We'll just see what we get cuz

5:54 there's no way. There's no way. This is This has got to be a fluke. It's a fluke. It's not it. That's not It's not possible. I refuse to believe that this

6:06 was possible. I'm like I'm almost more inclined to think like the models being served wrong on the back end than than it produced that for $136. All right. So, we have a fixed version

6:20 of our Nebulon OS after I had raged out a bit. Okay. So far it's more colorful. Let's just see like Oh my goodness. Okay. Can full screen now. Um

6:40 GTA 3D. All right. All right. You know what? We're just gonna run through this one. Okay. And we double click to All right. This is

6:51 not bad. It's like low poly and simple mail. Okay, we do have some mail. Launch build past inbox is seated locally and we can't actually unless these buttons are Nope,

7:10 they're just not. Okay, we have files. I'm not going to sugarcoat it. This is like like local

7:26 20 billion parameter model from like a year and a half ago. Quality. I'm shocked at how bad this is. Yeah, this is this is awful. This is genuinely terrible. I don't know why

7:38 it's this bad. And I don't think the previous ones have been this bad with this prompt. So, I'm going to refrain from making any judgment calls on this model's performance based off of one

7:47 simple result. But as of right now, things are not looking good for for this test. So, we'll try some more intricate things that maybe will allow capability to shine. All right. Next up, I'm going

7:59 to be trying the C++ skateboard game with the New York City map. Here, I have made it so it's in yolo mode, so it won't ask for any approval. And I've also initiated this just from within

8:09 goal mode. We see here we're using Spark 1.3 on ultra reasoning effort, so the highest potential performance. All right, so we've received our skate game right here and the goal is completed. I

8:20 don't know why these things down here are not checked off, but it is done as it mentioned done. Now, this was showing just some previews of it when it was autonomously play testing it. And I'm

8:29 very happy to report that this is a significantly better result than what we may have feared to get based off of the browser OS. So, H to close this card. Now, yeah, it's not like 100% wonderful.

8:41 However, I don't believe this model is super large and that's based on no actual knowledge. Just kind of on vibe. So, I would say this is actually this is definitely more in line with what I was

8:53 hoping to see. Now, I don't know that this would beat the Opus 5 or 56 Soul at this specific test. However, it has walking pedestrians. It does have shops. It has fountain spraying, which a lot of

9:05 them include in this specific test. It's just so interesting how different models from entirely different families will include such similar things on certain tests that are not included in the

9:16 prompt as like denoted to be. Can we Okay, so there's no interaction with pedestrians. L is for a melon. Okay, nice. Okay, grind logic does work. There is bell logic as well. I think this is

9:30 Yep. Okay, that's a ramp. So, it's simple, but it's actually relatively decently put together as a result. And this is not using Ray Lab or any dependencies. So, it is a more difficult

9:42 task to pull this off properly. See, like that right there was pretty cool. I think they also like a lot of them include that in the like game and that's not in the prompt. So,

9:56 it's just it's kind of weird. All right. Well, I'm happier with what I'm seeing right here. Definitely. So, let's take a look at what our cost is up to just after the browser OS, which we did an

10:06 additional try. Okay, so $4.20. So, it's up there. But again, keep in mind this isn't like this is a cheap to use model. The thing is we're seeing the actual cost of the API access to this. This is

10:18 not subsidized by some meta coding subscription plan like the OpenAI and Anthropic models are currently. So, if this model that's pretty cheap is racking up expense this much, it really

10:29 sheds some light on the day. if those plans stop being subsidized that heavily, it's going to be a disaster, I think. So, it's just something to be aware of, and that's why openw weight

10:38 models that can run locally that can offload a lot of the work are very cool to have and will save you big bills. Next up, we're going to be giving it the beautiful watch website where it needs

10:49 to make a good-looking 3D model of a watch, a cinematic panning camera movement in the hero section. And this is also slightly changed, so it must include an exploded view of the watch

10:59 controlled by a slider in a dedicated section on the page. I'm not running this with any goal or anything like that. Nothing special, but it is still on ultra mode and going to work to build

11:08 this. All right, in 4 minutes and 15 seconds, we received our watch website result. Okay, and that cost a bit more. It's possible some of the charges are lagging. So that may not have cost what

11:18 would that be? 97 cents, but uh just make judge this overall, I suppose, instead of task specific. So let's take a peek at this. I will say it's been very quick, and that's nice to see. It's

11:29 generation speed and token speed are nice. All right, I actually it's not great, but I'm going to say there are a few things here that I actually am impressed with that I wouldn't have

11:39 expected to be. The wood texture here is not bad. And also, it actually placed this on a mat. Now, there is a dial, but it's like it's moving,

11:52 but it's only pointed at like I don't know. That'd be an interesting watch for someone who just wants to be different. A single dial. The crown's actually not poorly done. The straps have some detail

12:04 to them as well. Good. We can move this around with the threads in them. There's a bit of wonkiness in terms of like the thing that the straps attach to. However, this is actually, I would say,

12:13 a respectable result considering now, let's scroll down and see. Okay, people. One bench each. Oh, okay. Basically just saying like it's high

12:27 quality. Pull the watch apart layer by layer. All right, let's check this out. Don't let me down. All right. All right. Now, we're starting to see some other watch faces. And sometimes that happens

12:38 where Oh, wait a second. Do you see this? It seems like Oh, it actually did a way better job, but it's being blocked by like this right here. Frustrating and in my opinion worthy of

12:52 a V2. If we keep going a bit more, we actually see it created a movement as well. But it is a bit difficult to see because we can't actually move the camera around up and down. So, we can

13:05 only get a partial look at the gears and movement that it would have included there. So basically this slab right here is blocking us from seeing the face. So this would have actually been an

13:15 arguably pretty impressive result. Can we change the color? Okay, we can't. But that's okay. Good. And good. This is not bad. And I'm going to give it a follow-up. All right. So I just dragged

13:26 the screenshot location to it so it will know where to look. And then also told it the watch face is being blocked by this black disc. And overall this could perhaps be a bit higher quality. I'm

13:35 have to say though, I'm much more bullish now than I was when I first looked at the browser OS. It's interesting and something I think I may have recalled

13:45 from previous Muse Spark models is when you actually given an image here to fix based off of the image, it seems to do a lot more work. We can see right here it's been working for about 12 minutes

13:55 when the entirety of this result took a little under five. So, it's just interesting to see. It's now doing a lot of like verify, recapture. So, it's doing a lot of fine grain visualbased

14:06 code modification. I almost want to wonder just has it actually made these changes live. Yeah, it has. And that's significantly better. Now, keep in mind it's not fully done yet. And the only

14:17 big issue here is that everything is kind of oriented like 90° clockwise or counterclock. I don't know. I'm I have that like left and rights are

14:28 difficult to parse. So, look at this though. the actual tick marks for the numerals. I've seen this before as well. They're off, but they're symmetric. Like, so they're not like totally

14:39 chaotic. They're just slightly off. It didn't fix the issues we had here with the pieces of this watch, but definitely we have a much better looking face as well as now the proper additional hands

14:51 of the watch right here. Is this showing the correct time in our local? No, it isn't. Okay. So, but definitely much better. Let's see right here how our exploded view is

15:02 looking. Okay, better. Kind of still a bit odd just in terms of where the straps go. And we can't get a good look at the movement unfortunately because we can't

15:14 move this view up and down. We can only kind of bring it to here. And we have the different models as well. So, that actually looks pretty good right there. Now, keep in mind it had not fully

15:25 finished right there, but I did want to just take a peek at what it was doing. I I will allow this to finish, but I don't know that we'll see any big difference between what we just saw and what it

15:35 will give us when it's done. Cost so far is up to around $7. All right, so the watch website is now officially finished. Let me just do a hard refresh. And I'm going to assume probably

15:46 everything we saw is what's going to be the same here. Yeah. So it still unfortunately has this oriented a bit of the wrong way, but I want to just ensure that nothing here has really changed

15:57 massively. No. So what we saw was more or less like the full finished version. And we can like barely see the movement down there is visible to a degree. So it did put some attention to detail in that

16:10 as well. Then we just have the different colored ones. So overall, this is definitely a more competent result. This is pretty solid. And again, I don't know how big this model is, but I think for

16:20 the size, these are some pretty solid gains in performance. So, next up, we're going to be giving it the subway FPS prompt. This is one that is slightly changed because the enemies need to

16:30 arrive via train. So, when wave 1 is cleared, a train will pull in the station and then the zombies for the next wave will come out. And basically, it's just a little stricter saying use

16:39 3JS and create the result as something playable in the browser. Now, the Fable 5.1 result for this was absolutely madness, but I ran it on Ultra Code. It took over two hours to actually create a

16:51 design markdown document for the game, which was thousands of lines of code. Folks had mentioned wanting that. I posted it in my Discord. There's a link in the description to the Discord, and

17:01 it's in the resources channel. Now, I actually gave that design markdown document to Quen 3.8 Flash Next at a Q4 quant running locally on the system behind me. It worked for probably like

17:13 12 hours and it basically replicated it about 85% of what the fable result was, which was really really exciting to see. So, there will be more on that at a later point, but I wanted to mention it

17:23 when running this as well, cuz it was awesome to see that. Let's take a look at our subway FPS. Okay, last train. This is actually pretty well done on first glance, but remember the enemies

17:36 need to come from train. So, do we have different weapons? Yes, we do. Okay, good. That's some pretty low poly bullet holes in environment, but they do appear, so I'm happy with that. Let's

17:48 try a different weapon. Okay, we have shell casings that Oh, okay. We need to make sure we make it to wave two. Good. All right. Is there like a place where I can get a health pack or something?

17:59 Let's quickly explore this station. Oh, look. Okay, good. And we see little people drawn in the windows of the train. Let's see. My guess is they'll just kind of like pop out from

18:15 Interesting. All right. Not bad. Not bad at all. I shouldn't have done that. Oh, is that a health pack? Yes, it is. Good.

18:30 All right. Good. And then ammo restocked. Oh, there's one more. Oh, more ammo. Good. Ah, the train doesn't leave and then Okay, so more just come out of the

18:42 train, so that's okay. Oh, there's a big one, too. It's interesting. Fable did this as well. These these are the only two models I've run this specific version of the prompt with. And

18:53 interestingly, Fable also included bigger like in stature enemies when the waves became more severe. Hostiles on platform. Good luck, MTA. All right, not bad overall. Like not

19:08 I wouldn't call this like an Opus competitor because we know what Opus 5 did. If not, go back and watch the video on Opus 5 and this result was like a freak of nature. It's competent, but

19:20 based on the benchmark scores, I don't know if there were no benchmark scores of that degree, I'd be like, "Hey, you know what? This is quite solid." So, make of that what you will. But not bad.

19:29 And that took 34 minutes and 50 seconds, just as a side note. And before we run our next test, let's just take a peek at what specifically the cost has gone to now.

19:39 $91 for all of the tests we've done so far. So the next one, and I notice sometimes folks don't necessarily like the 3D model of the engine, and I kind of get that. So I still want to have

19:50 tests with these where they have to create a programmatically done 3D CAD model that can be printed. However, I decided to do something hopefully a little more fun and visually exciting.

20:00 So, it's I want a 3D printed model of a 164th scale car, commonly like a Hot Wheels Matchbox car size. That's what that scale refers to. There should be at minimum the following two pieces. One, a

20:12 flat chassis, which the body mounts to using two screws from the bottom of the chassis to the body posts. The chassis must have slots for axles designed to accommodate the following wheels. And

20:21 these are actually tiny little Hot Wheels car like style size axles that I have lying around. So, it's just something to give it a real reference point of like an actual physical item to

20:32 design around. Then a 3D car body, which must be something that can be printed with little to no supports. The design of the body is entirely up to you. However, I would like a sentence or two

20:42 about why you chose that specific design and shape. In about 11 1/2 minutes, our CAD task has completed and our usage is up to $10.66. I like seeing a couple of things. is

20:52 it's just mentioning tolerance concerns for printing with like an FDM printer, which is just the one that's probably more commonly seen, not a stinky, disgusting resin printer with that. Um,

21:03 we have it here in the file. So, all right, let's just look at our preview. Okay, you know what? That's actually it's charming and it has carved out like things for the specific wheels and

21:15 things like that. Let's actually take a look at this in the SCAD file. Okay. Now, I said it should be printable with little to no supports. It kind of would be I'll find to be

21:30 that's passable. Now, something I don't see is a little cut there for the axles. So, depending on how it did the chassis, maybe that's not necessary as it's possible that Okay. Yep. So, that would

21:42 just go right on top of that. And then hypothetically, good. It does have tapered in screw holes for the body to mount. However, those should be oriented from the bottom up instead of that way

21:53 because if we were to put the wheels on this way, they would just drop down whenever the car was moved off the ground. Yeah. So, h interesting. Just kind of like a little miss from an

22:05 attention to detail, but let's see what this thing looks like. So, the car printed well. It didn't need any supports. And as we saw when looking at the renders, the big issue was the

22:13 chassis needed to be kind of placed in upside down. So, these screw tapers were not pointing in the right way or else the wheels would fall off. I did end up actually just gluing the chassis to the

22:23 body because I didn't have any of those screws on hand. The one problem was that the chassis was like kind of too short where the axles went in. So, the wheels like it's kind of hard to explain, but

22:35 that was one issue. And the other problem was just that the wheels were too large. So, they were rubbing in the wheel wells and it wouldn't actually roll. However, when it was all together,

22:44 it did look pretty good, actually, assuming the wheels were like pushed out to the right degree. And like what we see right here, it actually looks pretty passable and decently aesthetically

22:54 pleasing, especially with the chosen wheels. So, I have now hooked this up to the Blender and GDAU MCP servers. I have verified that it can actually see both of them. And once it said it does see

23:04 them, I said good, like, build this now. And this is the 3D wrestling game with the 1980s theme. So let's see what this thing can do when given some actual tools in the form of Blender for asset

23:14 creation and things like that and then GDO for essentially my game playing. So in 9 minutes and 5 seconds it did this game which required Blender and GDAU. I just followed it up with Bro, how do I

23:26 play this? In the meantime, let's check our usage. I think that's okay. So we're at $13 total for the entirety of the test. Speaker on '8s wrestling game primed.

23:38 Simple but not bad. All right. Normally we get like arrow keys maybe function. Uh oh. Oh, you have to select this button. Let's

23:56 just go with the this is actually uh it's passible I think. I don't know like All right. So, we're this. Let's I feel like this could be good if it was

24:12 I just Oh, cool. Who won? I won. All right. It's It's not bad. I want to preface what I'm

24:32 about to say with that. It's not bad. All right. So, I've given it a bit of feedback just basically saying, "Do you expect me to be impressed with this poultry result?" So, we'll see what it

24:42 does. 73 seconds, a minute and 13 seconds. It basically said it fixed the issues we were having with this retro slam game. It It did. And look, now we see the cowboy actually has a cowboy

24:55 hat. Um, they walk out of the ring, which is kind of frustrating, but I suppose this game the difficulty here is there's not much of it. Let's just special is for P. I

25:11 don't think that the O is to drop kick. Yeah, I feel like the individual moves don't actually make a difference. It just gives us that same like effect bubble there. Oh, come on.

25:30 This is actually It's not bad. It's I mean the models themselves are okay. It's basic but playable. Victory Lab.

25:47 All right. Not bad. And it finished it. Like the overall duration of this, including an addendum to make it better, was like 10 minutes. So, not bad. And it used Blender and GDAU. All right, I'll

25:59 take it. Next up, we're going to be giving this something that folks seem to like because when I don't include it, often times I see comments like, "Hey, I would have liked to have seen this." So,

26:07 this is the Krono City test where it shows the 3D city block over a number of different eras. Basically, the ones listed right here. Additionally to that, I have specifically mentioned it should

26:17 use 3JS because it will make it easier to compare it to prior results from other models as opposed to if it goes and uses Blender or something like that. Building 1945. All right, it's taking a

26:28 decent bit of time, which is okay. Um, it didn't put an import map in. All right, I had it add the import map to our script right here. Okay, you know what? This is actually acceptable. The

26:41 vehicles are moving along the wrong axis, so they're going sideways. However, the overall scene here is not bad, and it has a lot of tool tips to hover over. Keeping in mind, this is

26:51 1945. Now, I'm not hearing any sound, which would probably be cuz it's set to off right there. Good. Postwar hope 1945 ambiance. We have the diner war bonds radio. The outfits that the people are

27:02 wearing are kind of what one would expect for this time. And I do believe these are street car tracks as well. And I think the longer thing moving sideways here. Oh, following sedan. That's an

27:14 interesting feature. That's not something I've seen before. So, we can actually click on things and follow them at least. I I don't know how that happened then.

27:24 Plaza. Good. Rooftop. Good. Street. Ye. No. Not bad. This is good. Well, you know what I mean. It's It's acceptable. It's better than the browser OS. Oh, quality

27:38 high, low, high. So, basically shadows and like a bit more I think just shadows, but that's okay. It's cool that it even included that. That's not necessary. So,

27:50 we have a bit of info. All right, let's do our tour now. Nice. Good. That Look at the color palette change to the moon. Okay. Pastel pastel slabs. Oh, that's too fast. 85

28:04 has definitely the synth wave aesthetic that every single model when doing this prompt gives it. 2005 is more kind of sterile and bland. Buildings have become taller, which is

28:16 cool to see. Okay. More eco-friendly thing with like garden rooftops in 2025. Then let me guess, 2055. It's always like neon with like hover. Yeah. Okay. 2055. Every single model just basically

28:29 makes it look like a newer version of 1985. Okay, people have like glowing things on their heads. Interesting. There are things now hovering in the sky instead of where they were before. We

28:40 still have this normie space in the center. However, it's, you know, interesting. I wish I could pause it because I can't actually really click on any of these vehicles here. Crown tower

28:52 54 m. Harbor House 28 m. Look at that. It's even showing us how large these buildings are, which is not bad. All right, I can't actually click on any of these vehicles, so we'll just statically

29:04 quickly look back between some of these. Boba and Co. Phone Lab Urban Farm. Let's see what it's done in Oh, there are drones flying around. Do you see that? Of course you do. They're right there.

29:14 Look at that. That's actually not bad. I like that. Cloud gym oat bar. And then the vehicles are again, unfortunately, moving too quick to EV follow. Good. All right. Good. 2005

29:31 yellow taxis kind of the normie space has like food and things. 85. Yep. I wasn't around then, but I'm just going to assume this is what the 80s

29:44 looked like for everyone. So, and then 65 we saw I like the transition scene for 65. TV hi-fi starlight drive-in barber. Nice fountain in the middle. Nice pastel colors on the

29:59 vehicles. I I think they have Yeah, they do have like fins on the fenders. So cool. Overall, actually not too bad. And that kind of leads me into the closing here where it may come off that

30:15 like, well, Bjan, you were kind of harsh on this model. So, we'll get into that, but let's see how much the entirety of all these tests cost. So, everything we did today cost just under $17. Now, yes,

30:25 that's expensive considering, however, keep in mind this is an unsubsidized model being used. So, were we to have done this with something like Fable, assume that everything was the same, the

30:36 same amount of tokens were used, but at Fable pricing, which is essentially four times or no, no, I'm sorry. It's like over 10 times as expensive. So, this could have been like $170 or something

30:48 like that. Even with 56 soul or Opus 5, say 20 to 30. So five times as expensive that would still put us what's 5 * 17 $85 or something like that. So that would still be like it's pretty cheap

31:02 and we don't have the benefit of the subsidization of those plans like the chat GPT $200 a month plan or anthropic which would basically negate the entirety of this usage right here. But

31:13 those won't always be around. So really the reason that I may have been a bit harsher on this than one may have expected is truthfully because of the benchmark scores right here. So when we

31:23 take a look at the coding performance and scroll down to look at this benchmark on the steepswe benchmark which up until very recently I felt pretty accurately reflected my

31:32 experience with these models. The a 75.4 means like this is beating opus and soul and say like okay it's only like one or two away. It should be on par then. It's not. These models on

31:45 max and then Opus 5 on max will absolutely have decimated this in these specific tasks that we ran today. Which says to me the benchmarks unfortunately have become a bit saturated. This one in

31:56 specific. So make of that what you will. However, something we didn't really get to touch on specifically is these long context scores are also really significantly outperforming these. I

32:05 don't know if that would transpose into being what we noticed with coding where in real life it didn't necessarily stack up or not, but it's just something interesting to take note of. Now, for a

32:15 quick results overview, I suppose we could do. We had our Chrono City, which we just took a look at, so we don't need to look at that again. The watch website, when given a photo of the issue

32:24 and the ability and chance to rectify itself, it really did a nice job fixing this. I would say this result right here is probably one that is closest to what I would get from Opus or 56 Soul. So,

32:35 this was nice to see and it did a lot of work when I gave it an image and wanted it to fix it. Actually taking longer to fix this result than it took to even make the entire thing from the start.

32:44 The Subway game was competent. The big problem was just basically when you're looking at that benchmark chart, it's not stacking up to what those two models would have done. But were those

32:53 benchmarks not there, this would have been a completely fine competent game. So again, make of that what you will. And the cost to produce this was significantly less than what one of

33:02 those models would have run us. assuming we were using those through API as well. So, it had nice enough details and it did adhere to the prompt well. Then we had our wrestling game which we had it

33:12 use GDAU and Blender to create. It made some changes when we yelled at it saying there was too much bloom effect and it was too difficult to see. It basically zoomed us in. It was really basic. There

33:22 were actually like no special effects. There were no movement. Like the characters weren't rigged. However, it was done. The entirety of this, including fixing it, took less than 10

33:32 minutes to use Blender and GDAU for this as well. And it had some sounds and things like this. It was playable and honestly not bad. So, make of that what you will. I enjoyed playing this one. It

33:43 had the cowboy hat in like a cowboy. I mean, let's invert that. The cowboy was in a cowboy hat and then like the disco person was in purple and things like this. So, it was well enough and had

33:54 like some raster effect and things like this and some sass. The skate game again was one that if this was not benched that highly in comparing with Opus and Soul, this would have been a totally

34:06 fine result on its own. I think this is something that is probably going to be pretty similar to what we'll get from Gemini 3.8 Flash. So, it actually did okay with this. Everything was competent

34:19 enough. The fountain there, even the water effect is not bad. Grind logic worked. the tricks worked properly where the board moved independently of the player and it was overall actually

34:30 acceptable. Then we had our little 3D car. Unfortunately, it had some blatant issues where basically the chassis was flipped upside down and the tapers for the screws were on the wrong side, so

34:40 the wheels would fall out if you picked the car up. Other than that though, if we just put it together upside down, everything was fine and it was more or less okay. I do have to admit I

34:49 neglected to actually see the part of the prompt where I told it why did you choose this. So let's look at that. I chose a low 1970s style wedge supercar count/stratos lineage countage. A wedge

35:01 is all upward-f facing planes vertical flanks and an inward tapering cabin. So it prints upright with zero supports and its flat floor short overhangs and truncated cam tail suit. A flat 164

35:12 scale chassis and the narrow 18mm running gear. Okay, good. So that was its reasoning for that. The final thing is the browser OS which we'll never speak of again. So with that, it's going

35:24 to conclude our first look and test of Meta's Muse Spark 1.3. Yeah, they're getting better. And also the mention of some Muse Spark being open weighted at some point sooner than later is

35:34 extremely exciting as well. I'm extremely eager to see how large this model is in terms of size or how small it is. I would imagine were you to tell me like guess now

35:44 187 billione with 17 billion active. So uh yeah that's going to conclude our first look and test of this. There's a bunch of other models to test like 38 flash quen 3.8 max 0902 checkpoint and

36:01 then whatever comes out in the future. So, feel free to leave your questions in the comments and thanks for watching.

Frontier News · by Hyperjump Technology