Meta Muse Code Is HERE – Spark 1.2 & Meta’s NEW Coding Agent!

summarized

TLDR

Meta has released Muse Spark 1.2, an iterative improvement over its predecessor, alongside Muse Code, a terminal coding agent designed to work optimally with the model. The video tests both tools across various prompts, including games, 3D modeling, and web design, finding the model acceptable but not blowing away the competition, with notable improvements in handling complex cinematic prompts and a unique UI style.

Key points

  • Muse Spark 1.2 is an iterative update over 1.1, showing improved performance on coding and agentic benchmarks.
  • Muse Code is a terminal coding agent with async background agents that persist throughout a session, not just for individual tasks.
  • The model has a 1 million token context window and pricing of $1.25 per million input tokens and $4.25 per million output tokens, with a cheaper contributor version available.
  • Testing included a browser OS, a skateboard game, a 3D model of an RC car chassis, a watch website, a flight combat simulator, the "Street Yeet" game, and a cinematic "Steve the PC Repairman" game.
  • The Steve the PC Repairman test showed a significant improvement over the previous version, with proper cutscenes, dialogue, and gameplay elements.
  • Overall, the model performed acceptably but not outstandingly, with some tests like the flight combat simulator being disappointing.
  • The Muse Code tool features plan, grill, and goal modes, and the goal mode was notably efficient for a complex prompt.
  • Total cost for the testing session was approximately $24-25, including $8.50 in free credits.

Tools mentioned

Techniques

  • async background agents
  • plan mode
  • grill mode
  • goal mode
  • co-training of model with agent
  • reasoning modes (ultra)
Transcript (captions)
Why did it >> 40,000 grandmas? Yeah, I'm in. Run cleared. >> So, Meta has introduced Muse Spark 1.2, which is an update on its Muse Spark model. Additionally to that, and something that must be pretty important being that it's actually before the new model here in this header, is Muse Code. So, this is a dedicated coding agent for this model specifically. I don't know that I'd say it's an agent, but basically what it is, and I have it open right here, it's a command line coding tool that is specifically for this Muse Spark 1.2 model, and apparently it has some pretty interesting features, which we'll talk about later on in this introduction post. So, before we get into it, please do feel free to subscribe. We're about 30K away from that 100K plaque, so I am slightly excited by that. And let's get into it. So, the first thing we're met with in this blog post, which really isn't a very long announcement post, which I do kind of like, is that this is a beta version. It is a terminal coding agent powered by Muse Spark 1.2. So, this video is going to be testing both of these at the same time. I'm almost more interested in the new model, but I suppose if it's designed to excel in this specific harness, then I guess I'm excited about that, too. We can see right here it is currently available for Mac or Linux. Thankfully, Linux support is there as well, so I don't have to rage out about using a MacBook Air or something. And some interesting things about this. It can coordinate multiple persistent sub aents for each task, solving difficult problems faster, more accurately, etc. Yes, sub aents are not a new thing, but it seems like the way this is implementing them is a bit different than perhaps what we've seen normally. So, this talks about async background agents. Muse code operates with a simple agent loop plus a set of async background agents. These agents remain active throughout each session rather than being spawned for individual tasks which is something that I have not seen as often. Usually they get spawned for specific tasks. Then of course we have some specific demonstrations here of things the model has made. One of them or two of them are games which obviously have been directly influenced by my testing. But nonetheless, so the next big part of this announcement is of course the new model. Now, this is an iterative version improvement over Muse Spark 1.1 just based off of the like it went from 1.1 to 1.2. So, I would assume it's just a further improvement and that is kind of also reinforced by the benchmarks we can see right here. These are coding specific ones and agentic specific ones. We can see that the predecessor Muse Spark 1.1 is in the light blue charts while the new model which we will be testing today Muse Spark 1.2 is in the dark blue. Now, they do also have some more pertinent information about the testing methodology and evaluation methodology. And I wonder how much of these gains could be attributed to the model improving overall versus to the model being used now with this Muse code harness, which it's probably best at. And they also mentioned Muse Spark 1.2 has been co-trained with Muse Code just to ensure that when used together, they're at their peak potential performance. So that's also partially why we'll be using them in tandem today. And then finally, something interesting they show here is just a more longer and probably technically challenging benchmark than the fun games shown up here, which I do prefer. Finally, the thing we want to talk about, of course, is pricing and other model statistics. So, we can see right here from the API dashboard, MuseSpark 1.2 has a 1 million token context window. Input is $1.25 per million in and output is $4.25 per million out. Now, unless these figures for 1.1 have changed, it is the same price as its predecessor, which I suppose is good. Now, for those who are interested in testing this, we can see the contributor version right here is significantly significantly decreased in terms of the cost. 10 cents per million in, 20 cents per million out. However, that specifically is cheap because you allow them to train on everything you send it to. And it is also rate limited. So, for today's testing, I'm going to be opting to use this one right here, the more expensive but fully priced one without that training. Although, I mean, who knows? I mean, you're sending the prompt. So, but um I don't want to be rate limited. So, that's a big part of the reason why I'm opting to use this one right here. And we can see just at the beginning right here, we have $812 remaining in free credits. So, we'll see what our total cost today is as well per task and things of the sort. So, of course, the first thing we're going to do, and I'm not going to try any of the special modes or skills that this has right now for this prompt. This is the browser OS test v2.5, the traditional one, just to see how it does. But I am interested in trying the plan grill. And then there was another one, special skills that will hypothetically make the result better. So, we'll see what we get. I suppose I'll also just passively throughout the duration of the testing keep an eye on this new Muse code terminal agent because this is a new product as well that was introduced alongside the model. So, it's not just like we're testing a new model only with the same harness. It's two new things, which I guess is cool. All right. So, it's asking for some permission. Did it just already go to do this before I said yes? Nonetheless, okay, so it's trying to ping 8.8.8, which is Google's DNS server, and it's just looking to see if it has internet access, and then it will fetch these specific dependencies for 3JS. So, that's what it was just doing right there. Also, I had neglected to mention that there are selectable reasoning modes for this model. Perhaps I'll show them if I remember, but this is on ultra right now, which is the highest potential reasoning mode for this model. So, just keep that in mind. We have 5 1/2 minutes of working and our result does seem to be there. Our credits have gone down to $78. I have no idea what they were prior to this. I think it was like a dollar more than that, but don't quote me. So, let's take a look at our browser OS test v2.5. This is This is excellent to the point where I'm wondering if this has specifically been told to make this excellent because I will say on account of being an influencer. I'm going to go out on a limb and say this is probably our special feature, we have our singularity is currently stable. Our wallpaper is void mesh and we do have the time in the top right. I've been seeing more top right menu bars here. All right, let's try the big one. This will tell us if it's benchm or not. There's no right click, but maybe that would have been too obvious. So still benchmaxed. All right. So I'm seeing like hints of just this style with the italicized test seems text. It seems like the new version of like the purplish gradients. All right. This looks I mean this looks very good in my opinion. I like this style. It's like hacker style. New file hacker one. Very cool. And it Okay. New file created in this session. photo. Oh, I ruined it. Okay, so this is our second game. I apologize for my incompetence in carelessly clicking on it. Now, it's cool. It's a bit fast and it does seem like we may Oh, no. We can lose. That's actually kind of interesting. Look at the way it drew everything here. I'm going to turn the speaker on just in case. Next term. Oh, terminal. I should have bashish. I like that. Is that a hint of humor I detect? Okay, interesting. I'm going to wait to try any of the fun stuff like singularity because that may be partially like our special feature which we like to save for last. All right, nebula pad. I'm going to assume this is a drawing saved and the time stamp as well. Export to txt. Very good. Clear. Clear note. Okay, cool. Next up, settings. Make it yours. One file, six worlds. I'm definitely seeing hints of the new like cleaner LLM text style that you may be familiar with. Okay. Amber grid. We still do have our singularity core there. So, concrete, that is something that remains, midnight teal, and then line. Okay, here's inevitably some paint result from one of the other browser OS. So, the upload image feature does work. All right. Next up, we have Steel City, which is assuming gonna be our GTA clone. Okay, that's a questionable looking human. It looks like a um those little hot dog foods. You're probably familiar with what I'm talking about. All right, we just jacked this way. It is quite slow. Interesting because another model recently had this speed cap. It was Quen 3.8 Max. It was capped to 40 km or so as well and it was infuriating. Okay, no mesh colliders on the car. Are they on the buildings? Yes. I wonder if we can run faster than the car was going. It's about a tossup. Can we take this? All right. We do have a mini map. Let me try to get to something on the mini map. It's not great. I think the Quen 3.8 Max one was better. That's just the most recent one in my memory, but it had this disgusting sinewave sound that made me mad. So, it was better than Musepark 1.1, I believe. So, an improvement nonetheless. And this game we already saw. It was honestly actually kind of cool. The speed here is just like this would need a warning for like medical things. All right. Double click the core warp mode. Wallpaper shader adds radial streaks. The core flares amber lime. The window manager injects momentum and gravity. Drag a window. Fast release. It keeps sliding. Ah, why is this becoming more of a thing? Let's do it. Oh my. Actually, that's kind of cool. So, it did actually snap. You can tell that there's gravity. Let me All right, that was interesting. So, just to show the model selection real quick, we can choose the model from when we type model. And then you can either choose the super cheap one that's discounted because you can freely allow training data and it's rate limited or you can choose the full price one that's hypothetically not. And for effort, I don't actually remember how I changed that. It would probably be the effort tag right there. So, we can see that all of these are selectable options. I'm going to keep it on ultra for the entirety of the test just to see it at max performance. So, the next test is going to be the C++ self-contained skate test, but I'm going to begin this just from within this plan mode skill. Okay, don't check the parent workspace. Stay in your folder. We'll see what this does. It's in an empty directory, but it did check the parent directory. So, we'll see what it comes up with. It's thinking it needs to do. Okay, so it found like the one script in the parent directory that still has like semblance of a plan. It's for a physical arcade machine that has like real controls mapped to it. Absolutely nothing to do with this. So I'm going to just cancel this and then we're going to start it once more. And then I suppose putting it in plan mode here. And then maybe we need to specifically paste the prompt. Good. So this is the self-contained C++ skateboard game. I have changed it so it is the early 2000's New York City block that the location needs to be. And doing that with Quen 3.8 Max just produced a fantastic result. So maybe it's a fun map to have and a bit of variability across some similar tests. All right, I want to try this grill feature as well. It basically created a plan, put it in markdown, and it said it's not going to use ray, which is good cuz we want it to be more difficult and if it used that it'd be simpler. I'm also I want to try this grill feature. And the way this was phrased here, it says grill stress tests that plan until it holds up. So, it seems like something that you would run after creating a plan. I'm not 100% sure on that, but I figure it's a acceptable time to find out. Okay. All right. So, the grill feature there that we just checked didn't ask a ton of follow-up questions, and this was after creating this with a plan. So, I just wanted to try some of these other things being that we're also testing the new Muse Code like terminal. In about 9 minutes, it seems like it has completed. Oh, okay. So, 11 minutes and 41 seconds is our total time. My mistake. That was the current task that was ongoing for that time. Let's see. So, we have $441 now. So, we've dropped a couple to a few dollars there for that specific task. Nonetheless, let's take a peek at it and see how it did. It did do a lot of troubleshooting. not troubleshooting, but just quality assurance checking. So, I'll turn the speaker on. I don't expect there to be sound. And we'll see how our NYC skate game is. It's not great, but it's also not horrible. This is part of the reason why I personally like to not let the models go and try to troubleshoot and look at them themselves because it's easier for a human to look at this and be like, "Bro, the camera's getting stuck by the buildings and it's messed up." Okay. Is the board going to perform tricks like with the player attached to them? Yep. And that's a telltale sign of like a that's small model smell right there unfortunately is when the board does not move independently of the player in this result. It's kind of a it's just not the best. Now we did tell it it should have two blocks. It should have multiple pedestrians there. So there was a lot that was kind of denoted that this should have. It's not a great result. I will say let's give it some some reinforcement. I was going to say positive reinforcement but then I realized this isn't going to be positive. All right. So, I've given it some harsh feedback and we'll see what it does. I really want to push it to do better because I think it can and I think it's worth giving it an extra time. You're right on all counts. The current single file build cuts corners that hurt. Here's what I'll fix. No assumptions. All right. So, our improvement for this gate game brought us down to 40 cents remaining, which I mean, I'm fine with paying for this cuz, you know, that was still just leftover free from testing the first model, Muse Spark 1.1. This is definitely a significant improvement. I would say the graphics are overall still the same. However, it did seemingly fix some of the bigger issues, which were that the board was not moving independently of the player. There was no actual like body or arm or foot movement effects. Those have been fixed. Now, let's see how the rail grinding goes. Okay. And we do have bail effects. I want to just see if I can snap on one of these rails. Oh, no. Okay. So, we're foot grinding cuz the board got lost. All right. Oh, good. So, definitely better. It fixed a lot of the main issues that were just inhibiting playability. There is a lot of people walking around. There are a lot of people, I should say. This just makes me think of Street Eat, which we're going to do, of course. It's an improvement. It's still not great, but it's better than it was. So, next up, we're going to be doing a 3D model design test, but this is going to be rooted in reality a bit more than just telling it to create something from nothing. So, I have this little chassis that is from a tiny remote control car, F1 car. And the problem is it stopped working and it was old and it has like a cutout for AAA batteries. It's like 20some years old. So, I have taken a bunch of photos of this which are placed in this directory right here. They're imperfect photos, but part of testing these AIs is basically just giving them imperfect real life things exactly like this. So, this is somewhat properly placed, not really, on some measuring mat. And I'm basically giving it photos of this. It should use these centime squares to kind of figure out, okay, this is how long it is. This is almost as much a test of its visual capabilities as it is just of its ability to create a model. Now, for reference, this is something that these models struggle with. So this is the result that I received for that from GPT 5.6 Soul on one of its highest thinking modes. It was not quite right. These like it just it was off. It was totally off and it was kind of disappointing. Fable 5, same exact thing. Absolutely garbage result comparatively. And I even gave those a photo well I gave Fable like a graph paper drawing of it and it still just was embarrassingly pathetic. So for reference, this is what we're kind of working with right here. It's not all there now to show like okay well what would like a properly skilled human do like myself something more like this which is the one I've started making which actually has the same spot for the motor in it and this is just like hand drawing measuring and then extruding in cat. So I've started replicating this myself cuz the models right now just still are not 100% there. And again it is with an imperfect source photos. These are not like perfectly aligned, but that is part of using these things is seeing how well they handle like the real world. All right, our free credits have expired now and we've used 50. So, I'll put on screen just in editing when I realize what the actual cost was, but we do have some renders. Okay, I'm actually going to say based off these previews, this is arguably better than like a some of the results I've seen from more intricate models. Okay, maybe not. I think the previews were perhaps a bit misleading, but I'm going to say it. Not quite. But this is a difficult task, but it tests the AI in something that's more real world because the real world is messy. So like the photos were messy. They weren't properly taken like you would for measurements and stuff like that. But it got the overall shape more or less correct. I would have liked to have seen some better emphasis on the back right here. And the holes were kind of just all over the place. I do see some semblance of like symmetry put with them, I guess, which is good, but it's uh not 100%. And we have our render right there. That's actually cool. It did show it to us just on a piece of grid paper. Like the lines are right there. So, actually, let's see what this assess the overall length of this was. I believe it's somewhere around 160 mm or so. So, it was about 10 off, but still not horribly off in terms of scale. So, next up, I'm giving this a front-end web design test where it needs to create the watch website for the Slapus watch company. It needs to have a hero section with a good-looking rendered watch, a cinematic camera pan around it, and overall just an aesthetically pleasing website. All right, so let's see if it did do a 3D model for the watch. Very good. It did. Time is okay, that's going backwards. That happens more often than you would expect it to. Overall, it's nothing super special, but I swear the crown is actually spinning. That may just be some artifacting from the way the camera is orbiting. I'll say overall, I believe this is definitely improvement over the predecessor. I don't specifically recall if we tested this same prompt. Okay, this is giving me kind of like some motion sickness. Right now, it's going to Okay, it really took to heart when I said like it needs to have like cinematic panning and orbiting in the camera. It's all there. It's together. The tick marks for the numerals are in the right spot. the crown. I don't think it's spinning. I think that's just some artifacting from the way the camera movement is happening. But overall, this is actually acceptable. You know what I'm gonna say? This is probably one of the better dial drawings I've seen. Although the date, okay, that says the eighth. That's very good. Even if it's a simple 2D artifact. Lot of white here. Okay. Swedish precision softened by summer. Good. And in the actual product cards, we can see right here we manually are able to move around. It seems like the band is there, but it's unfortunately under the table. Just stuff that happens. Can we change the color? Okay, it just shows selectable color options. Good. And then we have another one. These are pretty expensive watches. 26,000 and 25,000. All right. It's overall it's not bad. And it did do 3D even though this was the prompt that didn't specifically specify it. And I like this area right here. Not bad. Now, in the meantime, I had also given this a test to create a 3D model of an inline 6 engine that's very famous. Something that must also hold a real RC motor in it, a very tiny one. Now, I'm not expecting greatness here just based off of how it did with our other more realworld 3D test. But nonetheless, we have a model to take a peek at, so I'm interested in seeing it. Here's our open SCAD. Okay, it's actually I don't have a bit of feedback yet. Let me look at these pieces individually. That's cool right there. Darn it. Okay, so it did a good job here up top. It actually did a strikingly good job up top with this right here. Fantastic. Very well done. 2600 for a 2.6 L engine. Interesting. It has some like some knowings. I don't know why that just came to mind. And it's odd how also, oh, that's just the SCAD model of the block, not the full assembly. My mistake. Had we opened to this, I would have been like, whoa. And it just put them side by side. Unfortunately, the big problem here is there is no cavity that would properly fit this motor. So, while it did exhibit some level of chops in terms of just like a proper model of portions of this engine, the main thing that it needed to have, which was fitting a real life part, unfortunately was a failure, which it happens in kind of like some smaller models. And I also realized I forgot to check price. Okay, so we've jumped up significantly. So, we've now spent $6.85. Keep in mind, we've been testing this and we'll continue to be testing this on the highest possible reasoning setting because I want to see max performance. Next up, I want to give this a really old school test. This is a flight combat simulator game that it needs to make where you can choose from three different planes and then it's just like a flight combat sim. I want to see a bit of like personality and fun from this model. I do believe it has it within. So, I want to try to massage that out of the model, if you will. And this is always kind of fun to play because they did showcase it pretty prominently in the blog post with some games. So, we are also going to do Street Eat, of course, but that's going to be the cinematic prompt. So, we'll save that for later. All right, in 9 and 1/2 minutes, we received our flight simulator test and our spending has gone up to $847. So, let's take a peek at what we got. All right, so far this looks pretty good, actually, just from a UI standpoint. Now, I have not tried this in a couple generations of models, I don't think. So, keep that in mind. But let's just start out with when we see. Okay, it's going backwards, but maybe we're the rear gunner here. Yeah, unfortunately, it's just it's not quite grasped what I was hoping for. You can see the other planes there, so we at least get a look of the models, but it's just not hyperplayable, I guess could be said. All right, cool. There's another plane there. Oh, okay. Okay, good. All right, so not all there. So, next up, I'm going to be giving this the full out street yeet prompt, and I'm also starting it from within plan mode. This is the one where it's the street eat game where you go around ye eating things. However, it starts out with a cutscene where Gary gets cut in line, someone takes the last slice of pizza, and then the yeeting begins. So, it's been given an API key as well for Open AI just to generate the speech assets that it's going to need for the cut scene. All right, we're at $11.13 of spend and hypothetically our cinematic street yeet game is now ready to be tested. So, I do have the speaker on. All right, so far it is in the style of what we want. I do believe Muse Spark is the first model I ever came up with this prompt for. So, all right. All right. >> That was the last one. >> Was it? That was the last one. >> All right. Oh, >> okay. Interesting. It's put music in here as well. Hold space and mouse to charge. Release to heat. All right. Interesting. Um, it's a bit different than the original, but again, this is a far more intricate prompt that needs that cinematic beginning. Okay, there's a dog in there, too. So, this model's sick in the head like Opus 5. Not bad. It's also it made the vehicles way too small, which interestingly is something I noticed with GPT56 soul. So, just interesting. Let's try to yeet someone where they won't necessarily hit a building. All right. How about this individual right here? Okay, they will. It's kind of like devoid of color terms of the buildings, but it's actually all right. Oh, that was a small yeet. Oh, cool. All right. There's street lights. It's It did what we needed it to, and it actually added in some music and stuff like that. I'm not going to eat the animal. Well, the dog at least. Oh, all right. So, somewhat. So, I've begun this in yolo mode, which basically means it will just like go all out. And I've given it a goal alongside the very intricate Steve the PC repairman cinematic game prompt. This is one that I did try with a previous version of Muse Spark. And unfortunately, we couldn't really get past the first cutscene, but it did look like it had some level of promise to it. So, I'm hoping that we get something better here. And I'm allowing it again to go in this goal mode. So, I don't know how long this will take, but it will generate a lot of different scenes. cinematic camera movements assets as well. So, it has to generate dialogue for the characters as well. And we'll see how well it tackles this. All right. So, apparently that goal took six minutes. I'm curious to see what our usage went to. I'm admittedly a bit confused right now cuz this is a more intricate prompt than Let's just see what's what's up here. Okay. Steve the PC repair man. definitely has the aesthetic of the browser OS. Interesting. It also has a skip functionality, which is one of the things that made me very angry with the first model doing this test. So, all right. >> Oh, Steve, you are a deer. My late Harold always said you'd keep this old thing alive. Does it still think it's year 2000? >> All set, Miss Ellis. Fresh CMOS battery. Clocks honest again. That Windows 98 will outlive us both. 20 bucks. Holler if it forgets again. >> You never overcharge like those big stores. God bless you, Steve. >> All right. >> A chin rub for me. Steve. Mr. Thomas, you are punctual. M Ellis to the car. Yes, I will wait. >> Be right with you, Mr. Thomas. Make yourself at home. Don't touch the soldering iron. It bites. >> Okay. Click to lock mouse and then to move. All right. Open Oleg's briefcase on the counter unseen. Okay, let's just take a look real quick. So, there's Hey, look at that. That's actually not bad. Doom. Okay, it's put in some Easter eggs here. That wall art is actually kind of cool. And then the cat is also hovering in midair, but this is definitely an improvement over the previous one. Anton Vulov calls himself silver support. Runs a farm of kids in headsets. Steals from old people. >> Okay. >> Fake virus pop-ups. Drains pensions. 40,000 victims. He lives in Zelatarve, Prague. 38th floor. Faraday cage. Air gap vault called Matrioska. Dead man switch. Bureau wants it erased. Briefcase burner phone. Handler is mirror. First class lot tonight. 2140. EMP pen. Cloner dongle. Black lot caddy. You copy the vault. Plant cinder. Walk out. No bodies. Steo. Your rule. >> Steo. Why did I >> 40,000 grandmas? Yeah, I'm in. Cleared >> 100% 180 already wired. Do not lose the caddy. And Steve Volov keeps Muz Ellis's file, too. Copy. Cinderan out. Tell Bureau I want his caller list burned to the ground. Steve Mirror audio check. You are on load 2810. Fov's penthouse is dark tonight. Only skeleton crew. Camera loop every 90 seconds. Guard patterns on your phone. Vault is behind the bookcase. Biometric spoofer won't work. Use the cloner on a guard with red batch. EM pen buys you eight seconds. Go. >> All right. So, we're getting some instructions from our handler. Okay. Is that a ticket? It's like a cutscene. Okay. Oh, I hate when this happens. We need to clone one of the guards. Biometrics required. Kit number two. So, whoa, look at that. Okay, that was unexpected. We need to get near one of the guards and press number two. I think I think it still has some remnants of like things that were in Steve's shop. Clone a red badge guard first. I'm trying to All right. Well, these guards are not very good. Skip to Prague. Oh sh darn it. Okay, we now have the Cadillac as well. I'm pressing too as that's supposed to clone the badge for us, but Oh no. Bypass. Hold E. I'd like to do that. Good. Oh, Max. No. Okay, good. >> Who is in my servers? Who touches Matrioska? Security. >> Okay, now what? Extract to the roof unseen. Okay, these cards are useless, so that shouldn't be a problem. >> Matrioska burned 40,000 files freed. Fov's empire is dark. >> That was antilimactic. >> Thank you, Steve. >> You did not kill. Good. M Ellis's file was first to delete. Yes. >> First to go. Come on, Colonel. Back to honest work. Batteries and bad capacitors. That's the good life. >> Stay on roof and look around. Oh, >> okay. >> 40,000. >> It was a definite improvement over the previous Muse Spark when I tested it with the same exact prompt. It became significantly significantly better. Now, it's just odd because I'm Yeah. All right. Maybe it took a little while to catch up. That makes a little more sense, I believe. So, I think that cost about a dollar. So, the final thing we're going to be doing is the subway station FPS because why not? I like doing this one and it's always fun to play, especially as the closing test. All right, this thing's been working for about 21 minutes and I'm honestly highly confused because the Steve the PC repair man cinematic game, different cutscenes, pulling assets from APIs took 6 minutes. Now, that was in the goal mode. This is not. But still, I'm very, very confused at the level of troubleshooting and fixing. This is still doing 20 minutes into what is not necessarily a very difficult prompt. It's just odd. Now, I believe we were at like 12 or $13 or something prior to running this. So, actually seems like it might be pretty good. You know what? This is actually really good. Well, it's not really good. It's just better than I expected. Oh. That's something now. Oh, reload. Does have a train. We do have casings. Um, a bit of oddness. Now, unfortunately, there are no holes in the environment that are being left, but what happens if we lose? Okay, very interesting. All right, he's he's done. All right, very interesting UI style. And our total Subway game cost was to that. So, basically the entirety of today's testing, assuming that this has caught up and is not lingering. I will check this again just prior to finishing this recording, but we started with about $8.50 of free credit and we're now at $14 spent. So that means we spent probably around $23, assuming that is the accurate final tally for today's test. So it's pricey and I wanted to run it without that very heavily discounted mode that allows them to use your prompts for training because it just gives a more realistic feel for what this is actually going to cost to use. I'm not huge on like testing things that are partially discounted cuz that's not the actual genuine price they're going to be. So I prefer to test things that way, I guess should be said. Okay. And we can see this is still creeping up. So we'll probably wait to go back to final cost. I think overall I'm not blown away, but we have to remember that Muse Spark 1.0 came out, then 1.1, and now we're at 1.2. This is basically like an entire new generation or product of model from Meta. They're still early in terms of competing with the likes of Anthropic, XAI, Google, OpenAI, and whoever else may come along. So, the leaps are definitely noticeable. I think one of the main things that I noticed that was actually a significant increase over the previous version was the Steve the PC repair man game. Now, it is possible that because I tried this with a previous version that it was I doubt it was made better specifically because of this test. So, this could probably just reflect some genuine increase in capability on the part of the model. It put in some fun Easter eggs and things of the sort. Basically, the Flight Combat Simulator result was just flatout disappointing. I would say it didn't properly show the plane models. Everything was inverted. The watch website was acceptable and kind of this site almost like accurately reflects my experience with this model where like it's acceptable, but again, I don't know for the price if acceptable is going to cut it given that the competition really is heating up. There's going to be a new Gro that comes out soon and I would assess it's probably going to beat this. I would pretty heavily assume that. Also, we have a bunch of other labs kind of throwing their hats in the rings recently as well. Even places like Inkling from Thinking Machines and things. Now, I think this is probably a better coder than that model, but still there is a bunch of stuff coming out. And I think pricing wise, this may be a little too weak to be competitive, at least in terms of the performance noticed in these specific tests. And that's my genuine honest opinion, having paid for this, albeit not a very expensive amount, at least in the scheme of things. But it did okay. I just wasn't blown away with anyone's specific result. I think the skateboard game, it showed a proper ability to improve its result based off of some maybe not super helpful feedback where I basically just yelled at it and was like, "WTF is this?" And it definitely took that to heart and improved the main things that we had complained about. I noticed though, something I'm okay with actually is it does seem to have a more like it has its own UI style. Like we notice those elements and aspects right here in the browser OS. We also notice them maybe not as much in the Sky Fury one, but just kind of down here in some of the text on the cards. Definitely right here in the Steve the PC repair man game. And definitely right here at the Subway FPS prompt like this is its own like unique style. And I kind of like seeing that because the models are independent and have their own styles, I guess could be said. So, the final thing I'd like to do is just refresh this one more time. We're at $15.25. I would assess it's possible this might go up a bit more. So, let's call the entirety of today's testing between $24 and $25, including the $85 of free credit that we went through. So, it was interesting. And we also got a look at the new coding terminal tool right here, the Muse code. I believe it is called. And it was interesting the way the sub agents work. And I wanted to be able to test all of the different features that they highlighted here, such as the plan, grill, and goal mode. It was interesting that when we enabled goal, it did the Steve the PC repairman game prompt in 6 minutes and didn't seem to use a lot of like tokens. So that was a questionable curious one. I would urge anyone interested in trying this. I would try a prompt with goal and see how it does for you. So, that's probably going to conclude today's test of Meta's Muse Spark 1.2 as well as the new Muse Code Agentic terminal coding. So, if you have any questions, please feel free to leave them in the comments.

Frontier News · by Hyperjump Technology