GLM 5.3 Is HERE – Is THIS the BEST Open Model Yet?

summarized

TLDR

GLM 5.3 is a 753B-parameter open-weight mixture-of-experts model that matches closed-source giants like Kimi K3 on benchmarks, but real-world coding and game-development tests reveal a gap between scores and practical performance. The model excels at complex reverse-engineering tasks (e.g., trying to turn an iPod Nano into a PC monitor) but stumbles on creative coding, leaving a mixed first impression.

Key points

  • GLM 5.3 is a 753B parameter MoE model with ~40B active parameters, significantly smaller than Kimi K3 (2.8T) yet competitive on benchmarks.
  • Benchmarks show strong coding performance, edging out DeepSeek V4 Pro and Qwen 3.8 Max in some areas, but real-world tests (C++ skate game, 3D wrestling game) were disappointing.
  • The model lacks native vision, relying on Zcode's tooling to verify visual outputs, which limited some tests.
  • A reverse-engineering task to display PC vitals on a 4th-gen iPod Nano showed deep technical capability but ultimately failed due to Apple's proprietary firmware encoding.
  • The model introduced vulnerability discovery data in post-training, raising concerns about potential misuse of open models for exploit generation.
  • Disabling thinking mode is no longer supported in GLM 5.3, and output tokens per task dropped significantly even on highest reasoning effort.
  • Weights are promised on Hugging Face in a few weeks, but the model is already accessible via Zcode's paid coding plan ($65/month).
  • The reviewer found the model's creative coding results (games, 3D scenes) underwhelming compared to expectations from benchmark scores.

Tools mentioned

Techniques

  • mixture of experts
  • post-training scaling
  • vulnerability discovery data
  • reverse engineering
  • firmware archaeology
Transcript (captions)

0:00 book the hammer Hansen. All right, he's in the ring. We can run. Oh, look. There's a subway red. I've never seen one of those before. I don't know where it went, but GLM 5.3 has been released.

0:10 And this is quite exciting because beyond just being the second big openweight drop of the day, this is a very, very powerful model that is on par with state-of-the-art closed source

0:20 models. And this seems to be a pretty potent jump in capability over its predecessor, which was, of course, GLM 5.2. Now, something I like to see just out the gate about this is it does seem

0:31 very compatible, not compatible, but very comparable to Kimmy K3 in terms of scores and benchmarks. However, this is a significantly smaller model. While it is still a bit too big for like 99.9% of

0:44 folks to be able to run locally, even while quantized, it is exciting because it's a subtrillion parameter model that can trade blows with Frontier state-of-the-art models across a number

0:54 of different domains. Now, while this model is not currently available to download, as we can see, the hugging face link is coming soon for the weights to be released. They do say in a couple

1:02 of weeks, additionally to this, following some more like safety and just proper preparation, those will be released. And I have no reason to doubt that they will in fact land on hugging

1:11 face. For some specific information about size, we can look at the predecessor because they do mention that scaling post training is all that they did for GLM 5.3. And it does seem like

1:21 that went rather well, especially considering some of these capability leaps between this and 5.2. This is a 753 billion parameter mixture of experts model. And I do believe it is around 40

1:33 billion active. So again, a pretty big model. However, compared to a lot of the models here, well at least the only other one that we know specifically the size of, this is significantly smaller.

1:43 So Kimmy K3 is a 2.8 trillion parameter model and this is under or around 750 billion. So to have performance that adequately can match that in a much smaller footprint is really like the

1:55 ultimate goal. Now this announcement post is pretty light in terms of what they say about the model. It just shows some benchmarks and then we have some more intricate and verbose benchmarks

2:03 here across a number of additional models. Things like Quen 3.8 Max which is pretty performant in and of itself. Deepseek V4 Pro0813 which the date right now is the 14th. So that model did just

2:14 come out. It is nice to see the benchmarks of that as well. And again, that's a 1.6 trillion parameter model, and it does seem like this is edging that out in coding, which is pretty darn

2:24 exciting. Now, something they mention here fairly prominently, which I do wonder if this could perhaps be twisted into some form of like fear-mongering against open models. And I do just want

2:34 to honestly mention that because it was the first thing that came to mind when I did see this section of the announcement post. They just say that they introduced vulnerability discovery data in some of

2:42 the additional post training that they were performing to create this model. And basically they say that it quickly continued to develop these capabilities as training scaled. It didn't simply

2:52 become better at identifying isolated flaws. It was reasoning across multiple stages of exploits forming coherent plans for complete chains. They do also mention a bit just about the training

3:02 process and perhaps some more technical bits of information there as well as some API changes. Now they mention here that disabling thinking is no longer supported by GLM 5.3. I can't

3:12 definitively state, but I wonder if it was supported by GLM 5.2. Nonetheless, that is just something to take note of as sometimes with the reasoning set to the highest, things can take quite a bit

3:22 of time. And there was a benchmark showing that. Although, something I also do like to see is we can see here the average output tokens per task compared to 5.2. Even on the highest reasoning

3:32 effort, that seems like a pretty substantial drop in output tokens. So, hopefully some of the results don't take quite as long. Which leads me into the next segment, which is how specifically

3:41 we're going to be using this today. So, I do have the GLM coding plan. This is the five times usage. And I basically think like I forgot to cancel this, but I believe in the meantime it actually

3:52 became one more expensive and two I could have sworn I heard something a while back where they were actually pausing new signups. No, that may have been for Kimmy. However, the point of

4:01 this is I have this coding plan. I pay for this myself out of pocket and it's apparently $65 per month right here. I do believe there's a bit of an increased usage aotment for the release of this

4:12 model that was referenced down here if we scroll all the way down. Yeah. So through August 31st, I guess there's this 1.5 time limited quota boost. So that's good. And how we're going to be

4:23 using this is through Zcode, which is their own coding application. It is very very similar to OpenAI's Codeex or well Codex is now defunct, the official chat GPT app, but it looks a lot like Codeex.

4:34 And honestly, I used this for the 52 test as well. and I absolutely loved it. It was really quite fantastic. So, that's what we're going to be doing for today's test. Now, I have already a

4:43 while back started the browser OS test with this because I just figured I'd kick something off so we could just jump right into taking a look at the results. So, with that and before we go further,

4:52 do feel free to subscribe so I can get to 75K and then 100K to get that plaque. And in 40 minutes and 48 seconds, we've received our browser OSV v2.7, which is the same except it includes that they

5:04 need to have procedurally generated wallpapers with movement and an email client. These were suggestions made two videos ago by one commenter. So, thank you for those because it's nice to

5:14 change things up a bit. So, let's take a look at it in our browser. All right, so far I dig it. Let me turn the speaker on just in case. I don't see a little speaker icon there, but single file,

5:24 zero depths. All right, good. Minimize. Unminimize. Full screen. X. And we do have sound effects. So, that's interesting. All right. The background's cool. It's definitely less hideous than

5:36 the Quen 3.827B, which is the first one I've done the 2.7 test with. We do have the time in the bottom right, as well as a couple other. Okay. Next wallpaper. And then I would imagine that is toggle

5:46 sounds. Now, the big one. Is there a right click? Yeah, every big model is pretty much doing right clicks now. So, make of that what you will. Oh, we just got a new mail notification. Disgusting.

5:58 Now, the first app here is, of course, the GTA clone. So, we're going to jump right into that. I just also want to look at the start menu. Without wasting further time, let's check out Omni City.

6:08 All right, this is this is good. No moving effects, but spaces to jump. Okay, this is competently put together. Simple. But all right. Can I enter this

6:26 vehicle? Yeah, I can. Okay. There is no There is no Okay. Um sounds of the engine. We have three stars right now. Let's see

6:41 if we can get it to five. Oh, nope. Okay. Oh, so I did see a I did see a police car chasing us. I didn't I had not um before this seen that the folks we hit

6:55 don't necessarily disappear. Uh-oh. Do Okay, busted. Fine. 100 total cash 50 spawning back at the point. Not bad. Not bad at all. I would like a

7:07 little bit of ability to cause problems without being in a vehicle, but it's okay. Still though, and this is crazy to say, the Quen 3.627B, so not the newest one. The oldest one

7:17 now or the older one had a better walking like thing where the legs moved individually. Can we double hit? Okay. No, we can't. But new mail from mom. So, we're getting

7:30 email notifications as well as part of the browser. It was all right. Not bad. And the mini map looked good up there as well. Void Runner obviously going to be some form of asteroid shooting game

7:39 because for some reason this is what they all choose. All right. This is good. I do like the like fluidness of our ship. Okay, I think I just hit something

7:51 that I was like a power up or something. So, look at the engine glow there. Oh, we can also drive with the uh cursor. Cool. Are these like Oh, okay. So, we're not supposed to hit those.

8:03 And all right. Share via mail Nexus. I'm just going to click it to someone at somewhere. That's funny. Attach any file. Copy saved to the send folder. And this brings us to our email thing. All

8:16 right. There's a lot in here. Basically, it just made like meta references to the Yes, Mom. I built an OS. It has a Mail Map app, so you can email me here. Nothing in spam. Only one in archive.

8:29 Okay. Next up, files. And this brings us to our Notepad app. Save as file.txt. All right. acceptable. The sound effects are

8:44 actually not infuriating, which is not something they would normally infuriate me is what I wanted to say. Help. Oh, wow. This is a very, very long list of potential commands. Let's just do the

8:55 traditional very Yesh matrix. Oh, okay. Not really as exciting as I'd hoped to be to see. We

9:10 saw notepad synth and it's even showing us the specific frequency of the keys. Not bad. Can we change the octave? No. Z and X to shift octave

9:30 square. sawtooth and we have adjustable delay. Yeah, we do. Nice. Then finally, triangle. Beautiful.

9:55 All right. Well done settings. All right. Let's take a look at some of our other procedural wallpapers. Warp Starfield. Yep. Liquid light. Oh, okay. I was hoping for more of like a lava

10:07 lamp aesthetic, but I guess that works, too. Neon horizon. Yeah, you should have started with this. You might have gotten insane in the title. No, I do actually like this aurora flow look. Oh, look.

10:16 The accent color actually changes the movable background as well. 24-hour clock is a selection. Wallpaper motion can be toggled on and off, and it actually saves where you pause it and

10:26 then initiates it back from there. That's pretty cool. The final thing we did not take a look at is of course the special feature which oh so the special feature is just basically like more

10:36 immersion and logs that have cross app like logging and information that gets saved to this virtual file system. So essentially when we were able to share or email our score of the GTA game that

10:50 was basically the special feature. Okay, very competent. Next up we're going to do our New York City skate simulator using C++. I have also appended to this prompt and we'll probably just keep this

11:01 do not use ray assuming we're working with a bigger more performant model like this. So this is the one that is essentially the same as the traditional self-contained C++ skate game we're used

11:10 to. However, the map needs to be an early 2000's New York City block. All right, so after 38 minutes and 48 seconds, we've received our skateboarding game test. I noticed and I

11:22 had totally neglected to recall that this model does not have vision by default. So, it was not really able to just look at this visually, but it made itself a little script just to verify

11:31 that all of the key presses and tricks worked. And then this has a tool built in using Zcode that will just basically send an image somewhere and then give it a like piece of feedback on whether or

11:41 not the image includes what's needed. So, just kind of one of the limitations of not having vision, but Zcode, at least in this subscription and harness, seems to work around it well. Let's now

11:51 take a peek at our H. Are you kidding me? This is just flatout bad. I'm going to be completely transparent here. This is worse than I expected by a significant significant

12:03 amount. I am quite disappointed with this because Quen 3.8. Okay, I guess I have to tamper my expectations because Quen 3.8 Max is a 2.4 trillion parameter model, but it absolutely absolutely blew

12:16 this away. This one like if I had seen this from 3.827b, I would have been like not bad. But from this, I am overall rather disappointed in the just result here. I don't know

12:31 why. Okay, cool. We can bail and the tricks actually do work where the board is independent of the player. That's not often always the case. It does have pedestrians. It does have cars we can

12:41 grind on. So, it added a decent amount of things. And I noticed it was debugging like the halfpipe tricks here, which were pretty cool. But overall, I'm going to say that I am a bit let down by

12:52 this result. I was expecting something significantly significantly better. It is possible that maybe we'll come back to this, but I would like to move on and just try some more tests. I guess the

13:02 only other thing is can we go in the water? Okay, we can grind on the fountain, which is cool. And trick-wise, I mean, like the tricks are decent. It's just I don't know. I mean, let me know.

13:14 Let me know down below what you think. No, I hate when people I hate when YouTubers talk like that, but I suppose in this specific scenario, it is uh a reasonable thing to say. So,

13:25 not it. All right. So, folks have mentioned that maybe it is time to include some new prompts that have not been able to make their way into the training data because they simply didn't

13:34 exist prior to these tests. So, that's exactly what we're going to do right here. So, this is I'm going to just read it all out so we can get familiar with this new test and you can let me know if

13:44 you like it or not. Let me know down below. No. Um, but this is a wrestling game. Create a 1980s themes themed 3D wrestling game. The game should be at the level of quality one would see from

13:56 a high-end indie game featuring a fun and engaging freeplay style. The game should begin with a start screen, procedural music, stylized elements, and a 3D orbiting of the ring. With the

14:06 start menu overlaid, there should be a simple toggle to turn the music on and off. The game loop is simple. The player will select one of four wrestlers to play as and begin. When the game begins,

14:16 the player starts in the ring with two other wrestlers. The potential list of moves they can use, their names, effects, and implementation are up to you, but they must be triggered by the

14:25 following key presses. And then I basically looked at the keyboard and picked like some random ones that were next to each other. The wrestler models should be realistic and have the attire

14:33 one would see in an 80s wrestling ring. Interactive props are always a plus, but the specific implementation and theme remains open-ended beyond the instructions above. Happy building. Go

14:44 all out. So, this will be just a fun new test. And I would imagine that this may be pretty darn funny. At least some models will probably make this hilarious. I would hope. All right. So,

14:56 after a total of whatever the addition of this number right here, this number right here, this number right here, and this number right here is, we do have our wrestling game titled Mega Slam 86.

15:09 Now, this was really frustrating because it kept having network errors. Their service is not very reliable. It has not been throughout the entirety of the time I've used it, which is beyond today's

15:19 test. It's in previous tests as well. So, that is something that is a little frustrating. However, we did ultimately get a result right here, which does very much have an 80s aesthetic. So, I did

15:30 basically just get so enraged with this consistently quitting out that I ended up just going to bed. But now we're back. So, apparently it is able to be like looked at here. So, let's just take

15:42 a peek. Here's Mega Slam 86. Oh my goodness. Buck the hammer hand. Interesting how that one is just not I'm just going to go with LT Gray.

16:08 Oh, this is actually I hope the lighting goes down. These are the two opposing wrestlers that we're going to be wrestling. Okay, this is us.

16:38 So, the big issue here is the lighting makes it impossible to see anything. And this model does not have vision. So I'm just giving it a description of the issue. I honestly

16:53 don't know if lighting bloom is the correct way to describe what the problem is. But I just told that it blocks us being able to see any of the wrestlers or the ring. So hopefully after a number

17:02 of other the match view is blown out. Cut the light intensities, tone down the bloom, and drop the exposure a bit. Then reverify in a browser. Okay. Well, I liked what I saw so far. All right. That

17:12 actually seemed to fix it really quickly, which makes me quite happy. though. Let me turn the speaker back on. We'll start a match. We'll go with the same LT Grey wrestler. Oh, wow. Okay.

17:23 Turn that down. Book the Hammer Hansen. All right. He's in the ring. There's a referee as well. Rockco Rebel. So, these are the two wrestlers we're going to be fighting. And then this is us. We do

17:35 also have more visibility into the crowd and they're holding signs. All right. So, okay, that's me. Uh I uh something is

17:51 D. The tiger torpedo. Oh no, that's O. This is I think this is troubled to be completely honest.

18:16 Okay, we just got yeated through the table. Oh, this is annoying. It's It could be really good, but like the the people are not quite oriented properly. I understand it's wrestling. I'm just

18:29 like spam clicking stuff. I we this is definitely one I'll be very interested to see the output from a number of different models. It's just something's messed up. Look at

18:40 the reflection or if that's like a big screen showing it up there. That's actually pretty cool. Okay. Yeah, this is press enter. Pin him.

18:52 Okay. It's unfortunately that was like 80% there and the 20% was the orientation. Still, it was a bit difficult to see. I think I'm going to move on, but this is

19:09 something to keep in the pocket, I guess. So, I also did do the test of the 3D printable motor. This is the inline 6 RB26 engine. This is something that took around 33 minutes. So, actually not too

19:21 bad considering that it's taken longer to do some, I suppose, simpler HTML results. So, this is interesting. It did do the job here, but it only gave us the SCAD model. It didn't actually give us

19:32 individual. Okay, so it's saying I'm just rereading this. You have to open it. Pick a part from the drop down. Press F6 to export the STL. Okay, so block. Ah, now we're

19:47 starting to get somewhere. This may be kind of in line. I'm just not used to ever seeing a result like this. Okay, now we're now we're getting good. You know what I

19:59 mean? That's a very very decent turbo model compared to what I've seen so far in this test. So, that's good. Okay, I'm so much more happy right now than I was a few minutes ago because it was just

20:12 rendering one piece at a time here. So, it's telling us that we need to basically export these plate. There we go. So, this is the entirety of what it's created for us. And this is where

20:22 things start to actually come together a bit. All right, not bad. I'm going to notice though that this is not really able to be printed without supports, especially because of this hanging piece

20:31 right here. That would cause like it's not possible to print this specific thing without having a gigantic amount of supports between here and the top of this. Like the cover looks good,

20:43 clean text is extruded well. And then we do have six individual things and an oil cap just to show us like how many cylinders there are. Even something like this, it's not really possible to print

20:53 this without supports because of these little nubs that get extended beyond. Were those not there, it would have been more in line with it. So, this was actually a pretty interesting summary of

21:02 things that it found because a lot of what we noticed like, oh, this won't print because these nubs are going below the bed or this is floating in midair. It also picked up on as well. So, if

21:11 even if we don't have things specifically to print, we're actually going to have proper renders now in STL. So, let's just take a look at these renders. That is a really really stellar

21:20 result I would say just seeing the render together for the first time. Really the only thing if I were to nitpick completely right now is the turbos are maybe oriented in the wrong

21:28 direction. Like they're flat but they should be kind of vertical. Though if we look at all these parts now individually, this is definitely seemingly a better job. I do also want

21:39 to check a couple of things that I know would have caused problems before. So let's take a peek at these STLs. one was in the block where it had that floating ledge kind of in midair. Now, I do see

21:49 that it mentioned that it should be able to bridge this. So, I guess that's possible, but still, I think that's a wide enough gap that generally you would probably see some form of support placed

22:00 underneath this piece right here. So, just keep that in mind as well. Excellent. So, it did fix that completely. Before this would have really had trouble and needed more

22:09 support. This did a good job. It really did pick up on a lot of things we noticed. Some other ones were the little nubs, I believe, sticking down from certain elements. I don't know that the

22:18 plenum was one of those and it did fix those. So, these had little nubs sticking out underneath them and it did notice that and remedy them. So, interesting. Even if I may have

22:26 mistakenly told it like this is seriously messed up because I didn't realize that you had to individually select the parts, it ultimately ended up in producing a significantly better

22:35 result. It does look very decent aesthetically and even just like the pulleys and belts and the front of the motor and things like that, not bad. and it showed some pretty interesting

22:43 capabilities just in terms of how it went about figuring out which parts need support. I was impressed with what I've seen right here overnight. I did also run the Slapis watch website where it

22:53 needs to create the beautiful website with the 3D watch render, a cinematic panning effect over it. And overall, it's more of a test of front-end design and 3D modeling capability. So, let's

23:02 just open this in a browser and take a peek. Okay, we have a nice loading screen. Not bad. The the one thing I'm seeing here that is a bit frustrating is that

23:14 the straps are oriented in the wrong direction rendered live in your browser. I also can't really like go up and down with the mouse in terms of controlling the way the camera is

23:25 panning. So, it's a bit difficult to actually get a proper look at this. I'm going to scroll down and we'll see. Good. Whoa. Okay, maybe good. Here's the Riviera Solstice. Look at the back. The

23:37 material as well as the text. I don't actually normally see anything written on the back like that. And then if we inspect the face real quick. Okay, Slapis.

23:46 Overall, this is actually I think it's pretty good. I'm having a tough time judging it just because the strap is kind of taking away from some of the impressiveness of it, but the crown

23:56 looks good. The materials look good, and the actual texture or like pattern on the face of the watch is pretty cool. There's a slight gap between the top and that, but if I were to nitpick, the back

24:06 though is something that is impressive. that I've not normally really seen before. The way it both has like this brushed material as well as the text written on it in a pretty darn nice way,

24:17 I would say, all things considered. Okay, there's some oddness to the second hand here, or one, two, three, four. This watch has four hands, but I guess that's okay.

24:27 Let's see there. That I like. This did a really good job with the strap material. Like the materials here are pretty well done. It's just the orientation of them is a bit odd, but it's possible it did

24:41 it purposely here just to showcase like the full entirety of this. Interesting. This watch does not have a date on it, but it does have the nice texture of the face plate as well. And then on the

24:50 back, we do also have that. So, I would say material-wise and some of the way it rendered like text on this is actually fairly impressive even if on first glance like the top hero section was

25:01 kind of mid. Definitely not bad. Like right here. And the the emissive properties of these materials are pretty cool. It it does give me more like keyshot lighting effects if you're

25:11 familiar with that. Just the way when you rotate a piece like this, depending on where like the point light or whatever in the scene is, it would exhibit some behaviors like this just of

25:20 the reflection. So interesting. Next up, I'm going to give this the Street Eat game, but the one that has the cinematic beginning cutscene where it has to generate some voice assets with an

25:29 OpenAI API key that I've given it. And then it just jumps into the traditional Street Eat game. All right, so in just under 40 minutes, we received our Street Eat result. And again, this is the one

25:39 where it has the cinematic introduction where Street Eat gets initiated because someone gets caught in line and the last slice of pizza gets taken from them. Now, I my computer kind of freaked out

25:49 during this, so I had to just force restart it, but it seems like that shouldn't pose too much of an issue. Street. >> That was the last one.

26:04 >> Was it? >> That was the last one. >> All right. Good pizza slice. Uh-oh. >> Okay. Okay. All right. All right. Click to

26:19 grab the fist. All right. Let's see what's up here. All right. It's I mean, the yeeting is definitely there. However, I have to say I'm not too blown away with the actual like

26:43 city and movement here. I should say lack thereof. I mean, it's a I did not mean to hit the Okay, we get it. This is almost like like forcing you

26:57 to like look at what you did. I meant to hit the guy there. All right. These are like like orbital yeetss as it says right there. Okay. It's

27:12 again I got to be honest and and let let me know down below no if I'm if I'm in the wrong but I'm really so far I've not been hugely impressed with any of these 3JS results or like games and stuff like

27:28 that. It seems like this thing's really just trying as little as possible. And I don't know why. I mean, the Quen 27B game was arguably like as good as this, maybe a little better because the

27:40 pedestrians were actually moving. So, I'm not 100% sure what's up with this, but it's it's not exceeding expectations. It's not even meeting them. So, I don't quite know what to do

27:51 because I genuinely am at a bit of an impass as to what to do. I don't want to continue giving it tests in scopes that it's consistently not really being impressive. So, I think I'm going to go

28:02 back to the C++ skate game right here. I've told it it's basically very low effort. This is akin to something we would see from a much smaller local model. I'm not really sure why. I know

28:11 you're capable of making this great. So, why don't you retry it? Go all out and make it completely high quality. In the meantime, while it does work on this, I may also try to have it tweak one or two

28:21 other results just to see if we can kind of nudge it into performing very well. But I think it's time to essentially think of a test that's going to allow it to showcase a different capability.

28:32 Maybe something more in the physical world. All right, so in 23 and 1/2 minutes, it improved the skate game. So we now have V2. And it mentioned that V1 deserved the criticism. Okay, we now

28:42 have audio. and I was hearing it make like some weird noises from the computer while I was just editing the portions of the video we have so far. All right, essentially it seems like we're going to

28:51 have just a significantly better overall experience here, which I am pretty hopeful to get. Oh, let's take a peek at it. Okay, so we're going to maybe just do

29:05 without the sound for the time being though. We notice that. Nope. It's still I can't actually move at all. The block is definitely a bit more lively, but it's arguably worse because there's

29:25 there's no functional gameplay now. Oh, I see the skateboard over here. So, the skateboard is something's just not quite right here. It did a better job with the street

29:42 textures and things like that, but yeah, this I just don't quite understand. Let me know down below what you think. All right. So, I'm working on trying to make

29:55 a difficult test for this that involves like old computers and then making newer electronics work with old computers. So, in the meantime, while I kind of figure out exactly what that entails, we're

30:05 going to do the traditional Subway FPS test. I would imagine this should be pretty straightforward and give us something good to play. All right, in 22 and 1/2 minutes, we have our Subway FPS

30:15 test. So, this is one that's probably going to be good, I would hope. And it is creepy. Good weapon model. Also some side to side movement. These vending

30:28 machines are arguably the best vending machines I've seen in this. All right. Okay. Good. Let's see. All right. We have to reload. We can run. Oh, look. There's a subway red.

30:45 I've never seen one of those before. I don't know where it went, but All right. This was I do believe the ammunition makes holes in the environment. Yeah, it does.

31:02 They have hit boxes, I think, in the center. Yeah. All right. It's a bit dark, but we do have the brightness slider, so we can remedy that. We're up to 220% brightness now.

31:16 Very good. graffiti, no train, but sometimes that's an interesting decision that they either do or don't include. All right, this is

31:29 it's good. It's not mind-blowingly so, but it is competent and everything is cool. It did also include some things I've never ever seen before, like stains that don't seem to go away between

31:40 rounds, as well as that rat, which was just kind of that that was cool. All right. And these vending machines were

31:54 were the best ones. All right. So, here's what we're going to do. I have one other thing that I want to give this that to be honest with you, I don't even know if this is possible. I don't know

32:05 what it would entail were this to be possible. But so far, the model has not really blown me away in the way that I would expect it to considering that it's benchmarked to be around the level of

32:14 like Kim K3. So, I have an old iPod Nano here. I believe this is a Nano. And I'm basically going to plug this into the computer and say, I want this to show the GPU utilization of this computer.

32:29 Like, go ahead, have fun, good luck. Because this is more in line with, I suppose, its abilities in terms of like exploit generation or not generation, but things

32:39 of the sort. So, I think this will be a more novel task that will show us some more capability. I don't know if this is possible. Again, I don't know what this entails, but it's something that I want

32:49 to try because I want to see greatness from this model. All right, so this is plugged in and I'm glad it still works. I would imagine it would, but I mean, you never know with older electronics.

33:01 So, I'm going to now just initiate this prompt. Again, I don't know. It's possible definitely, but I would imagine just being someone who doesn't know this off the top of my head that this would

33:13 require basically like doing some reverse engineering of the specific way that the display in the iPod actually shows things, some of its like compute capabilities and then also

33:24 figuring out how to transmit that data over the specific older iPod connection to display it on the screen of this device. So, this is going to be pretty difficult, and I don't know that this

33:36 will do it properly. I would imagine it's probably going to lean heavily on whatever open- source software does exist to do anything with iPods of this vintage. Nonetheless, see what happens

33:47 cuz I want to do something a little outside the scope. So, it's been quite a long time since I initiated the task of this turning this iPod 4th gen into a display that would just show telemetry

33:59 about the computer like CPU usage, things of that sort. Unfortunately, it never properly got it working and that is going to discount the level of effort and creativity that was displayed in

34:14 this test. So basically like this was a ton of just sitting here for hours power cycling the iPod for it and then having it try new things. It really I'm going to scroll up the entirety of this. I

34:25 exhausted my 5h hour limit twice just in this specific like test here and the weekly limit is up to like 41% use. So I mean look at the entirety of like everything it was doing here. It's

34:39 difficult to accurately like assess or at least to accurately demonstrate to someone who did not see the entire thing from start to finish, but it definitely did a lot of cool work. And really the

34:51 culmination of this because I think it's probably the best way to actually outline what specifically happened here is I'm telling it to make a beautiful front-end site in the style of Apple's

35:02 internal technical documents outlining exactly what we did. the timeline, what you figured out. Essentially, a concise debrief that also needs a 3D3JS mockup of this yellow fourth gen iPod. This was

35:14 a test that really would have put the micro reverse engineering capabilities and things to the test and there is a paragraph right here of like where I decided honestly should we call this how

35:26 long has it been? And it basically said yes. And here's the thing that is hanging it up. So, I'm just going to leave this on the screen. And just to kind of push back because I

35:40 know some folks will be like, "Well, you should have let it just go to completion." It got to the point where it was like, "I notice you have Arduino build tools

35:48 on this system. How about we like jump some of the pins between the cable for this iPod and then like the actual USB connection just to try to debug it on pins like 13 or 15 and 16." and it wrote

36:01 the Arduino script and I was like, I don't have cables that can like slot in there and hit it midway. And it was like, okay, well, we can upload it to the iPod and then once it's on battery,

36:11 we can just connect the pins. Unfortunately, the iPod has no battery like at all. So, as soon as it's unplugged, it dies. So, I tried it tried as well, like really getting some stuff

36:23 around, but this was just I think a bit more difficult. And I would say that I've had tests like this before where I believe I tried getting my Intel web tablet base station to communicate with

36:34 it, which required a lot of the same reverse engineering, understanding the protocols and things of this. I believe that was with GPT 5.6 Soul on Ultra and it unfortunately never got to a

36:45 functional result. So sometimes these harder tests don't always end like excitingly, but still. Oh, so my weekly quot is at 62% used. All right. So, here's essentially the entire thing that

36:58 happened that went on here for this reverse engineering task. One evening, one yellow iPod Nano believed to be a fifth gen proven to be a fourth. That's great. Thanks, cuz that was my mistake.

37:08 And a goal, make a 2008 music player display live vitals of the PC it's plugged into with zero hardware modification. Hours elapsed, power cycles, custom tools built. So, this is

37:18 where it starts to get interesting. Like firmware archaeology. Apple's CDN still serves N58 firmware bundle, bootloadader and diagnostics downloaded and decrypted on device then dissected. Modules map

37:30 display stack identified with a bit more information right here. Then we have payload engineering a position independent ARM V6 thumb module written from scratch EFI plumbing renderer

37:40 targeting Apple's display protocol and a pulled USB device stack answering control transfers wrapped in a handbuilt PE32 and sliced into Apple firmware volumes by custom Go tooling. We have

37:51 details the gauntlet. Okay. Roughly 20 power cycles, each returning one bit of truth. Each failure taught firmware's rules the hard way. FFS check sums live at header bytes here. Wrong offsets.

38:03 Corrupt GUID. Silent boot death. Every Apple image carries 16 byt something entry stub at this file offset. Without it, the boot ROM execute zeros. The boot ROM's DFU buffer tops out near 152

38:16 kilobytes. Uploads die at chunk 149. I recall when we found that out or when it found that out. The battery holds no charge. Even battery drain recovery was actually the watchdog or the user. So

38:26 basically just like me unplugging and plugging it back in. EFI dispatch order is positional. A payload spliced into the padding file never runs. The BDS app at position 28 seizes the machine. Final

38:39 blocker. Everything correct except dispatch. The DXE core starts modules from LZMA compressed sections only. or uncompressed variants boot cleanly but are ignored. The tool chains LZMA

38:50 encoder produces streams Apple's decoder rejects. And this is what basically was like, okay, we can't figure this out. Apple's exact encoder settings remain unreverse engineered. Colin restore

39:01 operator elects to stop. Okay, don't put that all on me cuz you were the one that was like before that it was opting to be like, you know, we could just like write down where we're at and go back to this

39:12 later. And this website is more just so I can better like explain the entirety of what it did. So right here, even serial debug console Arduino bridge on dock pin 13 would have replaced power

39:24 cycle experiments with a boot log worth building next time. It was just a very interesting and very in-depth reverse engineering task. This um visual right here is not necessarily up to snuff. But

39:39 everything else here is interesting, I think. Let me know down below. No. All right. So, that's that's going to conclude this test. I have to say I don't really have a wonderful takeaway

39:52 here because I find that basically a lot of the results we saw were just not as potent as I would have expected to see based off of the benchmarks to be completely transparent.

40:06 That doesn't mean they were bad. It just means that they were not really exceptional in the way that I had expected them to be. The thing that really was kind of very disappointing

40:15 and let's see some of these are just like uh folders that are not named what's actually accurately reflected inside them. So like for example the skate game that's probably in skate this

40:27 just was not good and the second chance we gave it which we actually did give it a second chance it made it worse because it actually just completely broke the functionality. So, I didn't expect that

40:39 and I was surprised at how poorly it did in that C++ task, which I would have figured this would have excelled at. Everything else was a bit simpler. The wrestling game was awesome, but the

40:48 problem was, and it probably won't load right here cuz I need to just run it in a server, but the issue was the odd orientation in all of the wrestlers, and that unfortunately kind of broke the

40:58 gameplay or really being able to do anything with it. Additionally, street heat. Again, it was just kind of like it was like, okay, these all have to be run in servers, so I'll just like summarize

41:07 them. It felt like with a lot of these specific types of results, it just wasn't trying very hard. And it was like, eh. And keep in mind, every single thing we ran here was on max thinking

41:18 mode. So, I can't imagine there's like an issue here, but you never can be 100% sure with new releases. So, I'll be very very interested to see what folks say about their first like Oh, it's still

41:32 working on this website, but that's okay. We saw enough of it just to get the information. I'm going to be very very interested to see what experience folks noticed with this model because I

41:40 want to see if it's reflected what I noticed as well or if folks are seeing significantly better performance. And again, I think in more tasks like this, not not this model, but like the

41:51 entirety of what it did here, it was very cool. It was smart and clever and it did a lot of like hardcore troubleshooting, reverse engineering, I suppose. So, overall, that's basically

42:03 where we leave off with this model. I don't 100% know what to make of it right now, and I'm interested to see what folks think of it. So, with that, let me know down below. No. Um, thanks for

42:14 watching. If you have any questions, please feel free to leave them in the comments and I'll have the Gemini 3.7 Flash video out probably Monday morning. And then there's Dots LLM3

42:27 by Rednote, which I did its predecessor and it was actually surprisingly hilarious. So, I'm very excited to test that one as

Frontier News · by Hyperjump Technology