GPT-6 Luna First Test – Is OpenAI’s CHEAPEST Model Actually Good?

summarized

TLDR

GPT-6 Luna is OpenAI's cheapest model at $0.10 per million input tokens and $0.50 per million output tokens, and it consumed zero usage during an extensive test session. Its performance is competent but not state-of-the-art, making it a practical choice for repetitive or low-intelligence tasks where cost efficiency matters more than peak capability.

Key points

GPT-6 Luna costs $0.10 per million input tokens and $0.50 per million output tokens.

The model has a context window of over 1 million tokens and 128,000 maximum output tokens.

It supports text and images but not video or other modalities.

Its knowledge cutoff is May 18, 2026, more recent than other GPT-6 models.

On the Deep SE benchmark, it scored 66.6% on max effort, up from 62.2% for its predecessor.

Tools mentioned

Transcript (captions)

0:00 Hey, can we win this? Come on. Yeah, Tiger Heart survives the three-way. The ring is yours. Run it back or pick a new legend. Okay, today we're going to be taking a look at the final GPT6 model

0:10 that we have not yet tested, and that is Luna. This is the smallest, cheapest version currently in the GPT6 family. And really, there's not a ton to say about this model. It's a replacement for

0:21 GPT56 Luna, which was also the cheapest in the 56 family. Really, the biggest takeaway here is the pricing. This model is incredibly cheap to use. It's 10 cents per million in and 50 cents per

0:33 million out. This is basically one of the cheapest models I think out there, period, especially from a company like OpenAI or Anthropic. So, while the performance of this model is likely not

0:45 going to be anything that blows us away, it's definitely worth testing because this represents an affordable option that still comes from a Frontier Lab. So, before we get into it, do feel free

0:54 to subscribe so we can get the 100K plaque. And let's start out just by taking a brief look at GPT6 Luna and then we'll jump into some fun and traditional testing. Now really they

1:03 talk about and basically every benchmark chart right here alongside what is spoken of this model shows it represents a slight increase over its predecessor. So right here on Deep SE we can see on

1:14 max effort it scored 66.6% whereas its predecessor was 62.2. And really the common theme in all of these charts based off of its predecessor is it does not seem like a gigantic leap in

1:26 capability. However, if it is slightly improved but significantly cheaper, that's still pretty good to see as that is the direction that we want things to be trending in cheaper and better. So

1:36 for a bit of additional technical information about this model, we can see right here it has a little over a million token context window, 128,000 maximum output. It is multimodal, but it

1:46 only takes text and images and it does not take video or anything else like that. Now, something that I do find to be kind of curious here is this model has a more recent knowledge cutoff than

1:56 GPT6 Astra and GPT6 Soul. We see right here, this is May 18th, 2026 for this knowledge cutoff. And if we look at either of those, we see it's April 20th. So, okay, maybe just like a month or a

2:09 couple of weeks in the case of Astra later, but still interesting to see that this is technically the most upto-date model in the six family when going solely off of the reasoning cutoff. Now,

2:19 there really isn't much else to say about this model, so we're going to let the results speak for themselves where we can see here we have initiated this with the browser OS test v2.7. I would

2:28 like to make note of a couple of things. One, we will be running this exclusively on max effort today. that is its highest potential reasoning effort. And in the benchmark charts for this, it always

2:38 scored the best on the highest reasoning effort. Additionally, the browser OS that it just did here, it did not take 45 minutes to do this. I believe probably 40 of those it was stuck on a

2:48 permission prompt that I needed to approve. So that is not reflective of the actual time that was taken to produce this result. And before we go further, let's take a peek at our usage

2:57 starting out this test. So all that's been done so far is the browser OS test. And this shows 100% left in our plan and we'll see what that goes to following the conclusion of all of our tests. I

3:08 would not really expect that to move very much given how cheap this model is. So, let's take a look at our GPT6 Luna browser OS. Okay. Uh first impressions, not bad. The calculator icon here as

3:19 well as some of this stuff perhaps gives indication of some simple code issues. I do actually want to check real quick. Okay, so yeah, we'll just see what we get. So far, we have no right click, but

3:31 we do have what is supposed to be a moving wallpaper. Now, I do believe the issues that we're noticing here are perhaps preventing the wallpaper from loading and maybe additionally some

3:41 other things. Let's just check. Yes. So, unfortunately, this is actually not a fully functional result. Not the best out the gate, but we have a lot of other testing to perform as well. All right.

3:51 Supposedly, in a little under 2 minutes, we have a fixed result. So, I suppose we can just reload it right here. Good. And we now are met with our moving background, which is what we wanted.

4:01 Let's see. Is there now a right click? Okay, there still isn't, but the background was fixed. And our calculator app is still funky, but we'll let that slide. We have the clock in our local as

4:11 well as a battery indicator, Wi-Fi indicator, and then our living wallpaper. No right click. So, let's see. Ah, interesting. Your browser, your little world. Okay, I think this is just

4:23 like a find an app. Oh, cool. And it opens our mail app. I guess we'll take a look at that real quick. First and foremost, it's actually pretty clean in terms of UI. Nothing special, but

4:32 nothing horrifically off here. And again, just matter references to specific apps in the browser OS. Next up, let's look at our notes app. We'll just go through these in the order that

4:41 they appear. The little desktop is yours. Take a drive around Neon City and then write down what you found. And it autosaves again. Clean. Nothing special, but nothing bad either. Calculator with

4:52 the funky icon. Okay, that is a very, very odd UI. Can we full screen and unfull screen? I wonder if minimize and Okay. Next up, our GTA clone. Yes. All right. So, this is one where we're

5:05 definitely going to start to see some of the limitations of this model. Being that this is the smallest, well, we don't specifically know the size of it, but we can assess this is the smallest

5:15 in the GPT6 family. And for the price, this is probably part of the course for what we would expect. We can't get out of the car. There is a nitro button. So, when we press shift, it perhaps makes us

5:26 go a bit faster. I think there is a a package there we're supposed to collect, but Oh, good. Okay. Well, the logic did work for that. Now, I don't know how to close

5:38 this. Okay. Next up, we have our orbit run game. So, we can't actually shoot. We're just supposed to dodge these things.

5:50 Interesting. This almost reminds me of like Guitar Hero the way you would see that UI. All right. Again, we can't actually close it, so we just have to refresh.

6:00 Focus orbit. This seems to be our special feature where it's just a focus timer. A Pomodoro timer that gently shifts the living wallpaper into a quieter pallet while you focus. Finished

6:09 sessions add to a local constellation on this device. Makes focus a part of the desktop's atmosphere while keeping your activity in private offline. Then files, um, fairly simple and it just shows us

6:20 our apps right there. Terminal, I believe, is the last one that we have not looked at yet. All right, overall kind of what was expected for GPT6 Luna. Let's now move

6:34 on to something a bit more taxing. Next up, we're going to do the self-contained C++ skate game. This is one where it cannot use Ray Lib and it has to have the early 2000s New York City block as a

6:44 map. I think this is something that will probably stretch the limits of this model. So, I'm more going to be focused on functionality rather than aesthetically polished results. So, in 9

6:54 minutes, it completed this task. And something I find interesting that's very infrequently, if ever seen, is it didn't actually compile this to an executable. It gave us the specific command to do so

7:05 ourselves. But in the prompt right here, it says requirements should compile and run as a single program, all code in a single file. So just interesting right there that it omitted doing that. All

7:15 right, it did compile without error. So now we can take a look at it. And again, I'm more that's actually significantly significantly better than I expected. So the question becomes, is it just

7:26 actually pretty good or is it perhaps familiar with this specific test? Fortunately, we do have some ways to uh work around that just by giving it a different C++ task, which is definitely

7:38 going to be warranted here because again, this is like pretty simple. But the actual design of the block and things like that is actually acceptable. Everything's clean. I don't see anything

7:48 that's wildly like off or not right here. The forward and sideways movement seem a little bit off. However, okay, K is a 360 and it did do that land clean to score. I'm actually more

8:05 impressed with this than I had expected to be, especially considering the browser OS result. I think this is a significantly better result than what we got with the

8:14 browser OS just in terms of what was expected from this model. I can't imagine we can interact with any of the pedestrians. Even the fire hydrant shooting water does actually

8:25 exist. Okay, no, there are mesh colliders there. This is definitely worthy of having some additional feedback given to it. I'm just giving it a pretty simple piece of feedback.

8:33 Basically saying make it better and we'll see what we get. But truth be told, that was a lot better than I had anticipated it to be. So, in 7 and 1/2 minutes, it seems to have gone with

8:41 basically like a synth wave aesthetic with a rainy dark look and neon streets. Okay. Do we have sound? We do. All right, it Hey, the cabs are actually moving now. All right, so this again is

8:59 really not bad. I don't know how large this model is from like a parameter count size. I have to guess it's probably not very big. So, I'll say this is actually a respectable result.

9:09 Everything kind of works very well. It's simple, but it's not bad. Aside from like the board having a some oddity to it, like the top of the board seems a bit separate.

9:22 Even though the rain is going backwards, we're just going to ignore that. See, like that was a sick trick. And the character model's not bad. There's even a little bit of movement when we jump up

9:33 and down. We can see that. So, acceptable. Next up, we're going to do the subway FPS. This is the newer version where the subsequent waves following wave 1 need to be delivered by

9:48 a subway train. This is one that will probably push at least one or two times beyond the first result, as I'm very interested to see what level of graphical detail and custom looking

9:58 enemies we can get from this model. This is 3JS as well. All right, in 12 1/2 minutes, we have the first result for our Subway FPS. I had sent it this follow-up question mark because it

10:08 seemed like the at least what it was showing it was doing had frozen. So, it looked like it was searching a folder for 5 minutes, which is why I sent that. And then it kind of snapped into its

10:17 current progress. So, basically, we can just ignore that. Now, let's take a look at our Subway FPS. Okay. Okay. Good. There is sound. It is quite dark and quite simple. Do we have

10:32 other weapons? Yes. However, they're tough to see. Okay, so that's inevitably a shotgun. Cool. All right, and here comes the train. Nothing really crazy. A lot of

10:49 times models will put like a red flashing light on the ceiling or something. The doors do open. Let's see how the enemies appear. Okay, so they just kind of spawn there. This is pretty

10:58 simple. Again, I don't see anything that's horrifically wrong with it. It's just quite simple. So, we'll have this do at least one follow-up on this. So, we'll send it a follow-up here to make

11:08 it much better. Doing a full visual overhaul, making it brighter, better models for everything, specifically the enemies, bullet holes in the environment, fine detail, and high

11:16 quality. All right, so in about 30 minutes, it completed the overhaul for this. That took quite a bit of time. And we can see here, first and foremost, the big thing that was a problem was the

11:25 station was too dark. Okay, we have a bit more visibility at least into some of the things. The other request I had was that there should be bullet holes in the environment. And I will say it

11:36 actually did a pretty nice job of implementing that. So far so good. I'm quite satisfied with what I'm seeing here. And our enemies are significantly

11:47 improved. I would say the fact that it really went above and beyond to make these look better, not just simple shapes, bodess well. This is something again that like 6 months to a year ago,

11:57 this is a cheap cheap model. like really cheap. This would have been absolutely mind-blowing state-of-the-art as a result. So, it is cool to see the capability

12:08 get better while the price goes down at least, you know. But then the problem is like the models that are state-of-the-art now make this look outdated. But still,

12:19 can we put bullets in the train? Yeah, we can. I have to say overall, I'm actually satisfied with this. This is something that definitely you could spend a lot more time on it. And I would

12:28 go out on a limb and say you could probably make this pretty darn good looking. All right. And let's see what happens if we lose. All right. And I had clicked in so it

12:41 restarted. Overall, not bad. And I just wanted to see what it would do if we pushed it a bit more. Next up, we're going to be doing the robot arm test where it has a camera attached to it as

12:50 well as a robot arm. And it is tasked with moving the orange Hot Wheels car to the other side of the mat. Unfortunately, we did not get a successful result for this. I did run

13:00 this on max effort, but I also enabled fast mode just because this is a task where speed is definitely a factor in completion. Now, something interesting I noticed was to actually get the camera's

13:12 view. It basically made a little web UI to be able to look and see what the camera is doing. I've not seen a model do that before when doing this test. Normally, they just say, "Okay, I'm

13:21 going to check what USB device this is." And then they pull frames from it in their own way. But the fact that it actually made a little web dashboard to be able to get images from it and did a

13:30 little computer use to do that was just interesting and not something I've seen before. And unfortunately, it was not a successful completion. We can see some of what it was doing and the way it was

13:40 getting tricked because that image looks very hard to tell. Okay, the claw is right over the car. So, why is this not working? And that is a limitation of this test. the angle of the camera and

13:50 there being one camera. Something I do on purpose because the models that can actually successfully perform this task with the static camera angle are very very impressive all things considered.

14:00 And I let it run for a little bit longer, but unfortunately it started moving the arms in ways that could perhaps have been destructive to the frame. So, I stopped it there. Overall,

14:09 I didn't really expect it to perform successfully, but it was just interesting nonetheless to see what it did. And just as a midpoint checkup, our usage is still on 100%. So everything

14:18 we've done up until now has not actually made any difference in that usage. Next up, we're going to be doing the 80s wrestling game. And this is the one where it has to create the assets in

14:27 Blender and then make the game using GDAU. So we'll see how well it can actually get some polished visuals when given a more powerful tool like Blender. All right, in 30 and 1/2 minutes, we see

14:37 we get our first spoiler look at this. I'm just going to try to ignore that. I don't know what the heck happened there in terms of where it put it. I created a wrestling folder for it, but it made it

14:45 in the output directory. So that's probably my mistake. All right. All right. It's quite simple. However, it's not badly done. It's just simple. It did make

15:03 these models itself, and they're really not bad. One of the bigger things that's a bit disappointing, but truthfully expected, is that they don't have any individual rigged movement or anything.

15:13 Something that I will say though, I'm quite disappointed in is there no sound effects aside from the music. So when you turn that off, there's no sound effects of the wrestling or when they

15:22 make contact or anything like that. It's got a big like springy feel to it. So Oh, okay. Thunder drops Tiger Heart. Run it back. And we have our commands right there. It's gone for

15:36 kind of like a cartoon. Actually, no, that's cuz I said '8s, so it just went for like 80s colors. That's on me. However, again, it's nothing really special. Hey, can we win this? Come on.

15:46 Yeah. Tiger Hart survives the three-way. The ring is yours. Run it back or pick a new legend. Okay, this was just so simple. It wasn't bad. It was just I I don't

15:59 really have a lot to say about this. What a suplex. All right. Again, it's just really simple. Next up, we'll try the watch website front-end test. This is

16:12 also the one that needs to have an exploded view of the watch, as well as with this photo, the special edition dubbed the Bjan, which needs to have this photo placed on the dial as well as

16:22 an aesthetic style that is inspired by this photo right here. So, we'll see what we get for this one. All right, I think this is done, but I have not like the laptop screen is bent, so I'm going

16:32 to open it. Okay, good. Because I knew it would have opened it in some preview here, so I didn't want it to be ruined. Oh, interesting. It is still working here. It's been around 17 minutes or so.

16:42 I'm going to just allow it to keep working then. Although, we probably see a good indication of what this is going to look like when it's concluded. Oh, good. All right. This isn't great, but

16:51 it's not horrible. The numeral markers are all messed up, but they are symmetric with one another. So, that's a good thing at least. And the surface it has chosen to place this watch on has

17:00 like, you know, some luxurious elements to it, I guess could be said. Now, the blatant issue here is there are no hands for the watch. So, we have no way of knowing what time it is, and there seems

17:12 to be some messed up geometry here and things like that. Let's just look at this and judge it maybe based on some front-end ability beyond just checking these 3D models. Okay, we have Esther

17:23 and Mayor view in 3D. I think that was the one it loaded in with. Let's check this one. Okay, good. I do like that where you can view a different one in 3D. And this one has some additional

17:32 little dials in it. None of which seem to have any hands. So that is definitely somewhat of an issue now. Oh. Oh yeah, I forgot about that. We also have the exploded view.

17:46 Good. Okay. And there are hands. It's just that they were hidden. So we can see them appear right there. That's something that happens pretty often where if something seems to not be

17:54 there, it likely is, but it's just covered by some surface. So again, this is pretty darn simple. And then we have the Bejian. Okay, I do have to say view the Bejian in 3D. I'm

18:06 actually quite a fan of the aesthetic that this chose for this. It definitely captured the Lacost tracksuit looking um pattern on the strap here. And then it did go with a gold look. So overall,

18:17 that one's actually quite nice. And again, this test is subjective. So that was definitely a nice surprise. Overall, this is pretty mid. I will say I'm not going to do a

18:28 follow-up to have it make this better because I just honestly throughout the entirety of this test, the one thing that keeps popping up in my mind is I would love to know the size of this

18:39 model. I'd like to try another C++ game because the Skate one it did was actually quite well done considering how the other results have looked. This is for a retro rally 3D game. It needs to

18:49 have a detailed interior of the vehicle, a low poly graphical style reminiscent of early Rally games. I have told it it can't use RayLib. And additionally to the detailed interior of the vehicle, it

19:01 also needs to have a chase cam and some sound effects as well. So, we'll see if the C++ capabilities seems to translate into a different prompt. All right. In 23 minutes or so, we received Dust Line

19:12 Rally. Let's see how it did. Okay. Whoops. That was a Freudian click, I think, just based off of seeing the car from the chase cam. All right. Yeah, this is

19:24 um it's it's I'm not actually steering so it just snaps it to the track. I will say the elevation in the track and things like

19:41 that are not actually bad. Do the gauges work? Okay, no they don't. But to put this together using C++ without RayB again is something that six months to a year ago would have absolutely been

19:54 mind-blowing. And now for a model that is this cheap to be able to put this out in 23 minutes is not horrible. So we uh Oh, okay. Loose gravel. And it just seems like perhaps some things aren't

20:07 necessarily rendering that would coincide with the elevation of the track. So it just looks like the track is floating in midair. As with pretty much every other result, I don't have a

20:16 whole lot to say about this, but I wanted to try another C++ result because the Skate one was surprisingly nice to use, and I wanted to see how well that translated over into something that's

20:26 perhaps less commonly seen in these model tests. I think overall that's probably going to conclude our first look and test of GPT6 Luna. Definitely a shorter video than normal, but I wanted

20:38 to just take a peek at this. Keeping in mind it's probably not as exciting to test as some of the big bad Frontier models. However, the pricing of this model makes it pretty darn attractive,

20:48 especially for something that's perhaps more repetitive, less necessitating high intelligence. And we can see right here, basically the culmination of this is our weekly limit did not move once. So, for

21:00 every single thing that we did in this video, not 1% of our weekly usage was removed. So, this is a very efficient model. The results are definitely not state-of-the-art. However, it's fully

21:11 competent. I didn't notice anything that was blatantly wrong with it. If we were to do a brief results overview, it's ironic that the first thing and the only issue we had with it was the original

21:22 browser OS result had some issues, which it quickly fixed, and everything was pretty competent, but nothing was really exciting. The watch website was kind of part of the course for what you'd expect

21:33 where some of the things were good, but it made simple mistakes like not putting the watch hands in and things like that. Though everything was just a good basis for improvement. And that's kind of my

21:44 takeaway with this model is it makes decent foundations, especially considering the price, but it's not going to wow you in terms of its aesthetics or style. I will say that I

21:53 think the winner of today's video, and keep in mind this is our second iteration on it, but not because it was problematic, just because I wanted to see how it would increase the fidelity,

22:03 the winner is definitely probably this skate game result. It was actually surprisingly together and competent. Yes, a far cry from the state-of-the-art results we're used to seeing. However,

22:14 considering the price of this model, it was absolutely a proper result. None of the text here was weird. The storefronts looked good. The textures of the stores were actually really nicely done. And

22:24 while simple, it was nicely put together and kind of cohesive in terms of its style. So, I was very, very happy with this. This was probably my most impressive result. Well, not mine, but

22:35 GPT6 Luna's. Also, we had it do a Blender and GDAU test in the form of the wrestling game. I'm actually going to have to find that one because it didn't put it in the folder. So, this was quite

22:45 simple. Really, the big thing that upset me about this was there were no sound effects outside of the music. So when you actually made contact with the other wrestlers, it didn't make any sounds. It

22:54 was simple. It made these models in Blender, put the game together in GDAU, and it did have some playability to it. They were basically just like chest bumping each other, and there was no

23:03 real rigged movement. But for this model and this class and hypothesized size of this model, that's totally acceptable. The Subway FPS was not bad. It's running it here in its own little browser

23:15 preview, but I think that's a much older version that we're seeing there. So, this did a much nicer job the second time. And I did like the implementation of the bullet holes. Of course, that was

23:23 something I denoted in the follow-up prompt that it should have. But even the shotgun actually has correctly like done bullet holes. So, and these 3D models that it created were

23:35 significantly better than the first one. I think this is something that you could definitely get a decent passable 3JS game result out of this if you worked with it for a while, which really would

23:44 make sense because it does not use any usage really. It's it's very very efficient. So overall, that's going to conclude our first look and test of GPT6 Luna. I have to say it seems like a

23:55 decent model and really it's tough to judge this because a lot of what's in recent memory is things that are jaw-dropping done by Opus 55 Fable GPT6 Astra. However, those models are $20,

24:08 $30, or $50 per million output tokens. This is 50 cents per million output tokens. So, the core takeaway here is this may not be a jaw-dropping model in terms of its aesthetic, but it is so

24:19 much cheaper than some other models, Frontier ones, that if you have a task that's repetitive and does not need that higher level of intelligence, this could be a very potent option. And it's

24:29 definitely worth keeping an eye on some of these cheaper models as well because they're not as exciting visually to test. However, they are more exciting when it comes into the pricing. So,

24:40 overall, that's going to conclude our first look and test of Luna. Again, it our usage did not drop at all. So, it stayed at 100% throughout the duration of everything we did in this video. And

24:50 I think that's probably the biggest takeaway here is pricing and efficiency. So, if you have any questions, please feel free to leave them in the comments. And thanks for watching.

Frontier News · by Hyperjump Technology