xAI actually did it... (Grok 4.6)

summarized

TLDR

xAI dropped Grok 4.6, a point release that leapfrogs coding and knowledge-work benchmarks, making it a serious third contender alongside OpenAI and Anthropic. The secret sauce is the Cursor acquisition — Cursor brought the coding data, xAI brought the GPUs, and the flywheel is now spinning. The real headline: 4.7 is already training with SpaceX data, and this whole thing is a case study in why coding-first strategy wins in AI.

Key points

  • Grok 4.6 is a point release (4.5 → 4.6) but delivers massive benchmark jumps across coding, knowledge work, and legal use cases, tying or beating GPT-5.6 Soul on several evals.
  • The Cursor acquisition unlocked the missing piece: Cursor had mountains of high-quality coding data but no GPUs; xAI had 200k GPUs sitting idle because nobody wanted the old Grok model.
  • On cost-per-task, Grok 4.6 is cheaper than GPT-5.6 Soul at equivalent intelligence, and far cheaper than Claude Opus 5, though it moved up in cost from 4.5.
  • xAI is now using Grok 4.5 to curate training data for 4.6 (recursive self-improvement), and Elon claims 4.7 is already trained and will add SpaceX proprietary data.
  • Grokbot, a no-code, no-model-selection agent product built on 4.6, targets a broader audience than just developers — think PowerPoint, Word artifacts — and strips away all technical complexity.
  • Anthropic is currently buying compute from xAI because they underestimated demand; the author suggests that deal is temporary and xAI will pull capacity back to their own stack as soon as they can.
  • The coding-first flywheel (better coding model → more usage data → better next model) has now been replicated by all three US labs: Anthropic started it, OpenAI caught up, and xAI is accelerating fastest.

Tools mentioned

Techniques

  • Recursive self-improvement (using Grok 4.5 to generate SFT trajectories for 4.6)
  • Coding-first product flywheel
  • Agent-based knowledge work automation (Grokbot)
Transcript (captions)

0:00 Over the past 2 years, it has become clear there are only two main players in AI, open AI and anthropic. But within the last couple of days, a new player has emerged and I am taking them very

0:15 seriously. Grock was late to the AI party and they came out with a bang but fizzled very quickly and basically haven't been doing much since. But this has completely changed. It all started

0:30 with Grock buying this company called Cursor, which gave them the missing ingredient they needed to become a dominant player. But it all starts with today's news. XAI just dropped their

0:43 latest model, Grock 4.6. I've got the blog open here, and I'm going to go over everything that's changed and why it matters. Introducing Grock 4.6. This is a phenomenal model competitive with the

0:59 top models out of OpenAI and Enthropic. And so let me show you everything about it because it's actually quite unique in a number of different ways. So this is a DOT upgrade iteration building on Gro

1:12 4.5. So this isn't a completely new training run or anything completely new at all, but it is a vast improvement on Grock 4.5, their last model. and they continue to make it better at coding and

1:25 knowledge work in general. This is their focus and it all stems from the cursor acquisition which I'm going to touch on a number of times through this video. All right, so let's get right into the

1:35 benchmarks because it does perform incredibly well and a lot of people really liked Grock 4.5 including myself. It became the default kind of sub agent model in cursor and I used it all the

1:48 time. It was really capable, very fast, and now look, we got massive improvements from Grock 4.5 to Gro 4.6. Basically, across the board, huge, huge improvements. And you're going to be

2:03 pleasantly surprised. It got even faster. And if you're a speed maxi like me, you will appreciate it. Now first is GDP val which is a benchmark from open AI which measures a model's

2:16 effectiveness on actual knowledge work tasks. GDP val 4.6 high has the number one score beating out 5.6 soul beating out fable 5 max. We have cursor bench which is

2:32 obviously cursor's internal benchmark surprisingly did not get the number one score but basically equivalent with fable 5 max and next we have deep suite which in many people's opinions is the

2:45 most accurate benchmark to how people are actually feeling with a coding model because usually we have these benchmarks in which you know some model performs incredibly well but then when you

2:56 actually go and use it for coding use cases it doesn't feel quite right and we kind of gravitate towards another model. That other model has been GPT 5.6 Soul to me for a long time, even more so than

3:08 Fable 5 in a lot of cases. But what we're seeing here, Deep Sweet, is it is in third place. So 65.9 GPT 5.6 Soul Max at 73. That is currently my favorite model and 70% for

3:24 Fable 5. So definitely not quite as good as these other models. And on terminal bench, it jumps from 15% to 26% from the previous version of Grock to this new version 4.6. Then we have Harvey Lab,

3:36 which is legal use cases, and it actually dominated 15.8% as compared to 2.5% and 11.3% for Soul and Fable, respectively. So very, very good. But if you've watched any of my

3:52 videos over the past few months, you know a benchmark only showing a quality score is not sufficient. We want to know how much does it cost per task. That's just as important as quality. Okay, so

4:07 let's take a look at artificial analysis, which is really one of my favorite websites nowadays. This is the artificial analysis intelligence index, which takes a bunch of different evals

4:17 and basically measures against all of them. And what we see here is Gro 4.5 is here and we have this nice jump all the way to fourth place where kind of crazy claude opus 5 is number one then fable 5

4:31 number two 5.6 soul number three and tying with soul is gro 4.6 six high but again that doesn't tell the whole story. How is it doing on a cost per task? Let's take a look. This is the chart

4:45 that matters. On the yaxis we have the index score. This is basically the quality or the intelligence coming out of the model. But just as importantly we have the cost per task. How much does it

4:58 cost to complete the task to get the intelligence score? And the way you should think about this is really it's two things. one, how many tokens does it use? And two, what is the price of those

5:13 tokens? So, when we look at, for example, Kimmy K3 Max, which is half the price of Soul and Fable, the reason why we're seeing it at about the same price as Soul is because it used twice as many

5:28 tokens. So, if it's half the price and uses twice as many tokens, it's effectively the same price. That is why the cost per task is such an important factor to look at. Now here is what we

5:40 get. Here is Gro 4.5 high which is coming in at about 36 cents cost per task and an intelligence index of about 55 56 and we get this jump. Now unfortunately it does get more

5:54 expensive. Now it's at about like 83 cents cost per task. But we also got a jump in intelligence from about 55 to 60. Now hopefully this price continues to come down over time. But we can see

6:10 the other models that are around it. We have 5.6 Soul Max which is basically equivalent in intelligence but Gro 4.6 is less expensive. We have Kimmy K3 which is the same price and slightly

6:24 less intelligent. And then of course we have Opus 5 way out here. Very expensive. Definitely more intelligent but much more expensive. And let's look over here. GPT 5.6 Luna Max coming in at

6:36 like 5 cents per task completed and it's at let's say 52 on the intelligence index. Really where you want to be as a model is as high up and to the left as possible. This green quadrant is the

6:50 killer quadrant. This is where you want to be. And so, unfortunately, Grock got more expensive. It got more intelligent as well, but that's how it all came together. All right. So, I gave three of

7:01 the top models the same exact prompt to generate kind of a designy cody single profile card. And this is Fable 5. It is so bad. Look what they decided to do as the image. This is just like the worst

7:18 SVG Microsoft Paint thing I've ever seen. And look how much better this is. This is GPT 5.6 Soul. They actually generated this image of this woman. And just design-wise, it's beautiful.

7:31 Everything looks really good. It's clean. It's visually appealing. I mean, just look at the difference between these two. And what is happening with Fable? That is horrendous. And now

7:42 here's Grock 4.6. So, pretty good. Not great. It also generated this image of a woman. Kind of the same looking woman, which is a little surprising, but the following button is kind of cut off.

7:55 There's no padding at the bottom. It's a little bit weird spacing with the text and this line break. So, overall, it's pretty good. I would say GPT 5.6 definitely won. Definitely won. And I

8:10 I'm I'm embarrassed for Fable here. And I know the design benchmarks are hard to measure. It's really subjective, but I think it's actually pretty clear which one won. All right, so Grock 4.6 is

8:21 available today. You can get it in Cursor. You can get it in Grock Build, which are kind of competing products at this point, and they're probably going to merge eventually. It's also available

8:30 in the API and Open Router, Verscell, and Cloudflare. Now, what's important is the price. $2 per million input tokens and $6 per million output tokens, which is fantastic. And as a reminder, GPT 5.6

8:47 Soul is multiple times more expensive than that. And Fable even more. So, this is a very efficient model, a very inexpensive model, and I continue to say Grock is this workhorse model for my use

9:00 cases. And they even have a fast variant which is twice the price but it is much faster which is crazy because it's already very fast and they're offering 2x usage double usage for the first week

9:14 if you're using it in Grock builds and cursor. And the last thing is it is also powering Grockbot which is Cursor's latest product that they just released yesterday. And Grockbot just came out

9:26 yesterday and we talked about it already today in our newsletter forwardfuture.com. And so I think what's happening here is postcursor acquisition, Grock is really

9:37 excelling now. We now have a third major US AI lab. We have OpenAI, Anthropic, and now XAI. XAI certainly still feels like they're in third place, but they are really close and accelerating very

9:53 quickly. And it all stems from their increased focus on coding use cases and knowledge work use cases. That seems to be the winning formula that was originally implemented by Anthropic.

10:06 They just went so hard on the coding use case, built this incredible flywheel where developers used Claude code, all of that data fed back into Anthropic, all of the revenue fed back into

10:17 Anthropic, and then they use that coding model to train their next model. And this incredible flywheel was then replicated by OpenAI. And don't forget, OpenAI was the original large language

10:31 model company, but they fell behind for a little bit because they were focused on so many different things. And now they are so heavily focused on codecs and coding use cases that they've caught

10:41 up to anthropic. And it just seems like the winning formula for large language models for artificial intelligence is simply go hard on coding. Coding is the key. And we're seeing that postcursor

10:55 acquisition. So XAI acquired cursor. They acquired all of this incredible coding data that they are now using to train and post-train their models. And so the Grock model is now really good at

11:10 coding. And they even say it in the blog post. They used Grock 4.5 to train Grock 4.6. That's the key. Recursive self-improvement. We might not have a fully closed loop in self-improving AI,

11:25 but certainly current models are training nextgen models from all of the major labs. And we're seeing it here with Grock as well. We then use graph 4.5 to regenerate the SFT trajectories

11:37 across reasoning efforts, agent harnesses, and domains such as STEM, software engineering, and knowledge work. So the important thing to note is they're using 4.5 to help curate the

11:47 data used to train 4.6. So now we have another incredibly powerful coding agent with Grock 4.6. Obviously, we have OpenAI's models and Anthropics models, but what really sets Grock apart is it

12:02 seems like they're very focused on the broader market, not necessarily developers. And that's actually okay. They have Cursor, which is definitely a product made for developers, but now

12:15 they have Grockbot, which is very clearly made for a much broader, less technical audience. you still get all of the power of a really good coding model. It can still get realworld work done for

12:27 you because essentially everything's built with code. And so that's what is so special about this new Grock model paired with Grockbot. They've stripped out model selection. There's literally

12:39 no way to tell even what model you're using. They don't show any code at all. So you're just looking at text and artifacts like PowerPoint presentations, Word documents, everything that could be

12:52 output from knowledge work. Even the way that they have these little guys right here, these little mascots, each thread is an agent and each agent has their own little guy. I don't even know what to

13:05 call it. It's a little guy. So this Grockbot plus4.6 is very powerful. And if you haven't had a chance to check out Grockbot yet, I definitely recommend it. It's really cool, very powerful, yet

13:17 very simple to use. And it also tells more of the story of the transformation of XAI. If you remember a year and a half ago, Elon was talking all about how the Grock model was aiming to be

13:30 maximally truthful. And that seemed like almost an impossible bar to reach. They built a model, it was pretty good, but they quickly fizzled out. Anthropic and OpenAI just accelerated forward had the

13:45 best models in the world. All the developers moved to one of the other products and I think they got the message. The only thing that matters is the coding use case because the entire

13:54 market of knowledge work is basically built on code. But to better understand how well XAI is positioned today, you have to understand a little about the history of how we got here. Just a few

14:06 months ago in April, SpaceX acquired Cursor. If you remember, Cursor was really the company that changed the coding landscape. They were the first ones to deeply integrate artificial

14:17 intelligence into an IDE to allow you to code in a new way using AI. They were the first and then Claude Code and then Codeex. But when Claude Code came out, cursor fell behind. Everybody was really

14:31 choosing Claude Code and then they started choosing Codeex. So this acquisition was just so good because cursor had this mountain of coding data, just incredibly valuable data, but they

14:43 didn't have a data center. They didn't have GPUs to train their own models. And then you have Elon and the XAI company, which in 122 days built out 200,000 GPUs, which is just insane velocity. But

14:56 the only problem is they didn't have a good model. Nobody wanted to use the model. So all of these GPUs sat idle. And so all of a sudden you have cursor sitting over here with a mountain of

15:08 data, XAI sitting over here with a mountain of GPUs. And what happens if you merge them together? You get something incredible. You get to use all of that data powered by all the GPUs to

15:19 create the next generation of coding model, which is what we're now seeing with Gro 4.5 and now Grock 4.6. Now I want to point out something here. Anthropic partnered with XAI. Anthropic

15:33 has one of the best models on the planet, but they underestimated the demand for their models. And so now again, you have XAI sitting on so many idle GPUs, they need to sell them. They

15:47 need to actually use the GPU. So they partnered with Anthropic. Anthropic is now buying compute from XAI. And this is like selling weapons to your enemy. This must have hurt Elon so badly because he

16:00 has said so many negative things about Anthropic and then had to sell them compute. But now imagine you're Anthropic and you're sitting here today and you're like, "Oh wow, Grock 4.6 is

16:10 actually quite good." And if they start to garner a bunch of attention and a bunch of demand, I guarantee at the end of the term for their deal with Anthropic, XAI is going to allocate all

16:22 of that GPU capacity to cursor and to Grock. Why wouldn't they, right? So, I would be a little bit nervous if I were Daario and Enthropic. And it's not stopping. They continue to get better

16:34 and better. Elon just tweeted, "Grock 4.7 is significantly better than 4.6 six and should be ready in 3 to 4 weeks. The initial training is complete. Now they're adding a massive amount of

16:46 SpaceX company data in supplemental training. It will be something special. I am so happy about this. We need competition. Competition is good. Open- source model competition is good.

16:59 Frontier closed source model competition is good because all of that means better models, less expensive for all of us. And I'm really excited about Grockbot. I have been using it. I made an entire

17:11 video about it, breaking it all down.

Frontier News · by Hyperjump Technology