Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
Four new AI models hit the market in three days — Grok 4.6, DeepSeek V4 Pro, Nemotron 3.5 Lightning, and Muse Glimmer — each showcasing significant post-training improvements rather than pre-training scale. The 30B parameter range is now the most competitive spot for local models, with Qwen still leading, but Nvidia is pushing speed with speculative decoding. The sheer pace of releases means most of these models will be obsolete within a few months.
Key points
- Grok 4.6 from xAI leaped ahead of its previous checkpoint, now competing with Anthropic's Sonnet at lower pricing.
- DeepSeek V4 Pro showed dramatic improvement from its April release due to a new post-training run.
- Nemotron 3.5 Lightning from Nvidia is architecturally identical to Nemotron 3 Nano but uses multi-token prediction and distillation for 30% faster throughput.
- Muse Glimmer from Meta is a new ~30B parameter open model, marking a push from US labs into the open space.
- Qwen 3.6 27B from Alibaba remains the best-in-class for the 30B size, with a new checkpoint expected soon.
- Post-training is the key driver of recent model improvements, often overlooked in favor of reasoning scaling.
- The model layer is changing fastest in the AI stack, with new checkpoints every 2-4 weeks increasing supply.
- Consumer hardware typically maxes out at 30B parameters, making this the sweet spot for local AI.
Tools mentioned
Techniques
- post-training
- speculative decoding
- multi-token prediction
- DSpark
- reinforcement learning (RL)
- distillation
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
So, we have four new models in just 3 days. And for someone who does YouTube full-time, this is still a lot for me to take in. What we're seeing here is a growing supply of models where so many
new models are flooding the market. Looking at the AI stack, more and more labs are releasing newer checkpoints every 2 to 4 weeks, which means supply is increasing at the model layer. And of
course, at the same time, demand is increasing at the application layer as well. I'm sure you've seen this chart by now where it shows you so many people in the world still haven't tried AI. let
alone a coding agent which is depicted right here. So the demand here is still increasing all at the same time margin at the infrastructure layer is decreasing because more models are
getting efficient. It puts a downward pressure in cost so inference providers have less and less margin to work with. And of course at the chip layer the supply of compute is constrained both in
the US and China. Recently, Corewave signed an A100 chip to be used through 2029, which goes to show how a 6-year-old technology is still being planned to be used in the future. Even
at the energy layer, we're now going beyond terrestrial data center and going to orbital data center as well. And out of this entire stack, the model layer is changing the fastest. In the past 3
days, we have Gro 4.6 6 from XAI, DeepSync V4 Pro Checkpoint, Neimotron 3.5 Lightning from Nvidia, and Muse Glimmer from Meta. And the first thing we need to look at are these numbers
right here. The numbers here represent a checkpoint, which means as companies improve the model through training in their lab, they decide to release a specific checkpoint to the public. And
the type of training I'm referring to here is post-training. We all know by now that training AI models is split largely in two phases, pre-training and post-training. And we're witnessing a
huge improvement being made right here in post training. Grock 4.6 leap frog from their previous checkpoint Gro 4.5 at an incredible leap. Deepseek V4 Pro, same thing in comparison. And Neimotron
3.5, the same story here. This kind of leap is made possible because of optimizations we're making in the post-training stage, which is kind of crazy. Post-raining has largely been
forgotten from public narrative ever since reasoning became the next scaling factor in intelligence. And yet, we're reminded that the potential that's inside the base model after pre-training
still has a lot more room to scale in post- training where this potential is actually realized. For Grock's case, the model has now entered into Fable's territory. Looking at the benchmark, XAI
said that they added additional pre-training with higher quality data and overall better post-training and RL to get the result that they got. The same thing for Deepix's new checkpoint.
And Deep 6's case is actually more dramatic when you look at the numbers compared to the model they released back in April. The flash and pro model showed an incredible leap in performance in
comparison following their new post-training run. Nvidia is the same here where Neimotron 3.5 lightning. The model is architecturally identical to the Neotron 3 Nano from December, but
with a specific pre-chain it added to include multi-token prediction. And Nvidia distilled a bigger Neotron 3 Super that's four times bigger. pretty amazing results in checkpoints looking
at the benchmark. Another angle worth considering here is when it comes to intelligence. Frontier or flagship models tend to have less constraint since their ultimate goal is to get
maximum intelligence almost at all costs. But when you go further into this mid-tier intelligence here, these models are often bound by intelligence per compute available. Emphasis on compute.
Typically, when frontier models achieve a huge leap in intelligence, they can often charge higher pricing until the rest of the industry catches up. We've seen this playbook with OpenAI1 and
later Enthropic Fable 5. However, in the mid-tier intelligence here, there's not a lot of room to hide since people often run these models locally. And the question is more about how much
intelligence can you pack into the model per compute available, which is why you have to shop in MicroEnter, who's sponsoring this video. August means we're nearing back to school, which
means we have some shopping to do. Microenter currently has back-to-chool tech deals running, including your hardware needs like monitors, SD cards, and printers. You know, something that
you might need for the rest of the year. Of course, I always look for hardware deals on graphics cards. And sure enough, there are deals on workstation grade GPUs like Nvidia Pro series. And
these are serious hardware that you can buy and run AI at home. Besides specific deals, Microsenter has a huge selection of GPUs from Nvidia cards that's meant to run at home. and AMD graphics cards
that are also listed here. As you can see, as AI demand continues to grow, shop at MicroEnter for all your computer and computer part needs. And right now, if you're near Austin, Texas, you can
sign up for a free 128 GB flash drive to redeem it in store when they open later this year. And if you're near Columbus, Ohio, the store is also getting remodeled and you can sign up for a free
128 GB flash drive on their reopening later this year as well. Microenter also has valuable blogs around AI, so you can check out as you experiment on your build. link in the description below.
And that's the difference you see between Gro 4.6, DeepSc V4 Pro and Neotron 3.5 Lightning and Muse Glimmer. For example, Grock 4.6 and Deepseek V4 Pro have the same strategy which is
offering frontier level intelligence at lower pricing. Grog 4.6 is not only cheaper than every other model in the same class by more than 50%. Its pricing is comparable to Anthropic Sonnet 5 at
the moment. Deepseek V4 Pro isn't at the frontier level yet, but the goal is the same in that they are aiming to bring the cost down. Now, on the other hand, models like Muse Glimmer and Neimotron
3.5 Lightning are fighting a similar, but certainly a different game. For example, I can run both Neotron and Glimmer here at home using the single DJ Spark. But the question that I'm going
to be asking is very different, which is how much intelligence can I fit into this thing since I'm working under this compute budget. Now, both of these models are around 30 billion parameters
in size. And out of the entire spectrum of models and sizes, 30 billion is one of the most competitive bands. That's because most consumer- grade computers top out at around 30 billion parameters
in support. And anything more than that means it's out of reach for most people. And at the same time, making the model smaller means getting worse performance. So 30 bit parameters is a sweet spot at
the moment. And despite heavy competition in this range, Alibaba is still the best in terms of how much intelligence they're able to pack given compute available. Quan 3.627B is a
great example. This model is already 4 months old and is still one of the best models given the size even compared with Neotron 3.5 Lightning and Muse Glimmer. And Alibaba is expected to release the
next Checkpoint Quan 3.8 tomorrow, which means it's expected to raise the bar even that much higher. But let's not sleep on Nvidia just yet since Neotron 3.5 Lightning shows its strength when it
comes to throughput. This is me running Neotron 3.5 Lightning here at home with DSpark on and the model is pretty fast. Nvidia is really doubling down on throughput by using speculative decoding
where they support multi-token prediction, DSpark and Dlash. Now, the technical detail on these weren't a video on their own, but I covered briefly on MTP in my previous video
explaining Neotron 3 models if you want to check it out, but according to Nvidia, they can achieve similar accuracy on Pinchbench while also getting the job done 30% faster than
comparable models like Quen 3.635B. Even looking at artificial analysis benchmark for intelligence index versus speed, the output speed is incredibly fast, nearly four times the speed. though it falls
behind in intelligence to Quen 3.627B, Quen 3.635B and Gemma 431B. Now in this same band of models, Muse Glimmer is also around 30 billion parameters in size, making the compute budget similar
to Neotron 3.5 Lightning. And the model still measures a bit shy of Quen 3.627B as you can see. But this is still certainly a moment to celebrate given that more US labs are now competing in
the open space. Recently, Thinking Machines released their first open models, Inkling, and Nvidia with Neimatron and now Meta pushing more models covering a wide range of sizes
here in the US as well. And this growing competition is really helpful all around as long as funding continues to grow. And ever since Zuckerberg rebooted their AI lab back in 2025 by acquiring Scale
AI, we are now finally seeing new models being dropped from them. Muse and Muse Glimmer being introduced in the model layer, further increasing the supply of models. We really are undergoing one of
the most incredible transformations in just how fast AI seems to be improving. Chances are all these models that you see here in this benchmark will likely be gone in 2 to 3 months. Which means
the video that I'll be making in December, you might not even see any of these models on this list in this benchmark. Especially as we wait for models like GPT6 and Gro 4.7, which is
supposed to be dropping shortly.