Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
GPT-6 Astra's 99.9% score on ARC AGI 3 came from OpenAI's own harness, while a neutral evaluation using Arc Foundation's harness gave 62.7%—a stark reminder that benchmark wrappers can dramatically inflate perceived intelligence. This doesn't erase the model's real achievements on other private benchmarks, but it does mean aggregate intelligence scores should be taken with caution.
Key points
Only one benchmark overlapped between OpenAI's announcement and the artificial analysis intelligence index.
GPT-6 Astra scored 99.9% on ARC AGI 3 with its own harness but 62.7% with Arc Foundation's harness.
The model achieved 97.6% on Frontier Math tier 4, surpassing Fable 5.1's 87.8%.
GPT-6 uses half the tokens of GPT-5.6 Soul for similar performance but costs more per token.
Tools mentioned
Techniques
- harness-based evaluation
- token efficiency
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
When we look at OpenAI's announcement for GPT6 Astra, out of the 14 benchmarks that they highlighted in their blog post, which can be represented here like this, only one of these benchmarks
actually overlaps with artificial analysis intelligence index, which has its own set of benchmarks that it tracks. Now, the problem isn't in the benchmarks themselves. In fact,
benchmarks serve an important role in measuring domain specific capabilities, and we really need them. The problem is the wrapper that sits on top of these benchmarks that tries to demonstrate how
intelligence should be measured. And every time we create an abstraction on top of these benchmarks, we're essentially creating a wrapper. Much like in programming where wrapper
functions job is to abstract the details of the sub routines implementation. And just like how we can have a really bad wrapper and a really good wrapper, composite benchmarks can also help us
focus on the right signal or become super misleading due to the noise underneath. And the fact that artificial analysis intelligence index puts GBT6 Astra in the fifth place which puts the
model behind Muspark 1.3 is a strong indication that artificial analysis intelligence index is a bad encapsulation of intelligence to extend an olive branch. That doesn't
necessarily mean that OpenAI's curated list here of benchmarks is also a good measurement of intelligence either. So in order to understand the true significance of GPT6 Astra, we have to
dig a little bit deeper into key benchmarks underneath and find out how Astra is actually making meaningful contributions in AI. One of the first benchmarks that I'm sure we are all
surprised to hear about is ARC AGI 3 where GPT6 scored 99.9% which practically saturated the benchmark. But hold on, didn't Nvidia recently also score 100% on ARC AGI 3? If that's true,
what's really the big deal here? ArcGI3 is revered as this highly abstract benchmark where you place the model through a series of games without giving them specific rules on how to actually
solve them. And as much as this looks really fun to play, the real evaluation is actually far from this prettyl looking UI. The evaluator from Arc Foundation sends a text prompt that's
64x 64 grid in text where every cell is one of 16 colors and the model receives that grid as a JSON object. So yeah, no screenshots are sent. Certainly not video, but a simple text object for
OpenIGBT6 to actually interact with the simulated environment and ARCI3 will send back the updated state in the greet format as the model sense action tokens similar to how VA's work and their data
set is split into three different ways public, semi-private, and fully private. And Nvidia scored 100% recently on the public data set that's available for everyone to see which is far different
than what GPT6 Astra scored on the semi-private data set. More specifically, Nvidia's contribution here was AVO, which is an agentic layer that serves as a harness for Opus 5, which
gives Opus 5 the ability to have persistent memory, context management, and execution environment. You know, something that every harness does these days. Now when we look at what GPT6
Astra achieved, their evaluation was on a semi-private data set evaluation where the environments themselves are hidden and procured by Arc Foundation but with the potential for data leak since it's
making external API calls to labs like OpenAI to evaluate. It means that the environment can still be mapped out and therefore not guaranteed to be private all the way through. But at least the
data and the environment for both semi-private and fully private are hidden from the public. It just depends on where the model is actually hosted for the evaluation to happen. So using
the harness that Arc Foundation hosts in their environment and only using Astra model through API, the model scored 62.7%. And when OpenAI used their own harness, meaning with its own context
management, the model actually scored 99.9% which is exactly what they announced. Now, this actually raises a really critical point here between practitioners like us who often use
models through specific harness versus people who are more curious about pushing the science forward. Benchmarks like ArcGI3 don't really carry a lot of impact to practitioners unlike other
benchmarks like Deep Suite and Exploitbench given that these benchmarks actually reflect more realistic situation that the model has to solve their way out of. So we have to be
careful not only how aggregate scoring might lead us to faulty conclusions but also how the underlying benchmarks actually measure different aspects of intelligence whether that's scientific
endeavor or actual usefulness in reality. Now ARKI3 isn't the only benchmark that GP6 saturated. Astra cleared Frontier Math tier 4 by scoring 97.6%
and also scored 100% on exploit bench essentially saturating them all at once. Let's look at the Frontier Math benchmark. What if one way to evaluate the model's intelligence is by grouping
more than 70 human mathematicians and have them write difficult math problems ranging from easy at tier 1 to extremely difficult at tier 4. This would cover anything from linear algebra and group
theory in tier 1, combinotaurics in tier 2, number theory and density in prime in tier three, and whatever this is in tier 4 that's probably beyond what most of us can comprehend. And just like ARC AGI 3,
Frontier Math is proctored by epoch AI who created the benchmark and the data remains completely private since OpenAI has nothing to do with epoch evaluation here. And GPD6 scored 97.6% on this
benchmark while Fable 5.1 scored 87.8% in comparison. Now deep sweet to me is an interesting benchmark where clearly we haven't reached the saturation point but you see so many frontier models
saturate around the 70% mark with no clear winner that makes a huge gap and ever since data curve released this benchmark models like GBD 5.5 that's nearly 5 months old now only score 7 to
8% below the current state-of-the-art looking at how GBT6 that scored 74%. Now looking at this benchmark, one trend that is becoming super exciting is token efficiency. When you look at the column
number of output tokens, you can see how GPT6 scored on number of output tokens generated. And this is actually pretty crazy. But first, a quick word from Zo sponsoring this video. It's pretty
obvious that more and more parts of our lives are being integrated with AI and often our interactions are scattered across Claude, Chachip, CEX, and Menace. So how can we have more ownership while
having a 24/7 agent? Zo gives you a dedicated computer on the cloud that's yours, meaning your agent is on standby for anything that you give. And part of being 247 is the ability to message your
agent through text. You can directly message Zo to have normal conversations through iMessage or ask about files that you have stored on your computer in the cloud. And better yet, build an
e-commerce website for vintage watches and host directly on Zo using their AI agent natively through text. What's cool about Zo is that you can vibe code and launch sites that are actually useful
because you can integrate all your contacts into Zo and run automations. Custom domains are also included at paid plan and all your sites are hosted on your Zo's cloud computer. Plus, they got
some neat tricks like the selector tool to edit exact areas to improve things or even add automations that plug into your website like a text anytime someone fills out a form. Try it out today. I'll
have the link in the description below. When we look at the Prito Frontier for GPT6 result on DeepSu and compare it to their previous model GPT 5.6 Soul, the cost curve compared to the intelligence
demonstrated here isn't all that different, which seems underwhelming at first. So the additional cost that you pay to raise a few% just doesn't seem to be worth it. But when you look a little
bit closer, GPT6 actually spent half of the tokens in comparison to GPT 5.6 soul in output tokens. What you're seeing here is a model that is not costefficient but token efficient which
there is a difference between these two. GPT6 is currently being offered at $10 per million input tokens and $50 per million output tokens which is not costefficient at all compared to models
like GPT 5.6 soul that is $4 per million input tokens and $20 per million output tokens for short context. And despite the huge cost difference, the fact that the model shows the similar parita
frontier just goes to show you this very difference between what a costefficient model and a token efficient model shows. You also see this token efficiency at work in artificial analysis where the
cost to run the aggregate benchmark that we talked about earlier in the video. GPT6 used the least amount of tokens to get the job done. Now I think this creates a really interesting tension in
the AI industry. If more and more models are producing fewer tokens to accomplish the same amount of work, then the effective supply of intelligence is increasing per token. So for labs like
OpenAI whose business is largely around monetizing tokens, the tension here is that making the model more token efficient means OpenAI needs fewer billable tokens to deliver the same
amount of value. Which means OpenAI practically needs less inference compute for the same amount of demand. And OpenAI could now capture this basically an efficiency gain that they just gained
per token and either pass it down to the user to lower the cost per task or increase the price per token which by looking at GPT6 pricing and how expensive the model is the efficiency
gain went directly towards OpenAI's profit margin for now assuming that the infrance overhead for GPT6 really hasn't jumped by that much. And for other labs that are competing with OpenAI, it's not
as simple as catching up to OpenAI in terms of capabilities, but more importantly doing as little work to get the same level of capabilities, which is something that we can directly see in
Deep Suite, which shows you the operational efficiency when we compare Gemini 3.8 Flash and Opus 5 that score similarly to Astra, but they're four to five times less efficient per token in
comparison to GBT6 Astra. Now, beyond all the nitty-gritty details that we just walked through, there are also some really exciting things that came with the announcement of GPT6. The most
obvious one being what they showed off in GPT6 computer use where people made some really cool things just using their voice. Certainly not AGI by my definition, but still a really cool
demonstration. While my personal experience with egentic computer use was largely underwhelming to say the least, computer use is just one of those things that I just wanted to be true, but we're
just not there yet. And given how GPT6 scored among many relevant benchmarks, it's certainly hinting that we're getting close to finally having Jarvis running at home. I can't even imagine
just how much more productive I will be as AI continues to make these things possible. And all of this causes us to really think about what a good model is. Few days ago, we just had Fable 5.1 drop
from Anthropic. And by certain measurement, Fable 5.1 comes across as a better model than GPT6. And the true definition of what a good model is certainly changing from simply an
intelligent model to a more useful model where token efficiency, cost efficiency, speed, and how well it scores on real use cases speak more about the model than the old way of measuring the
model's capabilities like how we see in artificial analysis intelligence index.