Opus 5.5 vs GPT-6 is racing to the bottom..?

summarized

TLDR

Opus 5.5 dominates cost efficiency over GPT-6 models at higher intelligence levels, but GPT-6 models lead in token efficiency. The cost of intelligence is falling, but that doesn't mean models are uniformly more efficient—subscribers and API users face different burdens from token inefficiency.

Key points

Opus 5.5 beats GPT-6 models in cost efficiency past a certain intelligence index.

GPT-6 models outperform Opus 5.5 in token efficiency across most of their range.

API customers pay for token inefficiency directly, while subscription users are insulated by labs.

The Pareto frontier for cost and token efficiency has shifted across 2026 quarters.

Anthropic's revenue is mostly API-driven, while OpenAI's is subscription-driven.

Tools mentioned

Techniques

  • token efficiency
  • cost efficiency
  • Pareto frontier analysis
  • 3D trajectory visualization
  • agentic use cases
  • daisy chaining agents
Transcript (captions)

0:00 Thanks to heavy competition, Opus 5.5 costs 40% of Fable 5.5 and GPT6 Soul and Luna now cost 50% less. What we're seeing here is the cost of intelligence continually dropping. But is it really?

0:12 We often use this dotted line here that follows the Parto frontier to track how models are actually progressing and we celebrate whenever we have a new model that breaks out of this line to chart a

0:23 new frontier in the model intelligence without sacrificing cost equally. For example, when we strictly compare the new GBT6 models with Opus 5.5, it clearly shows you that Opus 5.5 pretty

0:34 much dominates in cost efficiency when you go past GBT6 soul, as you can see. But what we often don't see is the Prito frontier in token efficiency. In other words, given two different models, we

0:46 measure how little work the model actually did to get the same level of intelligence. And we can measure how much token each model spent to get the same job done. And GPD6 models,

0:56 including Astra, basically dominates in token efficiency for as long as they can push their intelligence until Opus 5.5 takes over since it can push intelligence even further. But clearly

1:06 the trajectory of GBT6 shows more token efficiency than Opus 5.5. So what you're seeing here is a tension between costefficient and a token efficient model. So if both cost and token

1:18 efficiency is important, why can't we just see both of them at the same time in 3D? What you're seeing here is the trajectory of Opus 5.5 model across three axes. Costs in one axis, token in

1:28 the other axis and intelligence in the other. And it gives you an intuitive understanding of how exactly OPUS 5.5 pushes its intelligence, but also seeing the trade-off that the model makes along

1:39 the way. And you can see that when Opus 5.5 actually approaches the intelligence index of around 53, it starts to plateau. Now when we add this arbitrary line that scales more efficiently as a

1:51 baseline, we see it more clearly how the distance would diverge further from the baseline and Opus 5.5 starts to drift into being less token efficient. Now when I add Enthropic's previous models,

2:02 Fable 5.1 and their previous Opus 5, you start to see even more just how Opus 5.5 scales differently in the cost efficiency axis, while Fable 5.1 and Opus 5 scales more on the token

2:14 efficiency axis in comparison. Although none of these models still scale well in comparison to the baseline. Now this is how GPT6 soul scales in the same plot and how GP6 Luna scales as well. And you

2:25 can see that it has a completely different trajectory than Opus 5.5 with an important point that GBD6 Luna and Soul only go so far in its intelligence since GBD6 Soul tops out around 47.5

2:37 while GBT6 Luna at around 37.3 in the intelligence index. But when we look at it from the top down temporarily, not looking at the intelligence for now, we can appreciate just how GPT6 Luna moves

2:49 through the cost axis more effectively than other models that you see here. And Opus 5.5 on the other hand pretty much scales close to the baseline in both token and cost axis. Now, when we add in

3:01 GPT6 Astra, we see that the model scales more like Fable 5.1 than it does to Opus 5.5, which goes to show just how Opus 5.5 really scales more evenly in both cost and token efficiency compared to

3:14 the rest. Now, the question here then is how this actually looks like in real life when the models are actually taken into agentic use cases. What happens to the cost and token then? But first,

3:24 let's talk about Hyper Agent, who's sponsoring this video. I'm planning my trip to Italy soon and I decided to use hyper agent. So I created these four agents to help me plan my travel and the

3:33 goal is to have them daisy chained so that one specialized agent can call multiple other agents as needed. And this is my prompt to create Sophia who's going to be my travel planner and Marco

3:43 who helps me research flights, Giani who makes sure that I have videos and images to make this trip fun. And under each agent I can give access so that other agents can invoke the agent as you can

3:52 see. And now I can call up Sophia with my ideal plan to Italy with my dates and budgets and it will go and create the rest. As you can see here, you're seeing Sophia delegating multiple agents as the

4:02 entire trip is now being executed. And these are the results that I can see that shows you not only the detailed plan and images in highquality videos that go with it for the upcoming trip so

4:12 I can visualize what I might see. If you want to build your own, check out Hyper Asian Blows and sign up for a paid plan and you'll get a $100 in bonus credits. So from the consumer side, this must

4:21 mean regardless of token efficiency, what really matters is cost. Since Enthropic is offering per token at a cheaper pricing, right? Not quite. And the reason is largely because in the

4:31 application layer, there's a huge split between people who use the model through API versus people who use them through subscription. People often use the API route to build and use custom agents or

4:43 build custom applications. while a lot of integrated applications and coding agents will be through the subscription fabric. And when we look at the revenue source from 2025, OpenAI had majority of

4:54 the revenue come from subscription rather than API. Though it's opposite for Anthropic's revenue source where majority of the revenue comes from API and subscription limits are often

5:03 metered by token usage while API usage is ultimately metered by cost. And the difference here is worth noting because per token cost in API is actually cost that people pay while subscriptions are

5:15 measured by total token usage given a rolling 5-hour window with a weekly allowance. What that means is that when we actually bring back our comparison that we made earlier between cost

5:25 efficiency and token efficiency, the user experience is going to vary pretty heavily based on how token hungry the model actually is. and labs would often subsidize token usage at their

5:36 discretion depending on the competition at the time, data center availability, hardware costs, their margins and so on. So the tension here between users and inference providers is really who ends

5:47 up paying for the inefficient token that gets generated by the model. And for API usage, the token inefficiency directly gets passed to them while subscription usage is fairly insulated depending on

5:58 how aggressive or lenient the labs feel. So ultimately having a token efficient model is a great way to hedge against the growing competition because it gives them more flexibility to meet the

6:10 continual falling cost of intelligence as competition grows. When we look at other labs and look at their trajectory like Moonshot's Kimmy models, Z.AI's GLM models over time, DeepSync models,

6:21 Xiaomi's MIMO models, Gemini from Google, Miniax M series models, XAI models, and Meta. They're all collectively fighting in the cost axis, especially DeepSeek as you can see. So

6:32 looking at model releases that we saw from Opus 5.5, GBT6 soul, GBT6 Luna, it gives us an understanding of how models are actually scaling in intelligence, cost, and tokens. And understanding this

6:45 well will help us separate that the falling cost of intelligence doesn't automatically mean that the models are equally becoming more efficient underneath. One example is seeing how

6:54 the Prito Frontier moved from quarter to quarter in 2026. As you can see where you can see the clear progress from first quarter this year in the parto frontier and then how it moved up in the

7:05 second quarter and then so far in the third quarter as intelligence moves along with the parto frontier and approaching the arbitrary ideal case that I drew here and same case for how

7:15 token efficiency scale this year in the paro frontier as you can see and what we hope to see is the model scale as we go forward in meeting both the cost and token efficiency although most of us

7:26 might only look at the cost and intelligence. We also want to look at other axes like speed which can also change how we view the progress of AI and how the models are actually

7:35 advancing as we

Frontier News · by Hyperjump Technology