Fable 5.1, is it that good..?

summarized

TLDR

Fable 5.1 tops Artificial Analysis's composite benchmark but scores only 67.4% on Deep Suite, placing it near GPT-5.6 Luna and Grok 4.6. The discrepancy stems from different benchmark methodologies: official benchmarks pin the harness (e.g., mini suite agent) while third-party runs vary conditions like test repetitions and time budgets. Combined with high token verbosity and cost, Fable 5.1's practical value is undermined by cheaper, more efficient competitors like GPT-5.6 Sonnet and Gemini 3.8 Flash.

Key points

Fable 5.1 ranks first on Artificial Analysis but scores second, sixth, and seventh on individual benchmarks.

Anthropic's blog omits Fable 5.1's Deep Suite score of 67.4%, found on page 168 of its system card.

Benchmark scores differ because official tests pin the harness while third-party runs vary conditions.

Fable 5.1 generates 119,000 output tokens on Deep Suite versus GPT-5.6 Sonnet's 60,000 for a higher score.

Anthropic's $200 max plan offers only about six times the weekly allowance of the $20 pro plan, not 20x.

Tools mentioned

Techniques

  • composite benchmark scoring
  • pinned harness benchmarking
  • token efficiency analysis
  • cost efficiency comparison
Transcript (captions)

0:00 Enthropic holds the top three spots on artificial analysis now with Fable 5.1 sitting at number one. So does that mean Fable 5.1 is the best model? Well, this is where things get a little bit weird.

0:10 Underneath the artificial analysis benchmark that aims to measure intelligence, the scoring here is actually an amalgamation of nine different benchmarks that make up one

0:20 composite score. And when we take a closer look at Fable 5.1, it doesn't actually score first place across the board, but you see the model actually scoring second place here, sixth place

0:31 there, and seventh place as well. Fable 5.1 looks like the best coding model out there because it ranks number one on Sciode and Terminal Bench. But what about Deep Suite? Deep Suite tends to

0:41 get a lot of attention from developers since it mirrors really close to how software engineering works. But nowhere in Anthropic's announcement or its blog doesn't mention how it scored on

0:51 Deepswu. And we have to look that its 212page system card for Fable 5.1 and land on page 168. Do we finally find Fable 5.1 actually scores 67.4% on Deep Suite, which puts the model at a

1:06 much lower tier close to where GBT 5.6 Luna and Gro 4.6 is. Now, here's where things get really strange. There's a difference between what the official benchmarks say and what the labs

1:17 actually report in their announcement. Even though both of these are referring to the same deep sweep benchmark, they could totally be in congruent to each other. And the reason largely has to do

1:27 with the fact that we no longer really use just the model alone, but we often access the model primarily through the external layer that we call harness like cloud code, cursor, and codecs. And

1:40 benchmarks like Scode test primarily on the model without the harness since the benchmark asks the model to produce the code and the environment will then run the code it generated to run a pass or

1:52 fail test to actually evaluate the model's capability. And this is what some of the questions might actually look like where you have the question of the problem itself and the dock string

2:01 that describes the function and the dependencies that it has. But benchmarks like Terminal Bench and Deep Suite are different by nature since it measures the model by putting them inside of a

2:11 harness. And while both of these benchmarks allow anyone to run these benchmarks using whichever harness they want to use, the official benchmark typically freezes the harness layer to

2:22 one and only swap out the model so that we have a controlled variable in the experiment so that the model actually becomes the differentiating factor. While this sounds good in theory, often

2:32 times because we rarely use just the model itself, but instead we use the entire coding system that includes both the model and the harness. Simply having one model score better with a pinned

2:44 harness doesn't actually tell the full story. So when we look at the official benchmark from Terminal Bench, as you can see, it measures by both changing the model and changing the harness. But

2:55 when we look at the benchmark that we saw in artificial analysis adaptation of the terminal bench where the only changing variable in this case is the model while the model harness remains

3:06 pinned to the coding agent called Terminus 2. And here's another thing you might have noticed between the two. When we look at the previous model like Fable 5, the official benchmark from Terminal

3:15 Bench says Fable 5 on Terminus 2 scores 80.5%. And when we take the same model and the same harness and see how it scored on unofficial analysis that also runs the same controlled experiment, it

3:27 scored 84% instead, which is much higher. The reason why this discrepancy exists here is because even though the model and the harness and the data are all the same, they show two varying

3:39 numbers because they both have different condition largely on its budget. For example, artificial analysis will repeat the test three times to get the score while the official benchmark will at

3:49 least run them for five times. And also the time budget they are allotted to solve the problem is different as well as many other things like the sandbox and more. So even though they both are

4:00 referring to the same benchmark terminal bench, the numbers aren't really congruent to each other. You can see the same kind of thing when we look at deep suite as well where they pin the agent

4:09 called mini suite agent to make sure that all the models are tested using the same harness even though the benchmark technically allows other harnesses to be used like codecs cloud code and more.

4:20 That's why even going back to anthropics release of their deep benchmark score that says 67.4% 4%. We really don't know how Fable 5.1 will actually stack up with the official benchmark since in

4:32 this case it's completely swapped where the official benchmark here pins the harness to use the mini suite agent for fair comparison. As you can see, as users, there's a lot to keep track of to

4:43 not only keep up with new models that are being introduced and how these new models exactly pair up with different harnesses that exist at the application layer and how these models and the

4:53 harnesses actually score with newer benchmarks that are being introduced and keeping track of different variations of the benchmarks that people run differently on their own. Now, we can't

5:02 really talk about Fable 5.1 without talking about Enthropic subscription model that's notoriously bad when it comes to user experience. And what's also bad is when your internet traffic

5:11 is not secured. And speaking of protecting what you're doing online, this video is sponsored by Surf Shark. Now, as you continue to use your favorite agent in coffee shops or

5:19 airports, so you can continue building your software, something that we always think about the least is privacy. The websites that you visit and your browsing patterns are all logged along

5:29 with your IP and location, and you want to have a VPN that you can trust to make sure that your traffic is secure and safe. Now as we all know thanks to agents like cursor and codecs on your

5:39 favorite terminals, apps on your phone or even cloud agents. This means that we can use our laptops, our tablets and our phones so that we can continue building the next project with our agents.

5:49 Thankfully, Surf Shark provides services on all these platforms under one subscription. BeyondVPN, you can also get features like clean web to block ads, which can be annoying, or trackers

5:59 that follow you. An alternate ID if you don't want to hand out your real email and your personal information every time you sign up for something. Try Surf Shark today. Using the code Caleb's code

6:08 at checkout, you can get four extra months of Surf SharkVPN. Link in the description below. Since the release of Fable 5 back in June, this new checkpoint of Fable 5.1 took about 3

6:18 months. And during the 3 months, many people had a chance to try the Fable model. Since Enthropic looped the model into all paid subscription with its own usage window and now Fable 5.1 is sadly

6:30 no longer available in the pro membership that costs $20 a month which forces users to then either pay additional usage credit on top of their pro or just simply move up to a higher

6:41 plan. This is certainly an interesting move from Anthropic since in comparison OpenAI subscription model includes usage of their most capable model GBD 5.6 six soul on their chachi plus subscription

6:52 that cost the same. So gatekeeping their most advanced model like Fable is certainly where competitors like OpenAI XAI and Google can start attacking since more and more labs are now releasing

7:04 models that are extremely token efficient and costefficient which are two different things. You can have a model that's extremely costefficient like GPT 5.6 6 Luna and DeepSync V4 Pro

7:14 where the cost per task is significantly lower where enthropping models are just across the board terrible in cost efficiency and so much more for Fable 5.1 as you can see. But token efficiency

7:25 can be shown more when we look at benchmarks like Deep Sweet when we look at the column number of output tokens generated. You can see just how verbose models like Fable 5 is that generates

7:35 119,000 output tokens while scoring 70%. When we also have models like GBD 5.6 six soul that not only scored higher at 73% while also using 60,000 output tokens. Even when we look at the new

7:48 model from Google Gemini 3.8 Flash, their cost accumulation as it approaches top score is significantly shorter. And when we compare that to basically all anthropic models that you can see here,

7:59 they just really like to spend a lot of money. And we see this exactly mirrored in their subscription plan as well, where Fable 5.1 chews up token allowance like no other significantly fast.

8:09 According to Anthropic's documentation, it says that even on max plans, Fable 5.1 usage is included as a standard part of their plan and we could use up to 50% of the weekly allowance. And anthropic

8:21 sneaky marketing doesn't really help here either. We're even at their $200 max plan. The max plan doesn't really mean 20x the pros weekly or monthly allowance, but it's actually just making

8:33 the 5hour window more elastic so people can burn more tokens during a shorter burst, but drawing from the well of weekly allowance. And the estimation out there is that the max 20x plan is

8:44 estimated to be actually about six times the allowance weekly than the Pro. not really the full 20x since it only applies to the rotating 5-hour window. Now, whether Anthropic did all of this

8:55 intentionally or not, it's a great way to confuse all their users in what they're buying into. And I think a lot of people are noticing this as they use Fable in production and come to

9:05 conclusion that it's not very practical to use models like Fable that offers more intelligence, but also significantly more expensive than other models that are two to four times

9:14 cheaper, more token efficient, or both. So really the definition of what makes a model good is really changing here since people often paid more in the past to use a better model before because back

9:26 then what defined as a good model was simply intelligence. A good example of this is when OpenAI released their first reasoning model 01 where the model was significantly better than every other

9:37 model out there and OpenAI could charge a huge premium for their top tier intelligence. But back then, we didn't really have as many harnesses that we seem to have today. So, how useful the

9:48 model was wasn't really something that a lot of people thought about. But now, what makes a model good is a lot more than just what we see in benchmarks. And while Fable is certainly a good model

9:58 and it had a huge leap for some time, the rest of the industry not only caught up to Fable really fast, charging a high premium for this long is starting to make a lot less sense really quickly

10:09 since how we use these models in different harnesses available today needs to be more practical, which makes the concept of what makes a model good completely Different.

Frontier News · by Hyperjump Technology