Mythos 5 & Fable 5 Launched

summarized

TLDR

Anthropic launched Mythos 5 and Fable 5, with Fable 5 available to most users and Mythos 5 limited to select customers. The models show significant benchmark improvements over Opus 4.8 and GPT 5.5, especially in agentic coding, but are heavily safety-restricted, blocking even simple biology queries. Pricing is double Opus but less than half the Mythos preview, and Fable 5 will be removed from subscription plans after June 22.

Key points

  • Anthropic released Mythos 5 and Fable 5, with Fable 5 being a safer, generally available version of the Mythos class model.
  • Pricing for both models is double Opus 4.8 but less than half the Claude Mythos preview, attributed to increased inference compute from a deal with Elon.
  • Benchmarks show Fable 5 outperforming Opus 4.8 and GPT 5.5, especially on agentic coding benchmarks like SweetBench Pro and Frontier Code.
  • The models have strict safety classifiers that block queries related to cybersecurity, malware, life sciences, and attempts to extract chain-of-thought reasoning.
  • Anthropic will require 30-day data retention for all traffic on Mythos class models, a change from previous policies on third-party services.
  • Fable 5 is included in Pro and Max subscriptions only until June 22, after which API token pricing applies.
  • The safety system is extremely sensitive, triggering on simple biology questions like 'break down current Ebola outbreaks' and switching to Opus 4.8.
  • The speaker suggests using Fable 5 for planning and other models like Opus or Sonnet for execution due to cost and restrictions.

Tools mentioned

Techniques

  • Safety classifiers for cybersecurity and life sciences
  • Long chain-of-thought reasoning
  • Jailbreak prevention via input monitoring
  • Data retention for safety analysis
Transcript (captions)
Okay, so Mythos 5 is here or actually rather renamed rebranded perhaps even dumbed down a little bit Claude Fable 5 is here. And so this is the new model from Anthropic. We've had the whole sort of talk around Mythos and about how dangerous it's been. Now they've actually released it, but what we're getting is not the full Mythos 5 release. They're calling this a Mythos class model that they've made safe for general use. So they're claiming that this model's capabilities basically exceed those of any model that they've ever had generally available, which is kind of hinting that they've had a lot more models perhaps that internally that are even stronger than this. This will be interesting to see going forward, but undoubtedly launching this model is going to allow people to try it out and actually see is this model actually better for real world sort of use for coding for all these kinds of tasks. And in this video I'm not going to go deep into testing the model. I feel like it really takes a couple of days to actually sort of test the models now and actually to see the nuance of what they're actually good at, but also to see, you know, are they overly hungry on using lots of tokens or what are the trade-offs that you're actually making for the cost versus intelligence etc. So this video I thought I'd go and have a look at it. I'll also show you an interesting thing about how they're actually limiting the model which I'll show you in the demo later on. So they've actually released two models. Most of us will only have access to one of them being the Claude Fable 5. That said though, they are actually making Mythos 5 open to some of the people in the Project Glass Wing. So my guess is that that one is going to be very heavily limited to people who are strong customers and by that people who are spending or companies that are spending a lot of money and also have really good use cases for actually using the Mythos 5 model. Now, the big thing that always concerned me about the Claude Mythos preview was the actual pricing. And I got to say here that the pricing for this model for both Fable 5 and Mythos 5, while it's double what Opus 4.8 is, it's not as high as I originally expected. I actually expected that it might be quadruple the Opus price. And they do make the point to actually mention here that this is less than half the price of the Claude Mythos preview. So, it does seem that that deal that they did with Elon to basically get a lot more inference compute has actually helped them to probably be able to serve this at a scale where they don't need to charge as much as they were for Claude Mythos preview. Now, if we come in and look at the benchmarks here, undoubtedly it's doing extremely well, way better than not only Opus 4.8, but even better than GPT 5.5. I kind of feel it's pointless to compare it to Gemini 3.1. That model is already very old, and in some ways it might have been interesting more to compare it to the Flash 3.5 model. So, clearly this is doing a lot better. The question, I guess, is is it doing better enough to justify it being double the price for most use cases? And that's something I think that we're only going to know after using this for days, if not even weeks. On certain things like the agentic coding, which you can imagine is going to be key for things like Claude Code, etc., we've got a very substantial bump over the other models for things like SweetBench Pro, Frontier Code, which is a new benchmark from Ignition that just came out. If we just look at their tweet about this, they showed that, you know, really none of the models are doing very well at this except for Opus 4.8, which is getting 13.4. We can see that this is getting over double that. So, that is certainly interesting in there. But then when we look at some of the other benchmarks, we can see it's not hugely better than Opus 4.8. With things like the tool use benchmark, with things like computer use benchmarks, etc., it doesn't seem like it's going to be a substantial win for everything. But then certainly other things it will be like looking at this legal benchmark is kind of interesting to see that this is almost like 30% better over Opus 4.8, not to mention it being a lot better than GPT-5. Again, this makes me wonder how much they're actually cherry-picking which benchmarks they're doing well at versus which they aren't. For biology here, the benchmarks that they've got show that, okay, this does a lot better. But as I'm going to show you when I do the demo, try and ask it anything about biology and actually get an answer back, you're going to find that alone is going to be a very challenging task. So, even in the blog post they talk about Fable's new safeguards and, you know, how that actually works. I will show you some of that in a second. The safety classifiers that they have on this for cybersecurity and for any sort of they claim research biology, but what I'm going to show you is that even just simple things can actually trigger this off. It is going to be very interesting to see how do people actually jailbreak this model and has this system that Anthropic has put in there to prevent people from jailbreaking it actually worked or not. Another thing that's also interesting is that they're claiming that this model is causing them to have to basically change their data retention policy. So, you can see that they're saying that they will require 30-day retention for all traffic on Mythos class models on both first-party and third-party services. Now, that's going to be really interesting because as I understood it, if you use the Anthropic models on something like Google Cloud, one of the advantages for that was that Anthropic didn't get access to your data. I I don't know if this model's actually been launched on Google Cloud yet or on AWS, etc. But, it's going to be interesting that that that's quite a big change for certain customers who are just not going to want to be giving their data over to Anthropic. Now, they do claim that they won't be using this to train new Claude models or for any non-safety-related purposes. I'm very suspicious of a lot of these companies and how they do stuff, but then I've seen things in the past where people use derivatives of data, which they can then basically use for other use cases while they're not actually technically training on your specific data itself. But, it does seem that they want to be able to capture any sort of jailbreaks or any ways that people find to get around this so that they can plug those holes as quickly as possible. Okay, so like I mentioned before, the price is actually not as bad as I expected here. Both these models if you can get access to the early access program for Mythos 5, but for most people who are going to have access to Claude Fable 5, you're going to be paying $10 per million in, 50 per million out. Interestingly, that they talk about that okay, from now until June 22, access is included on the Pro and Max accounts. But, on June 23, they're going to remove Fable 5 from those plans and you're going to have to pay for the API token, you know, pricing just like companies are having to pay, you know, for this. It does seem like Anthropic is trying to basically flag that the all you can eat buffet plans for tokens uh perhaps more on the way out, not only for companies but even for individual users. Now, they do mention here that they will aim to restore Fable 5 as part of the subscription plans. My guess is it probably never comes back to Pro. It's only going to be for Max and higher and they intend to do this as quickly as they can. So, my guess is that will probably depend on how many companies decide that they're quite happy to pay for the API tokens for this model, how many people actually realize this model is going to help them do things better versus not, et cetera. Okay, so when you come into the actual Claude app yourself, you're met with the new Fable 5 option in here. It's going to cost you twice as much as usage of Opus. But, the other thing is that this model is definitely a lot more safety restricted. So, you'll see this where basically safety measures flag a message, automatically switch to a different model, but keep chatting. And then if you come to learn more, you can actually see that this is sort of uh a common thing that they clearly know with Fable to lock it down or perhaps to generate more publicity for their IPO, they're actually locking this down where, you know, that they're discussing that it comes with risks and that they want strong safeguards. So, therefore it's going to block anything that is sort of cybersecurity related, malware or attack related. This is not really that surprising. If you've hyped up Mythos so much that it can basically go and hack everything, probably one of the first things that people are going to try is, "Hey, go and hack something for me." So, I'm a little suspicious that this perhaps is not just from a a safety perspective, but also from a marketing perspective here. The other interesting thing is the whole thing around life science queries and asking things around that. And then finally, you've got basically anything that you're trying to do to extract out the summarized thinking, also known as the long chain of thought, in here. So, this seems to be Anthropic's way of trying to lock down the long chain of thought as much as possible. That the model itself, not only will you run into issues with it, but the model itself can actually block requests for anything that is related to trying to elicit out its chain of thought, which is very interesting because that may show that one of the key things that makes Mythos better is actually just getting it to make better chains of thought. And it does seem like they've trained some sort of classifier to basically sit there and either monitor your usage or to basically monitor the the API. You see here they talk about the checks review everything the model reads, not just your latest message. So, clearly trying to get around these blocks by uploading files or injecting things into memory are probably not going to be allowed. Now, it is very interesting that I have Claude Max, the biggest subscription you can get, and even here it's only saying that this is that Fable 5 is included to June 22. I don't know what happens after June 22. Does that mean suddenly people are going to have a new level of subscription that they're going to have to try things out on? And so, if I ask it something like this, break down the current Ebola outbreaks, what are the key risks for the World Cup and this? So, this is not asking anything dangerous. I'm curious to see like, okay, will the trigger just be triggered so easily? Or will it actually Okay, and you can see this already has switched to Opus 4.8. So, you can see here if I ask it why it switched to Opus 4.8. It basically flagged the messages, you know, biology topics. And it does seem that they've become extremely sensitive in what will actually trigger these flags, right? The fact that I haven't really asked for any sort of recipe about Ebola or anything like that. I've just asked it to break down with the current Ebola outbreaks, what are the key risks for the World Cup in the US and how they could be mitigated or not. Okay, so I've taken the or not off and trying with Fable again to see and no, it still gets triggered. So, it does seem that it is extremely sensitive to anything that even remotely would be in the same area as something that that's dangerous. So, I'm going to test this more in-depth over the next few days. I might make a follow-up video actually looking at this. It's certainly interesting that we've gotten this release. The big question I wonder is that based on the costs and stuff like that, is this really a model that many people actually want to use? And I do think this is going to force a lot of people to finally perhaps use a model like this for planning, but then use other models perhaps like Opus or Sonnet for the actual execution part. All right, if you've had a chance to play with the model, let me know in the comments what you think. This is really just an early take video of looking at it, trying it out, trying to work out what's going to be possible with this model going forward. And I'm sure it's something that I'll end up covering more over the next few weeks, etc. >> [snorts] >> Anyway, as always, if you found the video useful, please click like and subscribe, and I will talk to you in the next video. Bye for now.

Frontier News · by Hyperjump Technology