Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
Claude Opus 5 has been released with state-of-the-art performance on coding and knowledge work benchmarks, often surpassing Fable 5 while being cheaper. The model shows major improvements in agentic terminal coding, novel problem solving, and computer use, and is available at the same price as Opus 4.8. Its enhanced verification abilities make it particularly promising for agentic loops.
Key points
- Claude Opus 5 achieves 43% on Agentic Terminal Coding, compared to Fable 5's 33% and Opus 4.8's 21%.
- Opus 5 shows a dramatic jump in novel problem solving from 1.5% (Opus 4.8) to 30%.
- Opus 5 outperforms Fable 5 on agentic computer use and business workflows while being cheaper.
- Opus 5 is much stronger at verifying its work and iterating carefully until it succeeds.
- Opus 5 is available on all platforms at the same price as Opus 4.8.
- Fable 5 is described as a wise old owl for planning, while GPT 5.6 soul is like a Rottweiler that won't let go of a task.
Tools mentioned
Techniques
- Agentic terminal coding
- Computer use
- Agentic business workflows
- Multidisciplinary reasoning
- Verification and iteration in agentic loops
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
Okay, so Claude Opus 5 literally just dropped. Now, obviously, I'm going to spend the rest of my day playing around with it, comparing it to Fable 5, and I'm going to bring you guys an experiment. So, stay tuned for that later today. But I wanted to just real quick come in here and just show you guys this cuz I think it's really, really interesting. On several coding and knowledge work evaluations, Opus 5 is the new state-of-the-art.
So, look at this first benchmark right here, Agentic Terminal Coding. Opus 5 has 43%, Fable 5 had 33%, and Opus 4.8 had 21%. So this is a major jump in aentic terminal coding above opus 4.8 but also above fable 5. Same thing with knowledge work. This thing crushes opus 4.8.
I don't exactly know what the novel problem solving benchmark is solving but look at this. It went from opus 4.8 at 1.5 to 30% at opus 5. So anyways these benchmarks are insane. You can also see here computer use which I always would default to codeex. I always thought that GPT 5.6 or 5.5 with codecs was just the absolute best with computer use.
But look at this with Opus. Opus 5. I'm going to have to try it out with computer use. Anyways, what's really interesting to me is this efficiency because as you guys know, Opus is half the price of Fable and Fable will burn through your credits like nobody's business. But look at this.
Agentic computer use Opus 5 performs better than Fable and for cheaper. Look at this. Agentic business workflows. Opus 5 is better than Fable and for cheaper, which is just really, really interesting to me. Multidisciplinary reasoning by effort level.
Opus 5 does better than Fable and is cheaper than Fable. So, like I said guys, I'm about to run a ton of usage credits. If you look at this right here in my claude, I am already at my weekly Fable limit and I'm about to go into, you know, thousands of dollars in extra usage credits, but I'm going to run a ton of experiments here and I'm going to bring you guys my consensus. But, I just wanted to come in here and just like show you guys that real quick. It's really interesting.
So, if you real quick go into your VS Code, go into cloud code, update it, get on Opus 5 and just start running your skills and start running things and seeing how it feels because as we know these benchmarks are fun to look at, but always take them with a grain of salt. It always matters on your use case, the way you talk to agents, the type of work you're doing, things like that. So, Claude Opus 5 provides greatly improved performance for the same cost as its predecessor, Opus 4.8. Opus 5 excels at valuable software engineering tasks like the Frontier Bench on Cursor Bench. I really want to see how it performs on the deep suite because that's the one that's been like a pretty true like leading indicator.
But the whole cost thing is just really interesting to me compared to Fable 5 because I know a lot of you guys were kind of like, "Oh man, it sucks that we're going to be losing Fable or that, you know, we only get a certain different limit for it." But it's overkill for most things and it's really expensive. So if Opus 5 is better for like general knowledge work, then that's a huge win. And this is really exciting to me. Opus 5 is much stronger at verifying its work and iterating carefully until it succeeds. I heard this really funny analogy on X, which is honestly pretty true.
They basically said like, "Hey, Fable 5 is like a wise old owl. It's really good at planning and managing and ideulating and coming up with ideas, whereas GPG 5.6 soul is like a Rottweiler because that thing will grab onto a task and it will shake in and it won't let go until the task is done, which I definitely felt. It was spinning up more tests and it was better verification than Fable was, right? But now maybe Opus 5 saw that and was like, "Oh wow, we need to work in some of this GBD 5.6 soul verification." And now if we have that, that's huge because verification is kind of what powers all of these agentic loops that everyone's talking about. And if you're using a model that's really really good at verification, your loops are going to be better and your outputs are going to be better.
So I'm really excited to test out Opus 5 with some verification. And here are some examples where they gave Opus 5 some real scenarios and had it find bugs, verify, take different approaches, and ultimately give a good output. But anyways, this is now completely available to everybody on all platforms. Same exact price as Opus 4.8. So, like I said, I'm about to jump in right now and start testing the heck out of it against Fable 5.
And I'll bring you guys that full experiment later today. So, keep your eyes peeled for that video. But anyways, I'll see you guys there. Thanks everyone.