Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
Opus 5.5 won 7 of 10 use cases against GPT-6 Soul in this head-to-head test, consistently producing better creative outputs like websites, videos, and 3D worlds. However, GPT-6 Soul outperformed on a codebase analysis task, scoring 100/100 at 1/20th the cost. The presenter concludes Opus 5.5 is a major step up from Opus 5, while GPT-6 Soul feels like a step down from GPT-5.6 Soul, making Opus 5.5 the better choice for creative work and GPT-6 Soul better for cheap, specific execution.
Key points
Opus 5.5 won 7 out of 10 use cases against GPT-6 Soul, with one win for Soul.
GPT-6 Soul scored 100/100 on a codebase analysis task at 1/20th the cost of Opus 5.5.
Opus 5.5 is double the price of GPT-6 Soul per token.
Agent interference occurred in two use cases because prompts did not isolate sessions.
Tools mentioned
Techniques
- scroll-driven website layering
- parallel task execution
- sizzle reel creation
- 3D world generation
- browser use automation
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
So, I've been playing around with Opus 5.5 and GBD6 Soul all day long, and I just ran them across 10 different use cases. I'm talking things like building websites, editing videos, building slide
decks, browser use, tons of reasoning, and then a few other things as well. So, let's not waste any time and just get straight into this video. Okay, so before we start looking at all these
comparisons, I wanted to talk about the pricing of these two models real quick. So, Opus 5.5 on the left is double the price of GBD6 Soul on the right. You can see that Opus is four bucks for the
input and 20 bucks for the output, whereas Soul is $2 and $10. Now, the reason I wanted to bring this up is because when I first glance at these prices, my instinct is like, okay, well,
Opus is probably going to be a little bit better than Soul. That's typically just what you think when something's a little bit more expensive. Now, I know a lot of you guys might disagree with that
perspective, but the way that I like to think about AI models is I like to think about, okay, if I had a hundred bucks and I gave a hundred bucks to Opus and I gave a 100 bucks to Soul, which one
would give me a better output or better quality for that 100 bucks? So, that's why I think it's important to sort of look at that metric and sort of, you know, now I have a little bit of a
feeling of what I think these models might be like, but obviously that's why we run tests. So, anyways, now that we got that out of the way, let's jump into some examples. So, real quick before we
jump into the 10 use cases, I wanted to do a real simple example just to show you what they feel like and how they think. So, on the left we have Soul and on the right we have Opus and this is
the cloud desktop app and the Codex desktop app. And I shut off this prompt that says, "Pull in my two latest Fireflies calls for Q4 planning and ingest those transcripts into the wiki."
I spelled that wrong so hopefully it knows what I meant. And I said, "Then find me my last AIS Plus Q&A where I brought up NAND." So these are two different tasks. One is ingesting calls
into a wiki and one is finding something. So, it'll be interesting to see how they think about doing these in parallel. The other thing is this isn't super super simple. These two calls
combined is about 4 hours of transcript. So, it has to create all those links and the indexes and link everything together. So, it's not super simple, but I just want to see how these models feel
and how quick and lightweight they feel as well. And you can see so far Soul seems to be cooking way ahead of Opus 5.5. It's doing these both things in parallel. Like I said, it's looking for
the new Fireflies calls and ingesting them. And it's also locating the Q&A in parallel. So this is pretty interesting. It's been about 6 minutes now and Claude came back and said, "Hey, I noticed that
something else is also doing this." So it noticed Codex doing it and Codex so far has the answer wrong about when I last mentioned Nitn. It said August 17th. But the right answer is what Opus
found which was September 14th. So anyways, I said overwrite. It's okay. Both of these models are still spinning and soul just finished as you can see and it got this answer wrong. And that's
not good, right? Like it's not that hard for a model to look through a bunch of transcripts and find the most recent one where Niden was mentioned. It's probably because it was spelled out wrong like nn
or nadnen as you can see here when it was actually transcribed. But still I feel like that is a problem that soul should have found for me and they finished up in basically the same amount
of time but opus's answer here was definitely better. So this very quick experiment gives us a little bit of a baseline to take into account as we head into the next 10 actual use cases. So
let's jump right in. Okay, so the first one was a quick website design. I gave in all 10 of these examples I gave both models the exact same prompt. So here is one output. I won't tell you which one
this was yet. We have perk form. Your morning shouldn't require two drinks. We can see as we move our mouse, the keys and the can and the background are on a different layer. We start to scroll down
and we get this little animation right here. You can see the layering come into play. And now we have a different image down here. Coffee in one hand, protein in the other. Two purchases, two things
to carry, a shaker to rinse, blah blah blah. Or we have pull the tab. I like that little animation. That's kind of cool. Um, we get some smoke coming out. Your coffee and protein one can. We keep
scrolling down here. You have text animations with all the text coming into view. Pick your first can here. Okay, a nice little rotation animation. I do think that that's actually pretty clean.
And we come down here to some FAQ. And then we have out the door. I'm not sure if that's a glitch or it just leaves frame really quick. But anyways, this was one output. Let's take a look at the
other one right here. We can see if we reload, we get your morning shouldn't require two drinks. Not as strong of a hero image right now, but we scroll in, we start to get some different text
animations as well. we get like a little opening a vault. We have a weird packaging style image rather than like a realistic looking image. So, I think that was a bad design choice there. Um,
but you can see that both of these are keeping the same sort of like narrative and the same branding. We're getting all the same sort of like text animations. We're getting um once again the
packaging pictures rather than like real product images, which I don't love. And this overall, this one just seems a lot more simple. Definitely not as premium as this first one. So, this first one
was definitely the winner when it comes to design and style. And this one was Opus 5.5 and this one was GBT6 Soul. Real quick guys, by the way, if you want the skill that I use to generate scroll
driven and layered websites just like these examples that you saw, then you can grab this for completely free. The link for that will be down in the description. Let's get back to the
video. Now, when we take a look at the pricing, we can see here they both took about 35 minutes. Opus took a little longer and Opus also cost three times as much. So, $1832 compared to 5.89. But
right now, this one, even though with the price, the winner of this experiment is going to Opus 5.5. No doubt. So, 10 Opus 5.5. Okay. So, example number two is something that I think is just a
really, really cool benchmark that I might keep doing for all these models where I basically give them this Frame.io link with tons of different videos, 100 gigabytes of videos from our
recent AIS Live event. And I basically just ask them to create a sizzle reel, a 30-cond sizzle reel that helps us market the next event, which is October 17th and 18th. If you guys are interested,
definitely check it out. I'll put a link to the event in the description. But anyways, I basically want them to look through everything, tell a story out of it, and make a high energy sizzle reel.
So, let's take a look at both outputs. This first one will be Opus 5.5. >> Hello. Hello, AIS Live. My >> Oh my god, I'm so excited for this. >> Let's Let's get some energy going in
here. It's been a great two days. Wow. Okay, that was my first time watching that and I will say that that is so so impressive. Like that almost
feels more impressive to me than when Astra did that on its first try. So, that was really, really good. I mean, look at how it did the layering here. And, you know, some of it felt a little
bit static, but the music, the sound effects, the pacing, the way it came in and out, all of this layering, I mean, I thought that this was actually really, really good. Like, definitely better
than I thought it was going to give me. All right, now we will take a look at Soul's version. This has been so energizing. Cloud code is my primary driver.
I can go to bed, come back, and it'll still be working. Okay. Um, I mean, it's still impressive, but it does not give me the same energy. It didn't have the same just like
fast-paced feel. it doesn't get me as excited for the next event compared to Opus' version. So Opus 5.5 definitely wins that in my opinion. So Opus is two for two right now. And you can see here
when it comes to time and pricing, Opus was faster, 5 minutes faster, as well as only double the cost. So even though Opus was five bucks more expensive here, I'm giving this one to Opus. Opus is 2
and 0 right now. And I'm so so glad that we're actually seeing this because Opus 5 was really bad. Like I haven't used Opus 5 in a long time. I was using 4.8 8 and Opus 5 was just so so bad. It was so
slow. But Opus 5.5 so far, I've really really been enjoying it. So, I'm feeling good about getting back a little bit more usage on my cloud subscription now. Um because I really do like both of
these harnesses and they've got different good models. Like I do really really love Astra, but Opus 5.5 for the cost, I'm getting excited about it. But anyways, let's move on to use case
number three. All right, so for use case number three, we are doing an Instagram reel. So I told it to use hyperframes. I told it to edit it. I gave both models the same footage and the same prompt.
Let's this time start with GBD6 souls output. And I'm not going to play the entire reel, but here we go. Stop prompting Claude. Andre Karpathy thinks there's a much better way to work with
AI. And his method has three layers. Layer one is the spec. Instead of giving Claude a task and hoping it understands you, work with it to create a detailed spec first. Have Claude interview you
about what you're actually trying to achieve. Then break the work into smaller checkpoints before it starts. Layer two is the verifier. Before Claude does anything, define exactly what a
good result looks like. Then give it ways to actually check its own work. whether that's another AI model or real data or tests that it can run on its own. Okay, not impressed at all. There's
a weird like humming noise in the background. There's not really that many engaging animations or there's no music at all. I think that this is a very meh output. Astra does a much better job
with editing reels and editing videos. But this is just not something that I would have put on my Instagram account for real. I would have wanted to iterate a lot more than just what we see here.
And now let's take a look at Opus 5.5's version of this reel. Stop prompting Claude. Andre Karpathy thinks there's a much better way to work with AI. And his method has three layers. Layer one is
the spec. Instead of giving Claude a task and hoping it understands you, work with it to create a detailed spec first. Have Claude interview you about what you're actually trying to achieve. Then
break the work into smaller checkpoints before it starts. Layer two is the verifier. Before Claude does anything, define exactly what a good result looks like. Then give it ways to actually
check its own work, whether that's another AI model or real data or tests that it can run on its own. Okay, that was really good. Like, this is so much better. It's more engaging. Look at
these like typing animations, the sound effects you guys noticed. All of this was more engaging. It's still not something that I would probably put on my Instagram, but this is like without a
shadow of a doubt significantly better than what we just got from GBD6 Soul. And this version was pretty interesting with the pricing cost. It took double the time with Opus 5.5 and like four
times the cost. But once again, in this case, I would pay the premium there. The question would be, okay, so Opus spent 11 bucks and Soul spent almost three bucks. If we iterated with Soul over and
over to get to 11 bucks of spend, would it be where Opus is or would it be better? I think that's the interesting question to sort of think about, but also like you love the ability to just
oneshot stuff and they both use the same skills. They both have the exact same stuff. It seems like there's just more potential over here with a lot of this sort of like design and creative stuff
so far. So anyways, right now Opus is three for three. So, we're heading into the next experiment with Opus being up 30. So, let's go on to number four. All right. So, in this example, I gave both
models a ton of data, like a ton of data, and I said, "Hey, build me a Google sheet analytics, and then build me a dashboard as well." So, this was Opus 5.5's Google Sheet. This one looks
pretty clean. You can see that we've got these different boxes with different metrics like AIR, MR, customers, you know, gross margin. We have all these stats. We have these visualizations
right here. We can move into a P&L. We can move into SAS metrics with more visualizations. We can see cash and runway, unit economics, pricing, headcount, go to market, pipeline, um
customers, and then some data notes as well. I think that this looks pretty well. It feels branded. We have, you know, clearly some sort of color theme with like black or a dark blue and
orange. And yeah, this looks pretty good. We also got a pitch deck, which I can open up right here. And you can see that we have um like a logo has been created. Once again, we have sort of
like that navy blue and dark blue or sorry, navy blue and orange. What I meant to say, we have a picture here. We've got um this stuff is coming through pretty simple, but also pretty
clean. We have more data visualizations, some more B-roll or not B-roll, images being generated. Um it's not too bad. It's not too wordy, but it's not too premium either. So, it's just all right.
And then it also created us an analytics dashboard. We can see right here that we have I mean this looks very vibe coded, very AI generated, but hey, the dashboard seems to be functional. It has
different ways for us to switch through. This is obviously using the same data that we saw on our um Google sheet that it made for us earlier. I like this highlighting effect. Um same down here.
So, not too bad of a dashboard. Oh, we can also switch between different tabs and switch between different modes as well here. So, that's pretty cool. And then the last thing you see that it
created for us here was a landing page which is running locally. And we can see we have an image here with different like you know paper and clipboard things flying around. These are on different
layers. So, that's a nice little scroll animation. Here we've got a Monday brief. We come down here into a little bit of an animation with filling in some gaps. Here we have more text coming in
dynamically. We've got a little scroll animation here. So very clean. You can tell it's obviously AI generated. It just feels a little bit AI generated, but hey, there's some nice animations in
here and some nice layering that definitely gives it a little bit of an edge. So anyways, that was Opus. Let's go over to Soul and see what it was able to do with all of this data. Okay, so we
have once again a PowerPoint, we have a Google sheet, we have an analytics dashboard and a client page. So let's take a look real quick at the actual slide deck. We can see that we have very
similar. Like I wouldn't say that anything here is too different, honestly. Like this looks almost exactly the same as the one that Opus 5.5 generated for us. And you know what I'm
actually suspicious of? Is this the exact same deck? This might be the exact same one. They might have gotten confused and gotten in each other's way there and just like one of them grabbed
the other deck, which is pretty interesting. M. Okay, so somehow along the way, these two models started working together. They probably started to work inside of the same project
folder, and I didn't yet specify in this example. I didn't say, "Hey, make sure you're not overwriting any other agents work." So, this is a really interesting point. Like, I don't know exactly what
to do here. Um, I think they built everything the exact same, which is really, really interesting. They literally created everything together. And in this case, they both still spent
money. And Soul spent way more here, actually. So, it it ran for 8 minutes longer and it cost three bucks more. So, this is honestly a little bit of a weird one. I'm not going to like assign a
winner here, and I'm sorry if you guys are like pissed off that I let them do that in this in this specific example, but they took a ton of data and they had to build kind of a story out of it and
create four deliverables and they sort of collaborated on that, which I thought was interesting. Anyways, here's what we got on the the cost and the runtime for this specific example. Okay, guys, real
quick. In the moment of filming, I didn't think anything too much of it, but while I was editing, I was like, "Wait a minute, that doesn't really make sense." So, I went back to the session
and I did the math again, and I found out that this actually was only codecs $3.60 and about 9 minutes. So, the previous figures, I was like, "Wait a minute,
that just doesn't really seem to add up." So, at the end of this video, when we're going over the total cost, just subtract about 19 bucks from um GBD6 Soul's total cost. But everything else
was pretty consistent. Once again, apologize for messing this one up. Um, I should have put that in the prompt. I normally do. Didn't really think to do it for this one for some reason. But
either way, I fixed it. And just to show you guys what happened, I had both GBT Soul and um, Opus inspect the logs on both sides and they both agreed on what happened
here, which is really interesting. Basically, I gave Claude the request and it began creating files and then Codex got the same request. obviously in the same repo. But what happened is the
first time I sent this to Codeex, it failed because like GBD6 Soul was still rolling out and I was like, "Oh, I got to get on these prompts." So it failed. And then when I resumed it, what was
this 16 minutes later? Claude was already like pretty far along the way with building. So it basically just grabbed all of the stuff that Claude was working on and then started editing
those files on top while Claude was still working. So, I truly think that some of these deliverables would have been better if maybe if Soul didn't get in there and make some changes. But, you
know, obviously it's hard to say, but I think that some of the stuff that looked a little bit AI generated, I'm wondering if maybe that's because Soul got in here and made some changes. But, I do think
that the deck looks pretty good and I think that this Google sheet does overall look still pretty solid. So, anyways, we're going to move on to the next experiment, but I wanted to at
least show you guys like the order of events there and what happened because I think that actually is, you know, pretty interesting. But anyways, moving on to use case number five, we have this game.
So, we're going to first look at um GBT6 Souls version. It created a game called Small Hours. I gave them the same prompt obviously, and this thing is loading up super slow. My computer down there also
started to get quite loud. So, this thing must be a pretty large game. Okay, so here it is. We basically will just be walking around with WD, clicking to interact, and mouse to look around. So,
I'll hit begin. We see the Ren Museum of Miniatures closed for good at 6 o'clock tonight. Everything your grandmother built goes to auction in the morning. You sat down in her reading chair to say
goodbye. Okay, we got some sound effects coming in. Wow. Okay, so this is like really telling me a story here. Okay, cool. So,
this is me able to look around, move around. Honestly, not too bad from like a design perspective. Although my mouse if it goes out of frame, that's it. Like my mouse isn't locked. You see my mouse
is nowhere. I'm moving my mouse around, but like nothing's happening until I bring it back into frame, which is a problem because if I want to do a full 360, I can't. So, that's a little bit
weird. Little bit of an issue there. I'm not even sure like how I'd go about trying to fix that. But anyways, walking around here. Okay. Click to interact. Switch on the sewing room light. Uh,
right there. Right there. Switch on the parents bedroom light. Oh, that's just telling me what I can do. Interesting. >> Wow. I mean, this is honestly unplayable
because of this mouse thing. Like, I can't I can't turn this way if I want to because my mouse goes out of screen. That's not good. But honestly, like as far as the storyline goes and the way
this feels, like it feels smooth. It's just this mouse issue, which honestly, I can't even play this anymore because of that. So, I'm going to have to stop. We're gonna go over to Opus version. You
know what's awful is the exact same thing happened here as what happened in the previous version. So they literally once again started working on each other's files. And
obviously that's like a prompt error because I didn't specify in these prompts not to edit any other files. Usually I do say that, but in this one I just kind of was shooting off these
prompts. So unfortunately we'll have to take the outputs here with a bit of a, you know, a grain of salt because they started working on the same thing. But I don't think that this experiment is just
totally ruined because the pricing is still really interesting to me because somehow we still had Opus worked for two and a half hours almost with Soul working for an hour and Opus coming in
at four times more expensive. So unfortunately guys, I'm sorry that that happened. I mean we we have seen so far that Opus does seem to be better when it comes to actually giving us these sorts
of outputs, but I think it's important to understand like where would the handoff be between these two types of models. All right. Well, okay. Well, number six, these agents did not work on
the exact same deliverable. Thank goodness. So, let's go ahead and launch these up. Basically, what I had them do is look through my past 100 YouTube videos. I'll show you the prompt up
here. And I wanted them to create me basically a 3D playable world explaining all these concepts from YouTube. I thought this was interesting to see how they reasoned about this and how they
thought about telling a story out of this. So, let's open these up and take a look. Okay, so this one is Opus' output. We'll go ahead and start exploring. We can see that we walk forward with W. Can
we Okay, we can move like this. So, we've got kind of this cartoon 3D world where we're able to walk inside and take a look. I can sprint with space. This is our
guide. We can talk to Nova >> building. >> Okay, cool. I also can throw these balls for some reason. I can grab a ball and I can throw it.
Okay, cool. So we can see we can come into here how AI thinks and we have sort of like an interactive world where we can look at tokens. You can see AI reads word
pieces. We can look at things like the model harness and we can press E to learn. >> So basically I can come to all of these little dashboards and I can ask how it
works. We also see that there's some sort of game here. Use the next word machine. the cat sat on the interesting. Okay. So, it's like I'm
dropping balls in here to see like almost like how it chooses the next word. The way that an LLM generates the next word. So, there's obviously a bit of an issue up there because um the
balls were getting stuck. Probably because I dropped so many at once. But that is kind of cool. It's a cool like little simulation of how this should work except for the balls are getting
stuck too often. But anyways, nice little interactive way to explain these concepts. So, I'm not obviously going to go into every single one of these rooms, but there's quite a few. Right here, we
have the context window, and all of these have the same sort of explainers, but they all seem to have some sort of game as well. You can see I can add blocks, and I can like see how a context
window fills up, which is pretty cool. So, like, we're getting these interactive experiences as we're learning more about different concepts. So, I'm not going to go into a ton of
these, like I said, we have agent arena, we have prompt power, team of helpers, skills workshop. Let's actually head into the skills workshop real quick and see what the game is here. It's a recipe
robot. So, we basically build a tower. Um, I'm not really controlling anything here, so that's getting boring. But anyways, this is pretty cool how it was able to go through create this 3D world
for us, and we have all of these different concepts. Memory, automation, building, model garage. Anyways, that's not too bad of an output. And now we'll open up the idea
atlas from Soul. You can see this one's a little bit different vibe. Obviously, not as realistic at all with the way that you move around and you can just like walk through the fountain, but
similar sort of feel. We have real world results. We'll pop into here real quick. Um, also I can't get rid of this overlay for some reason, which is kind of annoying. Um, like this is just way too
crowded, I would say, of an interface. We can come in here for ideas, bottleneck, outcome, baseline, delivery. I think in here we can make maybe read uh it's not letting me read these
things. Click to read. How to help. Okay. I would click with my mouse rather than like interacting with um E. But anyways, this one doesn't obviously feel as good. Like this would be harder to
spend I don't know some time in. Like the other one I think I could have sat in there for maybe 15 minutes and before I got bored. But in here I'm like already a little bit over this. It's a
little bit overwhelming. It just doesn't feel as real. And this is where you can see sort of like the design and realism come into play. So like earlier that previous example where we were in that
museum, that was definitely Opus kind of like carrying the load there. That was definitely Opus making most of that code and designing most of that actual game. So anyways, this one's going to go to
Opus 5.5 as well. And now look at the cost on this one. For Opus, this was an hour and 45 minutes. For Soul, this was about 50 minutes and Opus was 60 bucks, whereas Soul was $8.55.
So I'm going to give this one to Opus. Even though it was way more expensive, the outputs are just coming back way better. I'm just I'm pretty impressed by Opus so far. All right, so this next one
we're going to start with Soul. It's pretty interesting. I said I wanted to take a trip from October 1st to the 30th. I want to go to super cool AI events. I want to see cool nature. I'm
in Chicago. I want to also get out of the country. And I just basically gave it this really vague idea and said that I wanted to turn my itinerary into this 3D world that I can explore. So I was
really trying to see what they could do with like building experiences from, you know, pretty vague ideas. And you'll see right here that it asked me a question. And that is one thing that I did notice
about Soul more often is that it actually stopped and asked you questions a lot more than Opus 5.5. Some people might say that's good. Some people might say that's bad. I personally love when
they stop to ask questions because it means that they're really trying to understand what I want. But sometimes when you do things like a SLG goal, maybe you don't want it to ask you a
question because you want it to keep going. So usually though with what I've noticed with Codeex is when it stops ask you a question, it doesn't actually stop. It'll inject a question into the
chat, but if you don't answer it, it'll still keep going. So, I do really like the way that Soul asks questions. So, anyways, here is the output we got from Soul. We have 30 days, one orbit. We
just have this kind of globe. So, I'm not sure I'm not really getting the 3D experience I was hoping to get, but I wasn't too specific. So, anyways, this is what we see here. Um, we can see
we're going to start from Chicago to San Francisco. It's giving me a link right here. This is going to kayak.com. So, it's going to a flight. It seems like this one is going to the Hyatt Regency
in San Francisco. So, it's providing me all these links, which I did ask for. I asked for all of the links. We can see the the order is San Francisco, Yusede, Amsterdam. Not going to say that. That's
an Iceland. I don't know how to pronounce these things. South Coast, Toronto, Aangquin, Niagara Falls, Vegas, Zion, Bryce Canyon, Grand Canyon, Red Rock Canyon. We have the entire plan
here. And all of these looks like it links me back to one of the places on the actual 3D globe, which I'm able to click in, and it basically like pulls up the little itinerary for that day. And
all of these have links that I could go book things or um grab tickets or whatever it is. So, it's really not too bad for a simple itinerary. It seems to have thought through a lot and planned a
lot of stuff out. So, it's not too bad. But, I wish the 3D globe experience was better. And here we have Claude Codes version with Opus. This one, honestly, very similar, although the map seems to
be better. It's also numbered here. So, I can click in. Okay, this is definitely more of what I was looking for when I said I wanted like the globe and for it to be interactive. And you can see that
if I'm going to these different stops, actually not there. Here it's actually taking me more inside of what I'm actually doing and where I'm going. This is definitely more of what I was looking
for. Um they're planning out very similar things. You can see the order is a bit different. But this one had San Francisco Tech Week, which is something that I was hoping it would actually
include. This is like the 9th through um I don't know, a few days later. But this one, I'm not sure if it actually included San Francisco Tech Week. Yeah, this one had us in San Francisco right
at the beginning of the month and there wasn't any sort of big event here. It was literally just going to San Francisco. So, this one did better research that I would have trusted more
for saying, "Hey, I want to go to cool AI events throughout the month of October." This one definitely was doing better. Um, it also picked out some other AI events in San Jose. And then
just the experience of this interface, besides the loading of the globe being pretty slow, the experience here is significantly better. So, this is going to sound repetitive. I think overall
Opus is just a much better model. I think that it was also like I don't know. I feel like GBD 5.6 Soul was much better than Opus 5, but Opus 5.5 is now much better than GBT6 Soul. They ran for
basically the same amount of time and Opus 5.5 here was three times the cost of Soul. But here's one that's really, really, really, really interesting to me. I'm I'm starting with the pricing
just to show you how interesting this one is. So, this was a codebase. I had Astro design this experiment and I kept this all neutral at the beginning. So it didn't know it was judging an open eye
model versus a cloud model, but I had it designed this experiment with a massive codebase and then had to actually make them like fix it and analyze stuff and and repair stuff. And what came back was
that Soul did did better. Soul got 100 out of 100 whereas Opus got 97 out of 100 and passed all 30 checks whereas Opus missed one of them. And what else is interesting about this is not only
did Opus cost basically 20 times more and was two times slower, but Soul actually stopped halfway through because of some sort of security thing. Like it wouldn't let me continue on. So I had to
pause both of them and then I had to tell Asher what happened. And then Asher said, "Okay, let's switch it up a little bit and then send this prompt back to both of them." But Opus didn't stop.
Soul stopped because of some sort of cyber security thing and safety thing. But Opus didn't. And I just wanted to call that out cuz I thought that was really interesting. But what ended up
happening was Soul somehow ran for $1 on this massive codebase and did better. And so I think like it's really interesting to see how different these models are and where they excel because
with a lot of these quick tasks or a lot of these more code I don't want to say code writing tasks but maybe like code reviewing and execution and repairing soul has been really good. But then we
also saw something like at the beginning of this video where it missed one of my Fireflies transcripts where I said nit and it just gave me the wrong answer. So obviously one test like this isn't
definitive, but in this case, you know, Soul actually does come out on top. And I would say by a wide margin, you know, for 120th of the cost and getting better output, that's pretty good. All right,
so these last two, 9 and 10, have to do with browser use. And this one's very interesting to me because historically, Codeex with pretty much whatever model has almost always beat Cloud Code with
whatever model when it came to browser use. And I've loved Codeex for browser use. So let's see how this one goes. I basically gave them both the same prompt. I gave them a course structure
with the videos and descriptions and everything and I told them to make the course inside of my free school AIS. So in here, here's the classroom. You can see if I scroll down and I give this a
refresh real quick, we should see we have two new classrooms in here. We have the sole AIOS and the Opus AOS. Let's click into Soul, we can see that it uploaded all the videos. They're all
labeled as drafts 1 through 15. Each video has the description, key points, key quotes, resources, and as I flick through, we can see that all of them hopefully are uploaded correctly. It
seems like it's got all of these right and we have properly linked, you know, documentation or resources and skills and they're all named. They're all linked. So, this looks pretty good. I
don't think we're going to notice anything significantly different from these two except for maybe like the cost and time. So, let's look at the opus one. We can see, okay, the only
difference here is that it didn't make each page as a draft. It just made the overall course as a draft, which honestly I think is a little bit smarter. But anyways, um introduction,
mindset, all of these have the same description, key points, the same things are um linked. Although this one linked it here, but also linked it here. It duplicated some work. It linked things
twice, which isn't as good, right? Like I don't want it to link everything twice. That just doesn't really make too much sense. But besides that, nothing really to pick on here with this course.
And I was expecting Soul to come in and be cheaper here. But in this case, Opus 5.5 was faster and cheaper and it had that tiny little mess up. So in this case, I would be giving this one to Opus
because the results were basically the same. Opus also didn't make each individual page a draft, which I thought was a nice touch and it was cheaper and faster. So Opus and Cloud Code have
actually won a browser use use case, which I haven't seen them do in my experience with Cloud Code versus Codex in a while. So let's move on to number 10, which is the last one. It's also
another browser use case. Okay, so this use case is slowly becoming one of my favorite like benchmark tests with new models. I gave them a reference image, which I'll pull up real quick. It was
this image right here. This was an AI generated image of me and Adam Sandler. And basically what I told them to do was I told them to open up Canva and use like the drawing and painting tools to
recreate this. And look what Opus 5.5 gave us. That is actually a really really good output. Like I wasn't expecting it to be this good because when I've done this in the past, even
with models like Fable, we got something back that was horrible. And this is actually better than Astra's output that it gave me when I did this a few weeks ago. And then when I gave this to Soul,
first of all, it just tried to create an image. And I was like, "Okay, that's cheating. That's not what I wanted you to do." So I had it like I just do did it again. I gave it another prompt, but
it ended up giving me this output. And I was like, "Okay, um I'm wondering if maybe it just needs to try again because the first time it it made an image. So I let it have another try. I spun up
another session and I told it to do it again." And then it gave me this type of output again, which is just comically bad. So I'm not exactly sure what happened in the degradation from GBD6
Astra to GBD6 Soul with browser use and with this Canva test, but this is just horrific. And obviously like we don't even have to debate that Cloud Code is going to take this one as well. Opus 5.5
has just really really impressed me so far. And yeah, um you can take a look at the speed and cost here as well. Opus was slower, you know, 30 minutes slower and about four times more expensive, but
the output is just so much better. So, this is where we're at after all of those experiments. You can see that we had 1 2 3 4 5 6 7 wins for Opus 5.5 and one win for Soul and then two that just
like didn't count because I guess I messed up. That was definitely my fault. So, 7 to one OBS 5.5. What about the total stats here? Well, OPUS 5.5 was 8 minutes or sorry, 8 hours and 40 minutes
across all of the experiments. Soul was 5 hours and 51 minutes across all the experiments. So, OPUS was a little bit slower and also cost about three times as much at 213 bucks compared to Soul's
$74.46. So, once again, we kind of came into this experiment understanding that Opus is um a bigger model. It's a little bit bigger, a little bit more expensive, and
my gut was telling me that okay, I think Opus 5.5 might be a little bit better than Soul based on the pricing. Um, but but 5.6 Soul was really good and Opus 5 was really bad. So, it's a big jump here
and I'm really excited to see what's going to come next from Enthropic with maybe their next Fable model or something because they're obviously they've been really grinding on the
feedback that they've been getting on X and just like their internal use and everything. Um, and I'm honestly a little bit disappointed in Soul. It feels more like this is a GPT6
Luna or Terra. Like this doesn't feel I don't know. 5.6 Soul is really good. I thought that GBT6 soul was going to be better. I do like how the pricing is is really cheap. That was like exciting,
but now we know why it's so much cheaper. So anyways guys, to sort of wrap this up into as concise of one sentence as I can try to make, I would say Obus 5.5 is a major step up from
Obus 5. GBD6 Soul is a step down from 5.6 Soul. It feels more like a GBD6 Luna that might come out. I would trust Claude Opus more for creativity and for some judgment calls and I would maybe
want to defer some work to GBD6 Astra if I know very very specifically what I want and I just want them to execute step one, step two, step three, step four. I think a pretty cool scenario
would be using Opus 5.5 as the orchestrator that spins up and sends off very specific instructions to a bunch of little GBD6 soul workers. What I think is very cool is that Opus 5.5 and GBD6
Soul are both cheaper than their predecessors with Opus 5 and with GBD6 or sorry, GBD 5.6. I'm about to lose my mind here. I've been up all day. Um, it's late and I have just like 5.5, 6,
5.6, just all these little numbers in my head, so forgive me, but that's kind of how I feel. Anyways, that's going to do it for this one. If you guys enjoyed, you learned something new, please give
it a like. It helps me out a ton. And as always, I appreciate you guys making it to the end of the video, and I'll see you all in the next one. Thanks everyone.