I Tested Opus 5.5 vs. GPT-6 Sol on 10 Real Use Cases

summarized

TLDR

Opus 5.5 won 7 of 10 use cases against GPT-6 Soul in this head-to-head test, consistently producing better creative outputs like websites, videos, and 3D worlds. However, GPT-6 Soul outperformed on a codebase analysis task, scoring 100/100 at 1/20th the cost. The presenter concludes Opus 5.5 is a major step up from Opus 5, while GPT-6 Soul feels like a step down from GPT-5.6 Soul, making Opus 5.5 the better choice for creative work and GPT-6 Soul better for cheap, specific execution.

Key points

Opus 5.5 won 7 out of 10 use cases against GPT-6 Soul, with one win for Soul.

GPT-6 Soul scored 100/100 on a codebase analysis task at 1/20th the cost of Opus 5.5.

Opus 5.5 is double the price of GPT-6 Soul per token.

Agent interference occurred in two use cases because prompts did not isolate sessions.

Tools mentioned

Techniques

  • scroll-driven website layering
  • parallel task execution
  • sizzle reel creation
  • 3D world generation
  • browser use automation
Transcript (captions)

0:00 So, I've been playing around with Opus 5.5 and GBD6 Soul all day long, and I just ran them across 10 different use cases. I'm talking things like building websites, editing videos, building slide

0:09 decks, browser use, tons of reasoning, and then a few other things as well. So, let's not waste any time and just get straight into this video. Okay, so before we start looking at all these

0:18 comparisons, I wanted to talk about the pricing of these two models real quick. So, Opus 5.5 on the left is double the price of GBD6 Soul on the right. You can see that Opus is four bucks for the

0:27 input and 20 bucks for the output, whereas Soul is $2 and $10. Now, the reason I wanted to bring this up is because when I first glance at these prices, my instinct is like, okay, well,

0:37 Opus is probably going to be a little bit better than Soul. That's typically just what you think when something's a little bit more expensive. Now, I know a lot of you guys might disagree with that

0:45 perspective, but the way that I like to think about AI models is I like to think about, okay, if I had a hundred bucks and I gave a hundred bucks to Opus and I gave a 100 bucks to Soul, which one

0:54 would give me a better output or better quality for that 100 bucks? So, that's why I think it's important to sort of look at that metric and sort of, you know, now I have a little bit of a

1:02 feeling of what I think these models might be like, but obviously that's why we run tests. So, anyways, now that we got that out of the way, let's jump into some examples. So, real quick before we

1:10 jump into the 10 use cases, I wanted to do a real simple example just to show you what they feel like and how they think. So, on the left we have Soul and on the right we have Opus and this is

1:18 the cloud desktop app and the Codex desktop app. And I shut off this prompt that says, "Pull in my two latest Fireflies calls for Q4 planning and ingest those transcripts into the wiki."

1:27 I spelled that wrong so hopefully it knows what I meant. And I said, "Then find me my last AIS Plus Q&A where I brought up NAND." So these are two different tasks. One is ingesting calls

1:36 into a wiki and one is finding something. So, it'll be interesting to see how they think about doing these in parallel. The other thing is this isn't super super simple. These two calls

1:44 combined is about 4 hours of transcript. So, it has to create all those links and the indexes and link everything together. So, it's not super simple, but I just want to see how these models feel

1:52 and how quick and lightweight they feel as well. And you can see so far Soul seems to be cooking way ahead of Opus 5.5. It's doing these both things in parallel. Like I said, it's looking for

2:02 the new Fireflies calls and ingesting them. And it's also locating the Q&A in parallel. So this is pretty interesting. It's been about 6 minutes now and Claude came back and said, "Hey, I noticed that

2:12 something else is also doing this." So it noticed Codex doing it and Codex so far has the answer wrong about when I last mentioned Nitn. It said August 17th. But the right answer is what Opus

2:22 found which was September 14th. So anyways, I said overwrite. It's okay. Both of these models are still spinning and soul just finished as you can see and it got this answer wrong. And that's

2:30 not good, right? Like it's not that hard for a model to look through a bunch of transcripts and find the most recent one where Niden was mentioned. It's probably because it was spelled out wrong like nn

2:38 or nadnen as you can see here when it was actually transcribed. But still I feel like that is a problem that soul should have found for me and they finished up in basically the same amount

2:46 of time but opus's answer here was definitely better. So this very quick experiment gives us a little bit of a baseline to take into account as we head into the next 10 actual use cases. So

2:55 let's jump right in. Okay, so the first one was a quick website design. I gave in all 10 of these examples I gave both models the exact same prompt. So here is one output. I won't tell you which one

3:04 this was yet. We have perk form. Your morning shouldn't require two drinks. We can see as we move our mouse, the keys and the can and the background are on a different layer. We start to scroll down

3:13 and we get this little animation right here. You can see the layering come into play. And now we have a different image down here. Coffee in one hand, protein in the other. Two purchases, two things

3:21 to carry, a shaker to rinse, blah blah blah. Or we have pull the tab. I like that little animation. That's kind of cool. Um, we get some smoke coming out. Your coffee and protein one can. We keep

3:33 scrolling down here. You have text animations with all the text coming into view. Pick your first can here. Okay, a nice little rotation animation. I do think that that's actually pretty clean.

3:42 And we come down here to some FAQ. And then we have out the door. I'm not sure if that's a glitch or it just leaves frame really quick. But anyways, this was one output. Let's take a look at the

3:51 other one right here. We can see if we reload, we get your morning shouldn't require two drinks. Not as strong of a hero image right now, but we scroll in, we start to get some different text

4:00 animations as well. we get like a little opening a vault. We have a weird packaging style image rather than like a realistic looking image. So, I think that was a bad design choice there. Um,

4:09 but you can see that both of these are keeping the same sort of like narrative and the same branding. We're getting all the same sort of like text animations. We're getting um once again the

4:17 packaging pictures rather than like real product images, which I don't love. And this overall, this one just seems a lot more simple. Definitely not as premium as this first one. So, this first one

4:26 was definitely the winner when it comes to design and style. And this one was Opus 5.5 and this one was GBT6 Soul. Real quick guys, by the way, if you want the skill that I use to generate scroll

4:36 driven and layered websites just like these examples that you saw, then you can grab this for completely free. The link for that will be down in the description. Let's get back to the

4:43 video. Now, when we take a look at the pricing, we can see here they both took about 35 minutes. Opus took a little longer and Opus also cost three times as much. So, $1832 compared to 5.89. But

4:54 right now, this one, even though with the price, the winner of this experiment is going to Opus 5.5. No doubt. So, 10 Opus 5.5. Okay. So, example number two is something that I think is just a

5:05 really, really cool benchmark that I might keep doing for all these models where I basically give them this Frame.io link with tons of different videos, 100 gigabytes of videos from our

5:14 recent AIS Live event. And I basically just ask them to create a sizzle reel, a 30-cond sizzle reel that helps us market the next event, which is October 17th and 18th. If you guys are interested,

5:23 definitely check it out. I'll put a link to the event in the description. But anyways, I basically want them to look through everything, tell a story out of it, and make a high energy sizzle reel.

5:30 So, let's take a look at both outputs. This first one will be Opus 5.5. >> Hello. Hello, AIS Live. My >> Oh my god, I'm so excited for this. >> Let's Let's get some energy going in

5:41 here. It's been a great two days. Wow. Okay, that was my first time watching that and I will say that that is so so impressive. Like that almost

6:11 feels more impressive to me than when Astra did that on its first try. So, that was really, really good. I mean, look at how it did the layering here. And, you know, some of it felt a little

6:21 bit static, but the music, the sound effects, the pacing, the way it came in and out, all of this layering, I mean, I thought that this was actually really, really good. Like, definitely better

6:30 than I thought it was going to give me. All right, now we will take a look at Soul's version. This has been so energizing. Cloud code is my primary driver.

6:56 I can go to bed, come back, and it'll still be working. Okay. Um, I mean, it's still impressive, but it does not give me the same energy. It didn't have the same just like

7:10 fast-paced feel. it doesn't get me as excited for the next event compared to Opus' version. So Opus 5.5 definitely wins that in my opinion. So Opus is two for two right now. And you can see here

7:21 when it comes to time and pricing, Opus was faster, 5 minutes faster, as well as only double the cost. So even though Opus was five bucks more expensive here, I'm giving this one to Opus. Opus is 2

7:34 and 0 right now. And I'm so so glad that we're actually seeing this because Opus 5 was really bad. Like I haven't used Opus 5 in a long time. I was using 4.8 8 and Opus 5 was just so so bad. It was so

7:43 slow. But Opus 5.5 so far, I've really really been enjoying it. So, I'm feeling good about getting back a little bit more usage on my cloud subscription now. Um because I really do like both of

7:53 these harnesses and they've got different good models. Like I do really really love Astra, but Opus 5.5 for the cost, I'm getting excited about it. But anyways, let's move on to use case

8:02 number three. All right, so for use case number three, we are doing an Instagram reel. So I told it to use hyperframes. I told it to edit it. I gave both models the same footage and the same prompt.

8:09 Let's this time start with GBD6 souls output. And I'm not going to play the entire reel, but here we go. Stop prompting Claude. Andre Karpathy thinks there's a much better way to work with

8:18 AI. And his method has three layers. Layer one is the spec. Instead of giving Claude a task and hoping it understands you, work with it to create a detailed spec first. Have Claude interview you

8:28 about what you're actually trying to achieve. Then break the work into smaller checkpoints before it starts. Layer two is the verifier. Before Claude does anything, define exactly what a

8:36 good result looks like. Then give it ways to actually check its own work. whether that's another AI model or real data or tests that it can run on its own. Okay, not impressed at all. There's

8:44 a weird like humming noise in the background. There's not really that many engaging animations or there's no music at all. I think that this is a very meh output. Astra does a much better job

8:55 with editing reels and editing videos. But this is just not something that I would have put on my Instagram account for real. I would have wanted to iterate a lot more than just what we see here.

9:06 And now let's take a look at Opus 5.5's version of this reel. Stop prompting Claude. Andre Karpathy thinks there's a much better way to work with AI. And his method has three layers. Layer one is

9:14 the spec. Instead of giving Claude a task and hoping it understands you, work with it to create a detailed spec first. Have Claude interview you about what you're actually trying to achieve. Then

9:22 break the work into smaller checkpoints before it starts. Layer two is the verifier. Before Claude does anything, define exactly what a good result looks like. Then give it ways to actually

9:30 check its own work, whether that's another AI model or real data or tests that it can run on its own. Okay, that was really good. Like, this is so much better. It's more engaging. Look at

9:39 these like typing animations, the sound effects you guys noticed. All of this was more engaging. It's still not something that I would probably put on my Instagram, but this is like without a

9:48 shadow of a doubt significantly better than what we just got from GBD6 Soul. And this version was pretty interesting with the pricing cost. It took double the time with Opus 5.5 and like four

9:58 times the cost. But once again, in this case, I would pay the premium there. The question would be, okay, so Opus spent 11 bucks and Soul spent almost three bucks. If we iterated with Soul over and

10:09 over to get to 11 bucks of spend, would it be where Opus is or would it be better? I think that's the interesting question to sort of think about, but also like you love the ability to just

10:17 oneshot stuff and they both use the same skills. They both have the exact same stuff. It seems like there's just more potential over here with a lot of this sort of like design and creative stuff

10:25 so far. So anyways, right now Opus is three for three. So, we're heading into the next experiment with Opus being up 30. So, let's go on to number four. All right. So, in this example, I gave both

10:36 models a ton of data, like a ton of data, and I said, "Hey, build me a Google sheet analytics, and then build me a dashboard as well." So, this was Opus 5.5's Google Sheet. This one looks

10:44 pretty clean. You can see that we've got these different boxes with different metrics like AIR, MR, customers, you know, gross margin. We have all these stats. We have these visualizations

10:53 right here. We can move into a P&L. We can move into SAS metrics with more visualizations. We can see cash and runway, unit economics, pricing, headcount, go to market, pipeline, um

11:04 customers, and then some data notes as well. I think that this looks pretty well. It feels branded. We have, you know, clearly some sort of color theme with like black or a dark blue and

11:13 orange. And yeah, this looks pretty good. We also got a pitch deck, which I can open up right here. And you can see that we have um like a logo has been created. Once again, we have sort of

11:22 like that navy blue and dark blue or sorry, navy blue and orange. What I meant to say, we have a picture here. We've got um this stuff is coming through pretty simple, but also pretty

11:30 clean. We have more data visualizations, some more B-roll or not B-roll, images being generated. Um it's not too bad. It's not too wordy, but it's not too premium either. So, it's just all right.

11:41 And then it also created us an analytics dashboard. We can see right here that we have I mean this looks very vibe coded, very AI generated, but hey, the dashboard seems to be functional. It has

11:51 different ways for us to switch through. This is obviously using the same data that we saw on our um Google sheet that it made for us earlier. I like this highlighting effect. Um same down here.

12:01 So, not too bad of a dashboard. Oh, we can also switch between different tabs and switch between different modes as well here. So, that's pretty cool. And then the last thing you see that it

12:08 created for us here was a landing page which is running locally. And we can see we have an image here with different like you know paper and clipboard things flying around. These are on different

12:17 layers. So, that's a nice little scroll animation. Here we've got a Monday brief. We come down here into a little bit of an animation with filling in some gaps. Here we have more text coming in

12:29 dynamically. We've got a little scroll animation here. So very clean. You can tell it's obviously AI generated. It just feels a little bit AI generated, but hey, there's some nice animations in

12:38 here and some nice layering that definitely gives it a little bit of an edge. So anyways, that was Opus. Let's go over to Soul and see what it was able to do with all of this data. Okay, so we

12:47 have once again a PowerPoint, we have a Google sheet, we have an analytics dashboard and a client page. So let's take a look real quick at the actual slide deck. We can see that we have very

12:56 similar. Like I wouldn't say that anything here is too different, honestly. Like this looks almost exactly the same as the one that Opus 5.5 generated for us. And you know what I'm

13:06 actually suspicious of? Is this the exact same deck? This might be the exact same one. They might have gotten confused and gotten in each other's way there and just like one of them grabbed

13:14 the other deck, which is pretty interesting. M. Okay, so somehow along the way, these two models started working together. They probably started to work inside of the same project

13:22 folder, and I didn't yet specify in this example. I didn't say, "Hey, make sure you're not overwriting any other agents work." So, this is a really interesting point. Like, I don't know exactly what

13:31 to do here. Um, I think they built everything the exact same, which is really, really interesting. They literally created everything together. And in this case, they both still spent

13:40 money. And Soul spent way more here, actually. So, it it ran for 8 minutes longer and it cost three bucks more. So, this is honestly a little bit of a weird one. I'm not going to like assign a

13:50 winner here, and I'm sorry if you guys are like pissed off that I let them do that in this in this specific example, but they took a ton of data and they had to build kind of a story out of it and

14:00 create four deliverables and they sort of collaborated on that, which I thought was interesting. Anyways, here's what we got on the the cost and the runtime for this specific example. Okay, guys, real

14:10 quick. In the moment of filming, I didn't think anything too much of it, but while I was editing, I was like, "Wait a minute, that doesn't really make sense." So, I went back to the session

14:18 and I did the math again, and I found out that this actually was only codecs $3.60 and about 9 minutes. So, the previous figures, I was like, "Wait a minute,

14:27 that just doesn't really seem to add up." So, at the end of this video, when we're going over the total cost, just subtract about 19 bucks from um GBD6 Soul's total cost. But everything else

14:40 was pretty consistent. Once again, apologize for messing this one up. Um, I should have put that in the prompt. I normally do. Didn't really think to do it for this one for some reason. But

14:48 either way, I fixed it. And just to show you guys what happened, I had both GBT Soul and um, Opus inspect the logs on both sides and they both agreed on what happened

15:00 here, which is really interesting. Basically, I gave Claude the request and it began creating files and then Codex got the same request. obviously in the same repo. But what happened is the

15:12 first time I sent this to Codeex, it failed because like GBD6 Soul was still rolling out and I was like, "Oh, I got to get on these prompts." So it failed. And then when I resumed it, what was

15:21 this 16 minutes later? Claude was already like pretty far along the way with building. So it basically just grabbed all of the stuff that Claude was working on and then started editing

15:32 those files on top while Claude was still working. So, I truly think that some of these deliverables would have been better if maybe if Soul didn't get in there and make some changes. But, you

15:41 know, obviously it's hard to say, but I think that some of the stuff that looked a little bit AI generated, I'm wondering if maybe that's because Soul got in here and made some changes. But, I do think

15:50 that the deck looks pretty good and I think that this Google sheet does overall look still pretty solid. So, anyways, we're going to move on to the next experiment, but I wanted to at

15:58 least show you guys like the order of events there and what happened because I think that actually is, you know, pretty interesting. But anyways, moving on to use case number five, we have this game.

16:07 So, we're going to first look at um GBT6 Souls version. It created a game called Small Hours. I gave them the same prompt obviously, and this thing is loading up super slow. My computer down there also

16:17 started to get quite loud. So, this thing must be a pretty large game. Okay, so here it is. We basically will just be walking around with WD, clicking to interact, and mouse to look around. So,

16:29 I'll hit begin. We see the Ren Museum of Miniatures closed for good at 6 o'clock tonight. Everything your grandmother built goes to auction in the morning. You sat down in her reading chair to say

16:40 goodbye. Okay, we got some sound effects coming in. Wow. Okay, so this is like really telling me a story here. Okay, cool. So,

16:52 this is me able to look around, move around. Honestly, not too bad from like a design perspective. Although my mouse if it goes out of frame, that's it. Like my mouse isn't locked. You see my mouse

17:02 is nowhere. I'm moving my mouse around, but like nothing's happening until I bring it back into frame, which is a problem because if I want to do a full 360, I can't. So, that's a little bit

17:11 weird. Little bit of an issue there. I'm not even sure like how I'd go about trying to fix that. But anyways, walking around here. Okay. Click to interact. Switch on the sewing room light. Uh,

17:25 right there. Right there. Switch on the parents bedroom light. Oh, that's just telling me what I can do. Interesting. >> Wow. I mean, this is honestly unplayable

17:39 because of this mouse thing. Like, I can't I can't turn this way if I want to because my mouse goes out of screen. That's not good. But honestly, like as far as the storyline goes and the way

17:50 this feels, like it feels smooth. It's just this mouse issue, which honestly, I can't even play this anymore because of that. So, I'm going to have to stop. We're gonna go over to Opus version. You

17:58 know what's awful is the exact same thing happened here as what happened in the previous version. So they literally once again started working on each other's files. And

18:08 obviously that's like a prompt error because I didn't specify in these prompts not to edit any other files. Usually I do say that, but in this one I just kind of was shooting off these

18:16 prompts. So unfortunately we'll have to take the outputs here with a bit of a, you know, a grain of salt because they started working on the same thing. But I don't think that this experiment is just

18:24 totally ruined because the pricing is still really interesting to me because somehow we still had Opus worked for two and a half hours almost with Soul working for an hour and Opus coming in

18:33 at four times more expensive. So unfortunately guys, I'm sorry that that happened. I mean we we have seen so far that Opus does seem to be better when it comes to actually giving us these sorts

18:42 of outputs, but I think it's important to understand like where would the handoff be between these two types of models. All right. Well, okay. Well, number six, these agents did not work on

18:50 the exact same deliverable. Thank goodness. So, let's go ahead and launch these up. Basically, what I had them do is look through my past 100 YouTube videos. I'll show you the prompt up

18:58 here. And I wanted them to create me basically a 3D playable world explaining all these concepts from YouTube. I thought this was interesting to see how they reasoned about this and how they

19:07 thought about telling a story out of this. So, let's open these up and take a look. Okay, so this one is Opus' output. We'll go ahead and start exploring. We can see that we walk forward with W. Can

19:17 we Okay, we can move like this. So, we've got kind of this cartoon 3D world where we're able to walk inside and take a look. I can sprint with space. This is our

19:28 guide. We can talk to Nova >> building. >> Okay, cool. I also can throw these balls for some reason. I can grab a ball and I can throw it.

19:46 Okay, cool. So we can see we can come into here how AI thinks and we have sort of like an interactive world where we can look at tokens. You can see AI reads word

19:56 pieces. We can look at things like the model harness and we can press E to learn. >> So basically I can come to all of these little dashboards and I can ask how it

20:07 works. We also see that there's some sort of game here. Use the next word machine. the cat sat on the interesting. Okay. So, it's like I'm

20:18 dropping balls in here to see like almost like how it chooses the next word. The way that an LLM generates the next word. So, there's obviously a bit of an issue up there because um the

20:29 balls were getting stuck. Probably because I dropped so many at once. But that is kind of cool. It's a cool like little simulation of how this should work except for the balls are getting

20:37 stuck too often. But anyways, nice little interactive way to explain these concepts. So, I'm not obviously going to go into every single one of these rooms, but there's quite a few. Right here, we

20:45 have the context window, and all of these have the same sort of explainers, but they all seem to have some sort of game as well. You can see I can add blocks, and I can like see how a context

20:54 window fills up, which is pretty cool. So, like, we're getting these interactive experiences as we're learning more about different concepts. So, I'm not going to go into a ton of

21:02 these, like I said, we have agent arena, we have prompt power, team of helpers, skills workshop. Let's actually head into the skills workshop real quick and see what the game is here. It's a recipe

21:13 robot. So, we basically build a tower. Um, I'm not really controlling anything here, so that's getting boring. But anyways, this is pretty cool how it was able to go through create this 3D world

21:23 for us, and we have all of these different concepts. Memory, automation, building, model garage. Anyways, that's not too bad of an output. And now we'll open up the idea

21:35 atlas from Soul. You can see this one's a little bit different vibe. Obviously, not as realistic at all with the way that you move around and you can just like walk through the fountain, but

21:44 similar sort of feel. We have real world results. We'll pop into here real quick. Um, also I can't get rid of this overlay for some reason, which is kind of annoying. Um, like this is just way too

21:54 crowded, I would say, of an interface. We can come in here for ideas, bottleneck, outcome, baseline, delivery. I think in here we can make maybe read uh it's not letting me read these

22:04 things. Click to read. How to help. Okay. I would click with my mouse rather than like interacting with um E. But anyways, this one doesn't obviously feel as good. Like this would be harder to

22:14 spend I don't know some time in. Like the other one I think I could have sat in there for maybe 15 minutes and before I got bored. But in here I'm like already a little bit over this. It's a

22:23 little bit overwhelming. It just doesn't feel as real. And this is where you can see sort of like the design and realism come into play. So like earlier that previous example where we were in that

22:32 museum, that was definitely Opus kind of like carrying the load there. That was definitely Opus making most of that code and designing most of that actual game. So anyways, this one's going to go to

22:42 Opus 5.5 as well. And now look at the cost on this one. For Opus, this was an hour and 45 minutes. For Soul, this was about 50 minutes and Opus was 60 bucks, whereas Soul was $8.55.

22:52 So I'm going to give this one to Opus. Even though it was way more expensive, the outputs are just coming back way better. I'm just I'm pretty impressed by Opus so far. All right, so this next one

23:00 we're going to start with Soul. It's pretty interesting. I said I wanted to take a trip from October 1st to the 30th. I want to go to super cool AI events. I want to see cool nature. I'm

23:08 in Chicago. I want to also get out of the country. And I just basically gave it this really vague idea and said that I wanted to turn my itinerary into this 3D world that I can explore. So I was

23:16 really trying to see what they could do with like building experiences from, you know, pretty vague ideas. And you'll see right here that it asked me a question. And that is one thing that I did notice

23:24 about Soul more often is that it actually stopped and asked you questions a lot more than Opus 5.5. Some people might say that's good. Some people might say that's bad. I personally love when

23:34 they stop to ask questions because it means that they're really trying to understand what I want. But sometimes when you do things like a SLG goal, maybe you don't want it to ask you a

23:41 question because you want it to keep going. So usually though with what I've noticed with Codeex is when it stops ask you a question, it doesn't actually stop. It'll inject a question into the

23:49 chat, but if you don't answer it, it'll still keep going. So, I do really like the way that Soul asks questions. So, anyways, here is the output we got from Soul. We have 30 days, one orbit. We

23:58 just have this kind of globe. So, I'm not sure I'm not really getting the 3D experience I was hoping to get, but I wasn't too specific. So, anyways, this is what we see here. Um, we can see

24:07 we're going to start from Chicago to San Francisco. It's giving me a link right here. This is going to kayak.com. So, it's going to a flight. It seems like this one is going to the Hyatt Regency

24:18 in San Francisco. So, it's providing me all these links, which I did ask for. I asked for all of the links. We can see the the order is San Francisco, Yusede, Amsterdam. Not going to say that. That's

24:28 an Iceland. I don't know how to pronounce these things. South Coast, Toronto, Aangquin, Niagara Falls, Vegas, Zion, Bryce Canyon, Grand Canyon, Red Rock Canyon. We have the entire plan

24:37 here. And all of these looks like it links me back to one of the places on the actual 3D globe, which I'm able to click in, and it basically like pulls up the little itinerary for that day. And

24:47 all of these have links that I could go book things or um grab tickets or whatever it is. So, it's really not too bad for a simple itinerary. It seems to have thought through a lot and planned a

24:56 lot of stuff out. So, it's not too bad. But, I wish the 3D globe experience was better. And here we have Claude Codes version with Opus. This one, honestly, very similar, although the map seems to

25:06 be better. It's also numbered here. So, I can click in. Okay, this is definitely more of what I was looking for when I said I wanted like the globe and for it to be interactive. And you can see that

25:14 if I'm going to these different stops, actually not there. Here it's actually taking me more inside of what I'm actually doing and where I'm going. This is definitely more of what I was looking

25:23 for. Um they're planning out very similar things. You can see the order is a bit different. But this one had San Francisco Tech Week, which is something that I was hoping it would actually

25:31 include. This is like the 9th through um I don't know, a few days later. But this one, I'm not sure if it actually included San Francisco Tech Week. Yeah, this one had us in San Francisco right

25:41 at the beginning of the month and there wasn't any sort of big event here. It was literally just going to San Francisco. So, this one did better research that I would have trusted more

25:48 for saying, "Hey, I want to go to cool AI events throughout the month of October." This one definitely was doing better. Um, it also picked out some other AI events in San Jose. And then

25:57 just the experience of this interface, besides the loading of the globe being pretty slow, the experience here is significantly better. So, this is going to sound repetitive. I think overall

26:08 Opus is just a much better model. I think that it was also like I don't know. I feel like GBD 5.6 Soul was much better than Opus 5, but Opus 5.5 is now much better than GBT6 Soul. They ran for

26:20 basically the same amount of time and Opus 5.5 here was three times the cost of Soul. But here's one that's really, really, really, really interesting to me. I'm I'm starting with the pricing

26:29 just to show you how interesting this one is. So, this was a codebase. I had Astro design this experiment and I kept this all neutral at the beginning. So it didn't know it was judging an open eye

26:38 model versus a cloud model, but I had it designed this experiment with a massive codebase and then had to actually make them like fix it and analyze stuff and and repair stuff. And what came back was

26:50 that Soul did did better. Soul got 100 out of 100 whereas Opus got 97 out of 100 and passed all 30 checks whereas Opus missed one of them. And what else is interesting about this is not only

27:01 did Opus cost basically 20 times more and was two times slower, but Soul actually stopped halfway through because of some sort of security thing. Like it wouldn't let me continue on. So I had to

27:12 pause both of them and then I had to tell Asher what happened. And then Asher said, "Okay, let's switch it up a little bit and then send this prompt back to both of them." But Opus didn't stop.

27:20 Soul stopped because of some sort of cyber security thing and safety thing. But Opus didn't. And I just wanted to call that out cuz I thought that was really interesting. But what ended up

27:27 happening was Soul somehow ran for $1 on this massive codebase and did better. And so I think like it's really interesting to see how different these models are and where they excel because

27:38 with a lot of these quick tasks or a lot of these more code I don't want to say code writing tasks but maybe like code reviewing and execution and repairing soul has been really good. But then we

27:49 also saw something like at the beginning of this video where it missed one of my Fireflies transcripts where I said nit and it just gave me the wrong answer. So obviously one test like this isn't

27:58 definitive, but in this case, you know, Soul actually does come out on top. And I would say by a wide margin, you know, for 120th of the cost and getting better output, that's pretty good. All right,

28:07 so these last two, 9 and 10, have to do with browser use. And this one's very interesting to me because historically, Codeex with pretty much whatever model has almost always beat Cloud Code with

28:18 whatever model when it came to browser use. And I've loved Codeex for browser use. So let's see how this one goes. I basically gave them both the same prompt. I gave them a course structure

28:26 with the videos and descriptions and everything and I told them to make the course inside of my free school AIS. So in here, here's the classroom. You can see if I scroll down and I give this a

28:36 refresh real quick, we should see we have two new classrooms in here. We have the sole AIOS and the Opus AOS. Let's click into Soul, we can see that it uploaded all the videos. They're all

28:45 labeled as drafts 1 through 15. Each video has the description, key points, key quotes, resources, and as I flick through, we can see that all of them hopefully are uploaded correctly. It

28:55 seems like it's got all of these right and we have properly linked, you know, documentation or resources and skills and they're all named. They're all linked. So, this looks pretty good. I

29:05 don't think we're going to notice anything significantly different from these two except for maybe like the cost and time. So, let's look at the opus one. We can see, okay, the only

29:14 difference here is that it didn't make each page as a draft. It just made the overall course as a draft, which honestly I think is a little bit smarter. But anyways, um introduction,

29:22 mindset, all of these have the same description, key points, the same things are um linked. Although this one linked it here, but also linked it here. It duplicated some work. It linked things

29:33 twice, which isn't as good, right? Like I don't want it to link everything twice. That just doesn't really make too much sense. But besides that, nothing really to pick on here with this course.

29:42 And I was expecting Soul to come in and be cheaper here. But in this case, Opus 5.5 was faster and cheaper and it had that tiny little mess up. So in this case, I would be giving this one to Opus

29:53 because the results were basically the same. Opus also didn't make each individual page a draft, which I thought was a nice touch and it was cheaper and faster. So Opus and Cloud Code have

30:03 actually won a browser use use case, which I haven't seen them do in my experience with Cloud Code versus Codex in a while. So let's move on to number 10, which is the last one. It's also

30:11 another browser use case. Okay, so this use case is slowly becoming one of my favorite like benchmark tests with new models. I gave them a reference image, which I'll pull up real quick. It was

30:21 this image right here. This was an AI generated image of me and Adam Sandler. And basically what I told them to do was I told them to open up Canva and use like the drawing and painting tools to

30:30 recreate this. And look what Opus 5.5 gave us. That is actually a really really good output. Like I wasn't expecting it to be this good because when I've done this in the past, even

30:39 with models like Fable, we got something back that was horrible. And this is actually better than Astra's output that it gave me when I did this a few weeks ago. And then when I gave this to Soul,

30:48 first of all, it just tried to create an image. And I was like, "Okay, that's cheating. That's not what I wanted you to do." So I had it like I just do did it again. I gave it another prompt, but

30:56 it ended up giving me this output. And I was like, "Okay, um I'm wondering if maybe it just needs to try again because the first time it it made an image. So I let it have another try. I spun up

31:05 another session and I told it to do it again." And then it gave me this type of output again, which is just comically bad. So I'm not exactly sure what happened in the degradation from GBD6

31:16 Astra to GBD6 Soul with browser use and with this Canva test, but this is just horrific. And obviously like we don't even have to debate that Cloud Code is going to take this one as well. Opus 5.5

31:29 has just really really impressed me so far. And yeah, um you can take a look at the speed and cost here as well. Opus was slower, you know, 30 minutes slower and about four times more expensive, but

31:40 the output is just so much better. So, this is where we're at after all of those experiments. You can see that we had 1 2 3 4 5 6 7 wins for Opus 5.5 and one win for Soul and then two that just

31:53 like didn't count because I guess I messed up. That was definitely my fault. So, 7 to one OBS 5.5. What about the total stats here? Well, OPUS 5.5 was 8 minutes or sorry, 8 hours and 40 minutes

32:04 across all of the experiments. Soul was 5 hours and 51 minutes across all the experiments. So, OPUS was a little bit slower and also cost about three times as much at 213 bucks compared to Soul's

32:14 $74.46. So, once again, we kind of came into this experiment understanding that Opus is um a bigger model. It's a little bit bigger, a little bit more expensive, and

32:22 my gut was telling me that okay, I think Opus 5.5 might be a little bit better than Soul based on the pricing. Um, but but 5.6 Soul was really good and Opus 5 was really bad. So, it's a big jump here

32:34 and I'm really excited to see what's going to come next from Enthropic with maybe their next Fable model or something because they're obviously they've been really grinding on the

32:42 feedback that they've been getting on X and just like their internal use and everything. Um, and I'm honestly a little bit disappointed in Soul. It feels more like this is a GPT6

32:50 Luna or Terra. Like this doesn't feel I don't know. 5.6 Soul is really good. I thought that GBT6 soul was going to be better. I do like how the pricing is is really cheap. That was like exciting,

33:02 but now we know why it's so much cheaper. So anyways guys, to sort of wrap this up into as concise of one sentence as I can try to make, I would say Obus 5.5 is a major step up from

33:13 Obus 5. GBD6 Soul is a step down from 5.6 Soul. It feels more like a GBD6 Luna that might come out. I would trust Claude Opus more for creativity and for some judgment calls and I would maybe

33:31 want to defer some work to GBD6 Astra if I know very very specifically what I want and I just want them to execute step one, step two, step three, step four. I think a pretty cool scenario

33:42 would be using Opus 5.5 as the orchestrator that spins up and sends off very specific instructions to a bunch of little GBD6 soul workers. What I think is very cool is that Opus 5.5 and GBD6

33:54 Soul are both cheaper than their predecessors with Opus 5 and with GBD6 or sorry, GBD 5.6. I'm about to lose my mind here. I've been up all day. Um, it's late and I have just like 5.5, 6,

34:07 5.6, just all these little numbers in my head, so forgive me, but that's kind of how I feel. Anyways, that's going to do it for this one. If you guys enjoyed, you learned something new, please give

34:14 it a like. It helps me out a ton. And as always, I appreciate you guys making it to the end of the video, and I'll see you all in the next one. Thanks everyone.

Frontier News · by Hyperjump Technology