Claude Fable 5.1 vs GPT-6 Astra. Which One Should You Actually Use?

summarized

TLDR

Claude Fable 5.1 beat GPT-6 Astra 84 to 81 in a 10-job test covering admin, email, research, website, content, tool building, contract review, application, trip planning, and video creation. Claude was more accurate and usable but needed more hand-holding; GPT was faster and more autonomous but made more errors. The biggest lesson: neither AI can reliably check its own work or supply missing facts, so human judgment is still critical.

Key points

Claude Fable 5.1 scored 84 out of 100 in a 10-job comparison test against GPT-6 Astra's 81.

Claude required six interventions while GPT required four during the test.

Neither AI could reliably check its own work or supply facts not present in the source material.

The quality of the prompt, especially defining the finish line and standards, was the biggest factor in output quality.

In the contract review job, Claude ranked risks by practical risk, while GPT did not.

Tools mentioned

Transcript (captions)

0:00 So, I gave Claude Fable 5.1 and GPT-3 6 Astra 10 jobs that most people would either lose a weekend to get done or pay someone else to do. 200 emails, 50 prospects, a website, a working tool, an

0:15 application, and an entire business trip. I'm going to give them the same jobs, the same files, and I'm going to give them no coaching. I will track what they finished, what they got wrong, and

0:25 every time they handed the work back to me. Because let's be honest, an AI that saves 2 hours but needs 3 hours of supervision has saved you nothing. And I expect that one of these 10 is going to

0:37 be broken, maybe by both of them. It is the one that I would have trusted them with the most, but I expect that they might have some struggles. So, let's get into it. Thank you to HubSpot for

0:50 partnering with us on today's video. So, here are the rules of the game. First of all, both systems get the same job, the same files, and equivalent test accounts. I will not coach one more than

1:02 the other. Number two, one attempt each. If the output is weak, the score is going to reflect that. There is one planned exception, and you will know which one I'm talking about.

1:14 Number three, neither system makes an irreversible decision for me. That covers emails, applications, purchases, bookings. And number four, every job is worth 10 points, [music] two in each of

1:29 the five categories. We have 100 points available, and I'm also counting interventions, and every unnecessary question will count as a point. Because every time the AI hands the work back to

1:41 me, I'm working again. If the scores tie, we'll see who had the fewer interventions. And there will be a thing running in the background. I'm going to ask both of them every hour to check the

1:52 Brightwater inbox, the calendar, and only tell me if something actually needs me. That last part is really important and that is the test actually because anything can send me a list. I want to

2:03 know whether it can tell the difference between something that matters and everything else. Now, we will find out later whether either of them comes back. So, let's start with the pile that most

2:13 people keep moving from one week to the next. And if you see me looking down, it's because my computer is right here. Okay? So, for the first three jobs, I am going to type basically the way that

2:24 most of us type when short instructions, no detail, no defined output. So, my entire instruction is going to be please sort out this entire week of unfinished admin. That's it, essentially. The

2:38 folder has a full week of loose ends, documents to organize, expenses to classify, dates to find, actions to extract, and things that need my decision. This is going to test whether

2:47 they can inspect a messy workspace, decide what belongs together, and make progress without asking me to design the process for them. So, let's see what they're doing. So, both of them are

2:57 thinking. All right, so Claude is already connected to the tools. It's not telling me anything about what it's doing. ChatGPT is walking me through the steps. Okay, that's not bad. All right,

3:10 so Claude also included the self-reflection in the end and it explained to me what it did and suggests to create a skill and notices that nothing needs to be stored in memory,

3:24 which is absolutely great. >> [music] >> It did produce more outputs. So, from Claude's side, we have a read me about what it did and what I need,

3:36 weekly admin routine, expenses tracker. So, let's see how the tracker looks like. I mean, it's pretty simple, but it does the job. It has the tabs on expenses, invoices, statements, as well

3:49 as open questions. That's one thing that I always like about Claude that it does not make any assumptions and it always asks me. Now, let's take a look at chat. So, after explaining step-by-step what

4:00 it was doing, it also put together one spreadsheet. Right, what our admin reconciled. Okay, let's see how that one looks like. I mean, this one is nicely formatted as well and it says actions,

4:13 expenses, bank, invoices and rates, file index and original tracker. So, this was clearly a pretty straightforward easy task and both of them made really good progress without much help. They didn't

4:26 hand back anything to me, but they did make quite a few assumptions. So, in terms of judgment, I will have to prioritize Claude in this case rather than ChatGPT because Claude explained to

4:39 me afterwards everything and also left quite a few things that it didn't want to decide on my behalf. So, let's go to our ranking. In terms of completion, I think both of them did a great job. I'm

4:51 going to give two points to both of them. In terms of accuracy, as far as I can tell, both look really good, so two points to both. [music] In terms of judgment, I will give two points to

5:03 Claude and one point to chat simply because I prefer the way Claude approaches this kind of situations where it explains to me everything that I need to decide on without it deciding for me.

5:16 In terms of autonomy, two points to both. And in terms of usability, I mean, they're not bad, but I'm going to give both of them 1.5. I'm already hehe going against my own rules, but I think both

5:30 of them deserve good score, but not perfect score because they are usable, but they're not something that I would send to anyone. So, 1 and 1/2 points. Now, clearly the

5:43 visible pile has gone. The question is how much of it has actually been resolved? And if this was one of my actual tasks, it would take me a little while to go through and check and

5:54 cross-check everything, but we will assume that they did a good job. So, let's move to the next one. So, next we have a test inbox with 200 emails. I'm going to give them the same

6:06 kind of instruction. Clear these 200 emails for me. >> [music] >> Okay? They are the same newsletters, the same replies, the same dates, tasks,

6:15 everything is the same. >> [music] >> And one of them says a supplier has changed their bank details. Now, let's see if either of the two is going to

6:23 notice that. Because clearly, that kind of email should never be treated like ordinary admin. A safe system flags it and asks for independent verification. Now, let's see whether they do notice

6:35 it. All right? Very brief prompt. We're going to talk about that in a moment. Both of them are working on it. Okay, so Claude is already asking me if I allow it to modify the labels in my inbox. I'm

6:49 going to say allow once. Chad is already going straight telling me it located the test set, it found all the emails, and of those 145 are routine or reference, and 55 need follow-up.

7:05 Okay. So, something very interesting happened because both of them were working on the same Gmail account with the same emails. Chad was faster, so it was able to actually go and do the work.

7:20 So, it cleared all the emails, it labeled all of them, but it did not notice that supplier email. At the same time, Claude is almost panicking here telling me that something has happened

7:35 and it did find the 200 emails, but they were unlabeled and now they're labeled, and that they were moved, but it didn't move anything. So, something happened and it's giving me a recommendation on

7:48 how to restore everything. So, I mean, clearly, speed-wise and completion-wise, Claude is getting zero and Chat is getting two because, clearly, Chat was able to go ahead and

8:00 complete the job. In terms of accuracy, I mean, both of them are right, but I'm going to give both of them only one point because Claude is accurate, but it didn't really do anything with the

8:12 emails, and Chat [music] is also accurate, but did not notice the supplier one. I can't keep Claude accountable for that because it didn't have enough time to read the emails

8:23 since Chat did it first. So, I'm going to give both of them one for accuracy. For judgment, I am actually going to give one point to Chat and two points to Claude just because they missed or Chat

8:36 missed the supplier email. However, when it comes to autonomy, Chat gets two points, Claude only gets one simply because Claude asked me already to give it access and also didn't manage to get

8:50 through and do what it needed to do on its own. And in terms of usability, I mean, I'm going to give Chat two points and Claude one point even though I probably should give them a bit lower

9:02 than that. Okay, that was a surprise. Clearing an inbox is easy to demonstrate with a number. The better test is what remains, what was misclassified, and whether the system knew when to stop.

9:14 And clearly, both of them have some limitations there. That could also be my prompt, so we're going to come back to that. Job number three, okay. We are moving from admin into growth. So,

9:25 Brightwater Operations is a one-person operations consultancy in Sheffield. And I'm going to ask it to find 50 people worth calling and tell me what to say to each.

9:36 >> [music] >> That's it. I'm not going to give any bio definition, any format, any explanation of what makes somebody worth calling. And we need to see what both of them do.

9:46 Okay, so both of them did return 50 names. >> [music] >> That I can give to them. But when you read through, you can see that Claude is

9:56 giving me names of existing customers and following up on invoices and telling me how to follow up and what to say to each of those people. And then Chat actually pulled together a list that is

10:13 totally unusable because it does give me some names, but there are no sources and no contact details. So, yeah. Based on that, I would probably have to give both of them zero or one across all

10:27 five categories. But I'm not going to. After these three tests, I believe Claude is leading, but I have realized something uncomfortable. All three of those instructions were the way that

10:37 >> [music] >> people type when they're busy. I gave both systems an intention and expected them to invent the finish line. That is one reason that outputs are so unusable.

10:48 And actually this applies to job number one and job number two just as much as it applies to job number three. But I'm going to do something different in a moment. Before I do that, one of the

10:59 jobs still coming broke both of them. You're going to see that. It is the one that I would have trusted them with probably the most and I'm still not sure either of them should be allowed near

11:10 it. Okay, so we just watched both of these models run whole jobs start to finish. And if you're sitting there thinking, "Right, I'm convinced, but what do I actually give it to do

11:20 tomorrow morning?" Then this free resource from HubSpot called 285 prompts to reclaim your time gives you the instructions instead of making you build them all from scratch. And it's more

11:30 than just the 285 prompts. It's 26 complete playbooks, and each of them is one whole task or even job on your team written out end-to-end, which is exactly what both of these tools are waiting

11:42 for. They're not waiting for a clever prompt. They are waiting for a job, and that is where most people are still stuck right now. Both of these will take a multi-step job, go away and run it,

11:54 check their own work, and come back with the finished thing. What neither of them will do is decide which job. So, you end up sitting in front of something that is so capable, asking it for a summary,

12:04 because that's the size of request people are used to making. But, that's not you. HubSpot already wrote the 26 jobs out for you, and every one of them runs in either tool. So, whichever of

12:15 these two is going to win and you choose to use, you're covered. Inside, you will find an administrative task slasher, which is the pile most of us are genuinely drowning in. There's an SOP

12:27 generator toolkit, so the thing you keep explaining to people gets written down once and stays written. And my favorite one is the hiring and delegation accelerator, to hire right and hand off

12:37 faster. So, once you're decided which tool you want to use, this gives you the jobs and workflows to start with. Grab it for free through the link below, and big thanks to HubSpot for partnering

12:47 with us on today's video. So, let's go back. I'm going to give both of them the same job again with a finish line, with source requirements, and a standard for good work. The original scores are going

12:58 to be there, but this is the control. I mean, the job has not changed, but the quality of my brief clearly has. So, let's see if they do anything different. Okay, both of them are thinking, which

13:09 is a good sign. I gave them more to think about. All right, so clearly now Claude gave me 50 prospects based on the buyer criteria. It went through lots of sources, it went through all of my

13:23 guidelines, and it gave me 19 fully verified, not 50, which is fine. I'm not going to be mad about that. Let's see how they look. I mean, they don't look impressive, but I have all the

13:37 information that I asked for. So, clearly the better prompting worked. Now, on the chat side, now it did a lot better. It has completed 50 researched candidates, and let's see how the

13:52 information looks like. Okay, clearly, now it gave me verified or unverified, it told me whether it qualifies or it fits with the concern. Definitely both of them improved, so great job on that.

14:07 Let's move to job number four. So, our case study for today, Brightwater, has no website. They have three services: the walk, which is 2300, the fix, which is 4600 per block, and the check, which

14:20 is 850 a month. Both systems get the same business documents, the same brand guide, the customer information, and the same brief. So, let's see what they do when I ask them to build a website. I'm

14:32 going to let them work, but what I'm interested in finding is can it understand a business from a set of documents? Can it make [music] design decisions? And can it write accurate

14:40 copy? And bonus points if it works when I click. Okay, so let's see what both of them built. Interestingly, although usually I prefer chats' designs and visuals, this time I think I like

14:56 Claude's one better. Clearly, [music] Claude understood much better what my request was and read more carefully the briefing. We have a menu at the top that says the walk, the fix, and the check.

15:11 It talks about the owner, whose name is Dave, and I think the copy is much better. On chat's side, I mean, it looks clean and minimalist. It does use the company brand guidelines, but I think

15:26 it's a little bit inaccurate because the pricing is off. So, let's score them. In terms of completion, clearly both of them did their job. In terms of accuracy, unfortunately, Chat is going

15:39 to get zero because it even missed the pricing. And Claude is getting two because it got pretty much everything out of my briefing. From a judgment perspective, one for Chat, two for

15:52 Claude. Autonomy-wise, I'm going to give two points to both of them. They did a great job. They didn't ask me for anything. But then in terms of usability, clearly zero for Chat and two

16:02 for Claude because we cannot go ahead with the wrong prices. But I mean, at the end of the day, both of them made websites that I could carry on working with. The difference came down to the

16:12 judgment and really getting the copy properly and the accuracy. Oh, the schedule task is here. >> [laughter] >> So, that is what we set up earlier. I

16:23 set it before we started job number one and gave neither system another reminder. And now, probably about 15 to 20 minutes in, a new lead landed in the Brightwater inbox. Neither of them was

16:36 told that that was coming. So, great job, Claude. Clearly, it ran at the scheduled time. It noticed the email coming in and it let me know. >> [music]

16:46 >> I didn't have to go back to the AI. The AI came back to the work. That is a small job and it changes the relationship. Now, let's move to job number five. I'm going to give them one

16:57 long talk plus the Brightwater brand and tone documents. I'm going to ask them to turn it into 30 finished pieces because generating a post is easy, but generating 30 without repeating

17:10 yourself, drifting from the source, or lowering the standard is much harder. So, let's see what they're doing. Again, they're thinking. So, Chat is asking me to upload the talk, whereas Claude

17:23 already found it. And from Claude, I'm already getting written pieces, whereas Chat has started generating visuals. Hmm. Okay, very, very interesting. First of all, I have to say that I am so happy

17:38 with the fact that Claude, because it's onboarded properly, gives me all the notices about whether we need to save anything to memory, whether we need to save a routine, or create a skill based

17:50 on the work that we've done. Whereas Chat does not do that. Now, I get a lot of content from Claude. I get all the written pieces, video scripts and captions, image captions, publishing

18:03 calendar, a read me with all the decisions, as well as all the visuals. Whereas Chat has put everything in here in the chat itself, which means if I need to share this with someone on my

18:17 team, I'm going to have to at least tell it to package it and send it wherever I need them to be sent. Now, let's see the quality. Visual-wise, I think Claude has done pretty standard work. It's not bad.

18:32 I'm actually quite positively surprised. I think if we were to use Claude design, it would do significantly better. And because this is a test account, there are no design skills. And if we look at

18:44 what Chat did, I think the visuals are better, but let's see about the copy. I mean, both of them have diverse copy. So, I'm not going to hold them accountable for diversifying, because I

18:57 think they did a pretty good job. I do tend to like the visuals that Chat created better than the ones that Claude did, simply because of the style. But I think both of them are quite usable. So,

19:07 maybe in terms of usability, I'm going to deduct 0.5 from Claude's side, but other than that, I'm going to give both of them two across the board. Let's move to the next job. Now, Brightwater

19:19 compares its available capacity against the work that different client projects need. That happens by hand every month. So, we need to change that. And by that, I mean we're going to ask them to build

19:30 a little tool. So, Claude is asking me to add the folder, whereas Chad is already going ahead and using everything in the drive. I mean, both of them look quite decent. I'm actually not upset at

19:46 all. I think they both did a really, really good job. Let me see again. I think maybe Claude's one is a little bit too busy, but I like the fact that it has the summary at the top, whereas the

19:59 Chad one is more minimalist, seems easier to use, but it doesn't give me the overview. Now, let's try something. I'm going to ask both of them to add a new functionality. Let's see if that

20:12 works or if it breaks anything. Okay. So, actually both of them went straight through. Um interestingly, Claude is telling me that

20:20 the fix does not fit as planned, and that five of six weeks break a limit, whereas Chad has not told me anything about that. So, we have all these conflicts and flags that Claude

20:33 identified, whereas Chad has not. So, I'm only going to deduct a point from Claude's design [music] and autonomy, and I will deduct a point from Chad's accuracy. Let's move to the next one.

20:46 >> [music] >> So, after six jobs, we have tested files, email, research, connectors, scheduled work, long documents, design, website design, content production, and

20:57 app building. The last four jobs need something different. They need judgment, because an output can be complete and polished and entirely wrong. So, job seven is going to help us decide whether

21:09 they're doing a similar job or maybe one is better. This is where comparison turned. An 11-page master services agreement that Brightwater has already signed. Schedules, appendices,

21:23 everything. So, let's see if they are able to advise us on what is right and what might be concerning about this master agreement. Both of them are reading. Okay. So, chat is much, much

21:36 faster, clearly. Um and it is telling me everything that it read and goes through the clauses and it tells me what's critical and suggests

21:48 what to change. Now, if we look at Claude's side, it also explains what it did and what is concerning and what is fine as drafted, but it's also giving me a separate document as the contract

22:02 review. Wow, and Claude even ranked them by practical risk, which I find extremely, extremely useful, especially in uh situations of negotiation, because you know which ones to insist on and

22:14 which ones you can be a bit more relaxed on. So, for this one, I am actually going to give two points across the board for Claude and I'm going to deduct a point from Chat's

22:24 >> [music] >> ranking for usability and for judgment. I am actually quite disappointed in um the way it approached this. And honestly, this is the job that I would

22:35 have been most tempted to trust because the output looked organized, specific, and professional, but that appearance makes omissions harder to notice. So, unfortunately,

22:47 >> [music] >> Chat quite disappointed me here. Now, let's move to job number eight. This is an application. So, the system has the form and Brightwater's document and the

22:58 evidence in the kit and I am going to ask it to fill in all the fields that need to be filled in. Let's see what they do. I actually don't expect them to fill in everything. Sometimes leaving a

23:10 field empty is the right decision. So Chat noticed that several required information and details are missing. Let's see about Claude. It's giving me three files. Let's see what's inside. So

23:24 let's take a look. Companies House number not evidenced not evidenced register address not evidenced. I'm glad that they did not make up any of the information and they did notice that

23:37 some of the important ones were not provided so they did not create anything from their own mind because dangerous AI fills in gaps

23:49 on its own. So this time I'm going to give both of them a two. We're getting close to the end. So job number nine is to plan the entire trip. So Pricewater has to get from Sheffield to Birmingham.

24:02 I'm going to give it fixed dates for meetings that cannot move. I'm going to give it a total budget of 450 pounds. So Claude is asking for further information and so is Chat. I did get an itinerary

24:17 from both of them. Let's see what they put in. Okay, Birmingham trip. It's giving me the schedule material access risk the exchange publishers Thursday building hours. Okay, so it did notice

24:28 that there is an access potential risk and then it tells me what remains blocked is the hotel check out both trains are selected at the contact information stage and not

24:39 payment. I like that. That's [music] great. Let's see what Chat did. Actually Chat did a much nicer job visually because it is able to integrate with booking.com and it shows more beautiful

24:53 outputs. So actually for this one I'm going to give Chat straight twos across the board. And for Claude, I am only going to give it a two for judgement, for

25:05 noticing that access risk, and I'm going to give it a one for everything else. I'm sorry, Claude. I think Chat did a much better job in fitting in the budget and finding good recommendations and

25:17 making sure that the schedule works. Now, let's move to job number 10. The last job uses the business both systems have already built. A 60-second promo video for the Bright Water website from

25:29 job four. I'm going to give them no footage to edit. They have to develop the concept, >> [music] >> create the visuals, write the script,

25:36 generate the video, and keep it all consistent with the site that they made. Let's see if they're able to do it. Okay, so Claude is asking me for access, but connected to Higgs field, it was

25:48 able to generate everything. Promo 60-second, it actually generated two versions for me to compare. Gave me the separate captions file, and the video looks quite good. Let's see about Chat.

26:01 So, it gave me the script and it gave me the video. The video itself is not that impressive, however. So, yeah. >> [laughter]

26:10 >> I think for this one, I'm going to give both of them a one across the board, mostly from a usability point of view. I think judgement-wise, they could have done better, too, in terms of picking

26:23 what to put in the video. So, yeah, I think both of them have a bit more work to do, and I think one shot is not ideal when it comes to video creation. Also, neither of the two gave me the shot plan

26:37 like I asked, so yeah, I think I'm comfortable giving both of them a one across the board. Now, after 10 jobs, three responsibilities neither system could

26:47 take from me. First, neither could reliably check its own work. Both inspected what they had made, both said it was finished, and we could tell in several instances, especially depending

27:00 on the prompt quality, that there were errors. After we improved to the prompting, self-review improved as well, and the output was better. It did not remove the need for someone else to

27:11 look, however. Second, neither could supply facts that were not in the source. That showed up in the application, the website, and the contract. A complete-looking answer can

27:21 still contain an invented detail or a missing risk. Number three, neither carries the consequence. The system does not lose the client when an email goes wrong. It does not miss the meeting when

27:33 the connection fails. It does not surrender its own intellectual property when a clause is overlooked. That stays with the person who presses approve or send. The goal is not constant

27:45 supervision. The goal is a clear brief, visible approval points, and the review that is proportionate to the risk. So, across 10 jobs, Claude Fable 5 scored 84 out of 100, and GPT-6 Astra scored 81

28:00 out of 100. Claude required only six interventions, I think. Astra required only four. Now, Claude wins because it completed more of the work to the required standard, even though it handed

28:17 a couple more decisions back to me. Personally, I prefer that. I don't want it to ask me everything, but I prefer that I make the decisions. The work that it did is more accurate and more usable,

28:29 even though it needed a little bit more hand-holding. Now, as you've seen, jobs 1 to 3 showed what happens when I type quickly and leave the outcome undefined. The rest showed what these systems can

28:42 do when they get the context, the constraints, and the standard the work actually requires. The winner tells you which system performed better in this test. The bigger advantage, however,

28:54 comes from knowing how to give either one work that it can finish, how to define good, and where human judgment is still important. That is the difference between using AI for isolated tasks and

29:05 having it take real work off your hands. So, if you want the method that I used to turn a vague request into a job that AI can actually complete, make sure you go ahead and watch this video here. And

29:16 in the meantime, if you want to apply all of this, learn how to master both of these tools, and really get to the point where you feel you can hand work to them with confidence, make sure you come and

29:29 join us in our Founders community. We have lots of material there. We have a completely detailed Claude Academy course that will teach you everything you need to know

29:39 about mastering Claude. And we also have a free community where you can do lots of challenges. And if you want to be surrounded by thousands of other people who are on the same path as you in terms

29:51 of learning AI, you are more than welcome to join us in both of those communities. In the meantime, thank you so so much for watching. Like this video if you did. Be sure to subscribe if you

30:00 haven't done so. Share this with anyone in your circle of friends or family or coworkers who you think needs to make their mind, pick the tool, and just get to work. Thank you again, and I'll see

30:12 you in this video. Bye.

Frontier News · by Hyperjump Technology