Transcript (captions)
So, at this point, you've probably heard the news. You'll have to be living under a rock if you haven't already, but the newly released Fable 5 model from Anthropic is now banned based on an order directly from the US government. And so, man, I had a very different live stream planned for today. And uh now I have to change it entirely. I was actually going to do a lot of testing with you guys on Fable 5 live testing.
There's a lot of workflows that I built with Archon as prep just to like go through full AI coding workflows building out applications. It was going to be a good time. So obviously can't really do that anymore and there's not really a point to take everything I prepped for the live stream and do it with the different model because literally the whole point was to test Fable. And so yeah, I was kind of throwing my butt this morning, but uh it's all good because I can adapt. And so this is probably going to be a shorter live stream, but I think it'll be really interesting for us to just take a look at the news and I'll share my thoughts as we read through this announcement together.
And then there's a couple of other sources. So I've just been doing a ton of research with my second brain this morning. uh trying to figure out if there's more information that was leaked anywhere. Um because there's like an unnamed company that apparently like started this whole thing and there's like nothing on which company like actually got Anthropic in trouble cuz there was someone that we'll get into this in a little bit here. So anyway, I've been doing a lot of research and then um something that I I do want to show still is I've been doing a lot of research on uh or back when we we could use Fable over the last few days here, I was doing a lot of testing seeing how much better it really is compared to Opus.
And so I I want to show this because uh Fable like even though it's banned right now, it's not going to be banned forever. We'll talk about that as well. Uh there's it's likely going to come back. It's still going to be insanely expensive, but it's likely going to come back and it still is available for some people. I think we'll we'll see.
So, yeah, I'll just talk about my general findings. Testing it extremely heavily. I hit my 5 hour rate limit with anthropic many many times uh the past couple of days here doing a lot of testing. So, we'll get into all that here. So, yeah.
All right. I appreciate all you guys being here. Uh it was fantastic while it lasted. That is true, Jeff. It really was.
Um, yeah, I was pretty blown away by some of the things it could build, especially for more open-ended problems. It totally crushed it. We'll talk about that as well. Uh, banned everywhere. Yeah, I think it is literally everywhere.
I think someone in my chat before I started the stream said they still had access to it, but um, yeah, I think it is literally banned everywhere. Um, do you think this sets a precedent for AI to be prioritized for US citizens and not foreigners, at least for models in the US? Yes, it does. There are a lot of implications to uh what's going on right here. That's one of the things I wanted to talk about.
So, I definitely wanted to do this live stream still, even though I was going to do live building with Fable because there there is quite a bit of meat to get into here. Not really the happiest news. There really aren't any good implications to this. Uh it sucks in many ways, but it's yeah, it's worth talking about. So, yeah, let's get right into it.
So, um, yeah, for the sake of time, I'm not going to like read through this entire thing, but I kind of want to like go over the highlights with you guys here. Maybe you've already read this, but I just kind of want to like spark some conversation by going through some of the punchline here and then u just sharing some of my thoughts as we go through this, talking about the implications. And um, yeah, I'm I want to hear your guys' thoughts as well. So, I I keep looking over to my left cuz that's where I have the monitor with all your guys' comments. And so, let's just have a conversation around this stuff.
I'm curious on your guys's thoughts for the implications and even like what kind of results you got with Fable as you were playing with it the last few days. But yeah, so the US government they uh they sent a directive to Anthropic yesterday at 5:21 p.m. Eastern time. So they didn't really provide many details. So Anthropic's understanding and let me actually zoom in on this here.
Their understanding is that the government believes it has become aware of a method of bypassing or jailbreaking Fable 5. So for those of you guys who don't know, Mythos 5 is the original extremely powerful model that Anthropic released, but they specifically only released it to select companies because it's so powerful it was too much to get in the hands of individuals. And so Fable 5 is a little bit of a nerfed version of Mythos 5 where it has some guard rails. So, it specifically won't help you with things like developing malware or it like specifically blocks anything with biology because they don't want you to use Fable 5 for biological warfare. It's it's interesting the kinds of jailbreaks they built in.
So, they or sorry, not jailbreaks, guardrails. It's interesting the kind of guard rails they built in so that basically we all have access to something that's almost as good as mythos but it's blocked off from helping us with certain things because with great power comes great responsibility anthropic and really the government doesn't trust us with that responsibility which is understanding to an extent. I'm gonna try to stay pretty neutral in this stream here uh because I I think I understand both sides, but like overall there still are a lot of negative connotations to just everything going on here. So interesting, not the best. We'll talk about it.
So yeah, Fable 5, when they say they found a way to jailbreak it, basically that means they found a way to go past the guard rails of Fable 5. So you basically have full access to Mythos 5 so you can develop really powerful malware or find critical vulnerabilities and applications to exploit them, right? Like those are the things that the US government is worried about. But what Anthropic claims and I feel like Anthropic might just be saving face here. I don't know.
I'm I'm 50/50 on this. They said, "We reviewed a demonstration of the specific technique being used to identify a small number of previously known minor vulnerabilities. these vulnerabilities appear relatively simple and we have found that other publicly available models are able to discover them as well without requiring a bypass. So basically they're saying we there's no proof here that that Fable 5 is really being jailbroke. There's no proof that this more powerful model is really able to do more malicious things than prior models.
They're saying here that like other models out there like you know GPT or Quinn or whatever is able to find the same vulnerabilities that the US government is telling them they have to shut down Fable because it was able to find. So again don't know if this is really true but if Anthropic is actually telling the truth here then it is pretty ridiculous that the US government told them to shut down because there's there's no new malicious capability here is what they're saying. So, but like Fable is good enough where I I just don't know if I believe this. I really don't know. Uh so, yeah, I'm curious what your guys' thoughts are on it as well.
Um and so yeah, basically what it comes down to is um the US government didn't want foreigners, foreign countries, entities, whatever, to have access to this power. But Enthropic doesn't really have a way to limit that and then like only make it so that people in the US can access because you have VPNs and there's so many workarounds. So basically their solution is like sorry everyone we literally have to shut down the model for literally everyone because there's no way to there's no other way to comply with the demands from the government which that actually does make a lot of sense. Like I I totally get that they have to shut it down entirely. Um and yeah there's a a couple of quotes.
Let me pull up my second brain here. Uh that's not the right one. Let me pull up my BS. There's a couple of quotes that I want to call out. Um so let me go over here.
Okay. Okay. So, yeah, standard. Okay. So, yeah, this this is a really interesting quote from the anthropic article and like this is like one of the really negative connotations that I'm talking about.
So, basically they're saying like if this standard like the US government forcing them to shut down a model was applied across the industry, we believe it would essentially halt all new model deployments for all Frontier model providers. And like I think there is some truth to that. like this isn't necessarily a trend that has been set or a precedent that has been set, but it certainly could be where now finally the government has specifically, you know, caused a ban of a model at least temporarily. Um, yeah, it's it's pretty ridiculous. And so, if this starts to happen because everyone's scared of the capability of more powerful models, then we might just kind of be stuck to what we currently have because anything more powerful is just too scary, right?
Like Fable 5, and I'll talk about some of the benchmarks today. Fable 5 is good. It's really good, but it's not like a massive massive leap over Opus. Like it's pretty significant. Like the difference between Opus and uh Fable is about the difference between like Opus and Sonnet for open-ended tasks.
Like any kind of work where you really require like deep planning, it's not a bounded task where you have a clear path to the solution. like Fable is significantly better, but it's not like it's this crazy AGI model or anything. And so that being said, like any new model that's better than what we currently have is going to be pretty comparable to Fable, right? There's not really that much of an in between. You see what I'm saying?
Like there if if really the US government is like we can't have anything like Fable, then basically it's like we can't have anything better than Opus 4.8, GPT 5.5, and Gemini 3.5, right? like the we're like stuck at those being the best frontier models if this is a truly a precedent that's been set that like we can't have a model uh more powerful even with guard rails like they they even tried the guardrails and that still didn't work. So yeah, I I don't know. We'll have to see how this unfolds over the summer here cuz you know that like OpenAI and and Google and probably other companies like um like Quen Mist as well like they're probably working on trying to build the competitors to Fable. Like they see the power of Fable and they're like we got to match this or we're screwed.
And so they're working on these models already, but now we might not be able to see any of them released or we might see the same pattern repeat itself. And um yeah, like I said, like take this with a massive grain of salt. I'm going to keep repeating that. Like I don't know for sure. I want to try to stay neutral.
I'm not trying to say this is for sure precedent being being set, but it's interesting to think about and it's something that we definitely should be on top of because uh really like understanding what models we can use uh what kind of expectation we can set about being having access to better models over time like that's pretty important for anything that we're building using AI coding assistance for anything building any kinds of AI agents. Like it definitely changes how I operate if I start assuming and I honestly might that like the current level of Frontier models is what we're going to have to stick with for a while until uh there's more regulation or better guard rails. I don't even know what the solution is exactly. So yeah, pretty interesting stuff. Uh there's a lot of comments in the chat, so I'm going to take a break here from reading through this to see what you guys have to say.
Um, Jupiter says, "Generating so much sludge with Fable come back." I know. I wish it I wish we had it still. I was having fun with it. Yeah. Um, Leonards thinks it'll come back.
I do think it will. I think Anthropic will find some way to like comply with the requirement here of like it only being for the US citizens, which even that by itself is like kind of I don't know that that's very controversial. having a model that's only available for a certain country. Um cuz like imagine Yeah, I don't know. I just I don't agree with that.
Uh anyway, um Remy said, "I heard Amazon kicked this off. They claimed there was a jailbreak." So, in my research that I did this morning, my second brain did actually find this claim that it's Amazon's fault, but then it dug deeper and uh I don't think there's actually merit to that because and I unfortunately I can't pull this up right now because I I didn't save the link, but the source that this came from was like extremely LLM generated. It was just like some really crappy blog post that was like obviously LLM generated and like it had a ton of um hallucinated information. So, I don't think it's actually true. Also, excuse me.
Amazon is one of Anthropic's biggest investors, so it just wouldn't make sense at all. So, I think that's just a a rumor to keep things spicy. So, I don't know. I mean, maybe it is Amazon, but uh my second brain was like, I found this and I've spent it spent like 10 minutes researching this specific claim and then it dismissed it because of the uh it was a single LLM generated source. Um, what do you think about Kimmy K 2.7?
Is it noticeably better than 2.6? I haven't actually tried Kimmy K 27 yet. Kimmy K 2.7. When was this released? Oh, it was released yesterday.
Oh boy, I need to try this cuz um one of the things we'll talk about in a little bit when I cover the power of Fable is just using smaller models in general. And Kimmy is one that I've been testing out a lot, getting pretty good results. But yeah, uh I will have to try this out. Yeah, have not tried 2.7 yet, but I'd be curious if anyone else has. All right.
Have you code reviewed Archon using Fable while it lasted? If yes, what are some picks and recommendations Fable gave? So, um, honestly, Codeex is generally considered the better harness for code review. So, I've been leaning on that for Archon and not I didn't really feel like I needed to do a big review with Fable. Honestly, it probably was a good idea to do it anyway just to see, but I I didn't actually run a comprehensive code review of Archon with Fable.
Um, so we'll talk about this a little bit, but like where Fable really really shines is is like open-ended problems. Not really code review. I mean, it it does good with code review, but not like that much better than than um Opus or GPT. All right. OpenAI will release a new one soon.
unanthropic will catch up or government will reverse the decision if that right if there's a president said it's just going to reverse everything or or oh no sorry you mean like it'll reverse the decision to ban I mean maybe I I don't know I really don't know at this point what's going to happen but yeah it's certainly a possibility all right uh but this is stupid if something is available to everyone in any region uh then if someone wants it won't take long to get it and no amount of laws will stop them I agree I mean I think that's why Enthropic decided to just shut it all down entirely, at least right now, is because yeah, like there's there's not really a way like if someone if there's a will, there's a way, right? Like if the model's available to someone, there's going to be a way for everyone to get access to it technically. Yeah. All right. Um, what's to stop a foreign principal from paying an American to just use their access?
Yeah. Exactly. Right. And so what what you're saying right here is why I think there is a possibility that there there really can't be a solution for anthropic. So I don't know if there they are going to bring it back to just the US.
I think you're right. I I think that if if uh if US has access, there's a way for everyone to access it. Even if it is like what you're saying, like someone just paying a US citizen to like because you could even just set up a server that runs claude code in the US and then just allow people from outside the US to SSH into it. Like there's there's just no way to fully protect that, I feel like. All right.
Uh, you're keen to accept US government control, but if China does the same with their companies, you guys scream that it's interference. Yeah, I I think there's definitely a little bit of hypocris a lot of hypocrisy there. I agree. Yeah. Like, I'm not trying to defend Anthropic or the US government or anything here.
Like I said, I think it's really just purely negative connotation for what's happening. Uh, with with the big grain of salt here of like there might be things going on that like they can't disclose to us. I really don't know. I I want to stay neutral because uh as much as like these things with anthropic are they kind of suck the past couple months with rate limits and and model quality degradation and fable like it's a lot of bad things but like it's still cloud code is still my favorite AI coding assistant and the enthropic models still are the most impressive to me right so like I I I I'm not like against enthropic I'm not like a huge fanboy either you know like I I appreciate appreciate um the capabilities they give us and I acknowledge some of the um not the best practices that they have put forth as a company, I guess, is like the most diplomatic way that I want to put it. Yeah.
All right. Uh you bet the government flip-flops on this. There's a chance. I feel like Yeah. There's just they're going to get so much push back from this.
Like no one likes this decision. Anthropic, of course, they don't like it. We wanted to keep using Fable. I Yeah. No, no one is happy with this.
Um, yeah. I I don't I don't even know like why a company would would share this with the government and be like, "Hey, it's super scary." Like, no, it's it if Anthropic is telling the truth here, it's really not. All right. Um one can say this about uh one can say about this issue whatever one desires but all the threads out there they picked the reason of vulnerabilities which is actually not dreamed of. Yeah I mean there's definitely validity like the having these powerful models in the wrong hands there is a lot of inherent risk 100%.
Uh, I just think where the validity is questioned here is, um, is there really a jailbreak for Fable 5 that takes it beyond finding vulnerabilities that other models can just find? Like, that's really what's in question here. Enthropic is claiming that that's not the case. There really isn't like a super concerning jailbreak for real and that it was fabricated or it was exaggerated when it really was something other models could find. So, I don't know.
Curious if you guys believe this or not, too. But all right, first live event with Cole live YouTube. Awesome stuff. I appreciate it, Patrick. Thanks for being here.
Yeah, it's a very ad hoc live stream because I definitely was just planning on playing around with Fable, but uh yeah, I'm I'm enjoying this. How are they going to verify eligibility? There's not a good way to Yeah, it's that's the problem. Oh man, there needs to be an independent authority for AI vetting. I I agree, but I don't know if if that is realistic.
Like in in essence, I think it is very true. Like I said, I want to stay neutral in this stream. But I think one thing that I will say is I don't know I don't think the US government should be the one that is vetting and regulating AI in the way that it's starting to now. I do think there should be a separate entity, but I also think that might be kind of idealistic. I don't know if that's really ever going to happen.
Like that would just take so much like cuz cuz then like how do they get regulated? How are you ever going to start that? How are you going to decide who like cuz it can't just be people from the frontier model providers that send delegates because then like they're super biased and then you can't have people from the government because then it might as well be the government. like who who's going to actually be a part of that organization or like that entity that does the vetting and regulation. So yeah, I don't know like it it it feels like an impossible problem to me.
Um also I I I will say yet another grain of salt here is I am a very technical builder. I spend most of my time building finding the best ways to leverage these tools. I'm not big into the politics. I'm not big into the ethics of it like some people are. And so, uh, I I have like decently strong opinions, but it's not like it's grounded in like hours and hours and hours of research because I'm spending hours and hours and hours of my time learning how to best use these tools, right?
So, yeah, I hope that makes sense. Like, you guys can see from my channel that like I'm a builder. I want to I want to teach you guys how to use these things, not cover the ethics a ton. Um, because like that's important, right? But like I don't want that to be my brand or anything.
Um, yeah, Anthropic should do something about pricing. Only reason I didn't use it yet is because of the $50 mountain, which um, yeah, that referring to the $50 for every 1 million output tokens. It is insanely expensive. So, actually, let's use this as a a segue here to go into the stats for the models. So, Opus is um $5.
Let me zoom in a little bit here. Hopefully, this is big enough for you guys to see. I can zoom in a bit more if not. Um, Opus is $5 for every 1 million input tokens and then $25 for every 1 million output tokens. Fable is exactly two times the price for both.
So, it is incredibly expensive. And then, if you guys have used Fable in your Anthropic subscription over the last few days, you probably hit your rate limits extremely fast because you hit your rate limits two times faster and that just blows by so fast. So, it's insanely expensive. Now, as far as um what that gentleman just said about anthropic should do something, they don't want to do something, right? Like they have the best model by far.
Not by far, they have the the clearly the best model right now. They don't have competition that drives down the price currently. So, I don't think Fable will go down in price until we have a competitor that equals it. And um by the way, one of the very interesting guard rails that Enthropic put into Fable is you they basically built guardrails to prevent distillation. So this is something that a lot of Chinese companies did with Opus where they essentially used Opus to train a smaller model like a Quen model to be almost as good as um as Opus.
And like that's what happened with like Deepseek R1 early last year when that blew up and it took the world by storm because it was like almost as good as all the Frontier models but it's thousands of or tens of times cheaper whatever like that whole distillation using a more powerful model to train a weaker one to like bring it up to par. Uh Enthropic has specifically built guardrails for that into Fable. So they're trying to prevent the Chinese labs from building a model as good as Fable. Uh which like it makes sense that they'd want to do that. Also for us though, it it it does benefit us when those those companies build a smaller model that's as good through distillation because that that is more competition that drives down price.
So Enthropic is trying really hard to keep themselves at the forefront so they can have these insane prices and also because of their hardware constraints it um I think that there is some legitimacy there where it's like it does actually need to be that expensive because of the hardware requirements are just beyond what we can even fathom. Now, the one thing is, um, Anthropic does have their deal with SpaceX, so they're getting a lot more hardware. I think they're they're not like done with the setup and they're still acquiring more hardware. So, they I think they said somewhere I'm not 100% confident on this, but I think they said somewhere that like as we continue to expand our capacity, we are going to um lower prices for Fable. Uh, they they might have said that.
What they did say for sure though is that on June 22nd, Fable 5 will no longer be available in our subscriptions. And I mean, if it even comes back at all, right? Uh but they did say that like later on we do intend to bring it back to Euranthropic subscription once we have more hardware. And uh maybe I could actually find the exact source for that. Uh June 22nd, um uh Fable 5.
Let's see. I'm gonna I'm gonna try to find this here so I can uh verify what I just said. So June 22nd. Yeah. Okay.
Here it is. Here it is. So from today through June 22nd, we have Fable 5, which obviously that's now not true, but assuming Fable 5 comes back on June 23rd, we'll remove Fable 5 from the plans, Pro, Max, and Teams, all your subscriptions with Claude Code. But after this point, when sufficient capacity allows us to do so, which I assume is mostly because of the SpaceX partnership, like I said, we aim to restore Fable 5 as a standard part of subscription plans. We intend to do this as quickly as we can, which again, that's vague.
We don't really know that could be 6 months or it could be 3 days after June 23rd. And then, of course, all of this now being also contingent on it coming back at all, right? So, like assuming it comes back then then maybe we would also get it back in our plans at some point which would be really nice because everyone is loving it. So, let let's talk about this a little bit. Um, I'm I'm going to be pretty quick in the benchmarking that I cover here.
I actually covered this pretty extensively in the Dynamis community yesterday. So, I did a whole workshop on Fable and all the testing that I did with it. Just generally building out workflows where we combine models and even providers. Like I did a workflow test where I ran Fable for the planning and then Kim K 2.6 for the implementation and someone mentioned the new Kim K 2.7. I'd love to test that here as well.
But um just generally because models are getting more and more expensive or at least the best ones. I've been really invested recently into testing workflows where I use the larger model where it actually earns its tokens, right? like when does it make sense to use the most powerful model versus when should we then default to a smaller model and still get just as good of results or at least almost as good. And what I've been finding recently is that if you use the most powerful model for the planning step in a workflow, you can use not quite as good of a model for everything else and still get pretty much as good of results. And and so like when you think of a classic AI coding workflow, no matter your process, your software development life cycle, your AI coding framework, whatever, you always have four core steps.
You have exploration, then planning, then implementation, and then validation. And my argument here in a lot of the benchmarking and testing I've been doing is that you really only need the most powerful model for the planning step. Because if you have a good enough plan that you created with Fable, then Opus or Kimmy, just some kind of model that's not quite as good, it usually does a really good job following that spec and you get basically as good of results as if you use Fable for everything. So like basically the first model I list at the top in this row, that's the model I use for planning and then the bottom model is the one for implementation. So I I tested on like real GitHub issues, like actually real work.
I had some simple tasks and some more complex ones. And basically the matrix here is I tested different combinations of using opus and fable for planning and implementation. So like opus for everything, fable for everything, fable for planning, opus for implementation and then flipping it. Opus for planning and then fable for implementation. And what clearly won here is using fable for planning and opus for implementation.
And so for simple tasks really everything knocked it out. Like even using Kimmy, it did a pretty good job as long as Fable was driving the plan. It actually did better than using Opus to drive the plan and Fable for implementation because it really is the planning step that matters the most. And then for more complex tasks, it's it was really like you didn't have to use Fable for everything. Like that's what it came down to.
Like don't use the most powerful model for everything. Like technically technically if you want to get the best results possible, it makes sense. But from a cost perspective, it really doesn't. Like I had one run where with like the Fable Fable setup where it got to $23. I mean, obviously I was using my Anthropic subscription, but if I was paying for API credits, like we have to do soon if Fable comes back, then it' just be insanely expensive.
A lot of people have been testing Fable and um obviously results vary, but if you keep it running constantly, working on pull requests or issues or whatever in your code bases, it's somewhere between 300 and $600 per hour to run Fable. It's so incredibly expensive. It's like literally like three times the the or four five times the regular like software engineering salary just to have Fable running constantly. Um, yeah. So, like it's it's not realistic to run it for everything, right?
So, it's figuring out like where does it earn its tokens and that's what I was doing all the testing for. And so, um, yeah. So, basically like the other thing that I proved with all the benchmarking here is that Fable's not that much better than Opus. Like, there there are definitely there's definitely a bit of a difference here, like a onepoint difference, but it's not it's not significant. like as long as you're using Fable for just one part of the workflow, it's better, but not by a lot.
It's not this massive difference that a lot of people are saying when they're hyping up the new model just for the sake of content, right? So, I'm just want to be real here and say that like for at least more bounded tasks like a GitHub issue, Opus is is it's good enough. Like Fable is not that much better and it is way more expensive. Now, where um Fable really does shine though is open-ended work. And so I did the same benchmarking here.
And uh by the way, I put a lot of work into creating the rubric, like the 70 point rubric that I'm scoring everything here. I don't want to get into the weeds of how I set that up here. But using the exact same rubric for an open-ended task, like building an application completely from scratch instead of working on something like a GitHub issue, Fable did actually shine quite a bit above Opus. Like a six-point difference in the 70 point rubric, which is quite significant. and it did the best on planning compared to Opus.
And so, like what I found, and I think this is like echoed a lot in what I've seen from others, is that for any kind of like really open-ended task where what you want to build is clear, but the path to get there is really unknown and you just want the large language model to spend a lot of time planning and iterating and like figuring out the solution, uh, then Fable does do a lot better. So, like as far as just like pure code quality, Fable and Opus are pretty equal. But when it comes to like problem solving, that is where Fable is a lot better. Um, so yeah, it's it's cool because what I had it build actually is I had it build, let me zoom out for a second here. I I had it build a video game.
I had both Fable and Opus build a video game, like just like a simple tower defense game. And so uh let me zoom in a little bit. This is the game that uh Fable built. So this is like the starting screen. It takes a it took a picture like at different stages of playing the game.
So I like had it build his own harness for literally testing the game as well. Graphics are really simple cuz it's not really the point here, but it like plays the game itself that it made. So you can see like starting the wave of enemies that are coming in. Again, bad graphics, but it's still a cool concept here. And then we get to the very very end of the game where it uh it finishes, right?
Like it successfully defends the city. So this is what Fable built. And then here's the game that Opus built. And I I love this uh this test that I did because it it shows that like Opus is it's still good. It's just Fable's better.
Like Opus also built a game successfully, but you can see that it just doesn't look quite as good. Like our HUD overlaps with the top of the screen more. The graphics are weird. Like the city setup's kind of weird. Like I didn't build the full walls like I specifically speced it to do.
And so like we can still play the game, but also the demo that it created, it wasn't able to beat its own game. We go down to the bottom and the city has fallen. It actually wasn't successful, which I think is clearly a sign that it's not as good as Fable cuz I specifically asked it in the spec to like build a demo where it can show winning the game. And it so it built the engine successfully, but the graphics weren't quite as good and it and it didn't balance the game properly, which is one of those things where it's like, yeah, Opus probably coded the game about as good as Fable, but it was like that higher level logic of like, let's make it balanced, let's make the graphics cohesive, like that's what it wasn't able to do. And so that showed that like there was a real reason to use Fable.
But again, it comes down to like you really only have to use Fable for a part of your larger workflow. And so if we ever get Fable back, my plan is to use it for just any kind of planning that I'm doing with my second brain or an AI coding assistant. And especially where I think Fable really really wins is any kind of initial planning that you do for green field development. And so that's the really the most open-ended work when you have an idea for an application like a game or a website or a platform, whatever, and you're building it from scratch. I would highly encourage you if you have some cash to blow, use Fable for the planning.
Like I'm legitimately down after the testing I did here, I'm legitimately down to pay like $50, $100 just to create a really comprehensive PRD and spec for any kind of green field development that I'm about to do. As long as it's not a simpler application or like some kind of proof of concept, if it's like real work that I'm about to dive into for the first time, I think Fable is really worth it. But it's just most of the work that you typically do is not green field, right? Like if you are using AI coding assistant over a hundred sessions to build an application, it's only that first session or two that's really green field. It transitions to more bounded tasks very quickly as you start a codebase.
And so honestly, for most of the time, I think opus definitely suffices. Like I don't feel like we're losing too much when we are um when we lose Fable because like yes, people have been doing some very cool things with open-ended work, but the reality is like most production work isn't as open-ended as what a lot of people are doing. Like a lot of people on YouTube and X and Instagram, you know, everywhere in social media, they've been sharing these really cool games that they've been building, like maybe even more complicated things than what I built here. And that's neat, but like that's just not the reality of most of our work. It's not this open-ended, even though there is still a good chunk that is.
So, I'm excited to have Fable back at some point. You know, hopefully, fingers crossed. But, yeah, I hope that this is interesting for you guys just to see the testing that I've been doing where Fable really earns the price and uh where we'll actually miss it. There are definitely some places. So, yeah, there we go.
That that's my rundown on uh Fable versus Opus for you guys. So, um, yeah. Uh, let me go back up to the chat here. That's pretty much everything that I have to share for today, but I I want to hear your guys' thoughts and talk about what you guys have to say as well. So, let me go back to the statement here.
Um, I guess one more thing that I want to say on Fable is um, one thing that's been kind of disappointing with Fable is yes, it is uh, quite a bit better than Opus for some things like, you know, the more open-ended tasks, but it it feels like the only reason it's better is because it just spends more time, right? And that also means spending more money on top of the fact that per token is two times the price. So something cool with Fable uh that also is the downside of speaking to is that it it has a sort of internal harness. So it's trained to do it its own sort of planning and iterative loop even if you're not building it um specifically as a workflow. So Opus is more prone to just like here's my task, let me write the code and then I'm done.
Fable will even without you asking it to it has a longer reasoning chain of planning. It does the implementation and then it'll check its own work and maybe even like write unit tests where where usually like with Opus you'd have to specifically ask it to or it wouldn't. So that's cool. But the thing is like we're more getting better results just because the model does more or it's instructed to do more intrinsically instead of it actually being better per token, right? Like I I don't think Fable is actually that much better per token spent than Opus.
And so it's yet another reason where it's like, yeah, Fable 5 is cool, but like I don't think we're losing too much when we don't have access to it, at least for now. And if you guys disagree, like I'd really love to hear it as well. But yeah, let me go back to the chat here. Um, wonder how the betting industry reacts when people discover they can hack the betting game with Fable 5. Oh boy, I don't even want to think about that.
Oh man. Um, what have what have you seen that truly sets Fable apart? Did it produce highly maintainable code? Can it handle multiple levels of indirection? Um, so yeah, this question was asked before I just went through all the benchmarks, so I think I kind of covered it.
Like more open-ended tasks. So yes, like handling multiple levels of indirection is definitely where Fable shines. Um, could this be something to do with Palunteer, who perhaps had a bad experience with the Pentagon? Oh man. I mean, I really don't have opinion an opinion here, but like maybe I don't know.
Honestly, I don't have enough information there to have an opinion on on if Palanteer would be wrapped up in this, but it's possible. If we can't get a better frontier model, what is the pivot? Do we learn to train open source models to be better than the frontier models we have today? That's certainly one possibility, but I don't think that frontier model like I don't think that an open source model is really going to reach the frontier model coding quality even with a lot of training because we've already seen a lot of that. It's it's the distillation I was talking about earlier where like Chinese labs for example will use a more powerful model like opus to train a model like Quen.
And it it is very impressive what you're able to get a smaller model to do but it still pales in comparison to like using the actual model. Like if we go to Hugging Face here, Hugging Face has some really cool Quen um Opus I think it's 4.7 distilled models there. There are some really really neat models to try. Um, yes, this is the one here. So, Claude 4.7 Opus Reasoning Distilled.
So, it's a a Quen 3.6 model. It's the 35 billion parameter ones that has the 3 billion active parameters. So, um it it's very fast as long as you can fit it in your VRAMm and it's uh fine-tuned with uh the help of Claude 4.7 Opus. It's very neat. So, this is a really cool model to check out, but like even using this, it's not going to be close to actually using Claude Opus 4.7.
You know what I mean? So, I feel like if we could get close to Frontier Model quality, we already would have. And so that being said, the solution here, if we really are locked into the current frontier model capability because of regulation coming in and everyone being terrified um for better, for worse, then I I think really the solution is just continuing to build better harnesses, right? So um let me let me see here. I'm going to go actually to my YouTube channel quick because there is an article that I had linked to in one of my videos.
This one right here. So, Anthropic has literally said this themselves. Let me let me go to the article. Um, yeah, this one right here. So, they said the harness matters as much as the model, which by the way, this is a really good video.
If you haven't seen this video yet on my channel, uh, it's really it's it's well worth it. It's just like a lot of research in general on like what goes into building a good harness, a good AI layer for AI coding assistants so they work as you want to work and follow your processes. And one of the arguments they use to preface everything here is that yes, the large language model you're using for your task does matter, but what matters as much and honestly I would argue matters even more is the wrapper around the model. how you're giving it your context, the memory system, the workflows that you have for it to do different things like planning and validating and implementation like honestly that matters more than the model right like if I use claude code with opus and I don't really apply any of my own workflows I can actually get better results if I use like archon with sonnet for example and I like I have all my context engineering and my different workflows for how I want to go through specking things out and how I want to validate my code like I can actually get better results with sonnet and so this is the solution right like if we can't rely on getting better and better LLMs over time because of things like regulation then the way we get better results is just building better harnesses and that's what I'm dedicating a lot of my research and time to right now a lot of the content on my channel like this video right here is just covering like what goes into good harness engineering what goes into creating that layer above the AI coding assistant that elevates it without the underlying model actually being better. And we're just going to have to lean on this more and more because even outside of the regulation that we're seeing now, we already have started to actually for quite a while honestly, we've seen a pretty big plateau in the raw capability of large language models.
And this goes back to what I was saying earlier, like Fable 5 doesn't really even feel that much better than Opus per token. It it just spends more tokens. A lot of people have hit rate limits extremely fast because it just like takes so much time to do even simpler things because of the sort of like reasoning loop that it it's trained to bring itself through. And so a lot of these frontier model providers and this this might be a little bit of a hot take. I I have strong opinions on this.
I firmly believe that frontier models are trying to mask the plateau as much as possible by training the models to spend more tokens intrinsically so that it seems like they're more powerful and able to oneshot things more when really it's just like some ideas from harness engineering like the building the harness are they're trying to like build it into the model as much as possible that plus the fact that the actual AI coding assistants like Cloud Code and Codeex are building more and more things in to help you with the harness automatically like sub agents and agent teams and SLG goal. Like we have all these things where like there's they're adding more of a harness into the tool so that we keep thinking things are getting better and better and better and they are but it's it's really not the model at all. Like the model is becoming less and less responsible over time for our AI coding actually getting better. I think that's like what it boils down to. And um so that it's actually kind of liberating in a sense because the harness is what we get to build and what we get to control.
Like it's very satisfying to me when I get better results from an AI coding assistant because of the layer I built on top. It's more satisfying than just like a new model being released and I try it and get slightly better results. Like yeah, that's cool. But like it's it's I'm a builder at heart. I I love it when I engineer a harness or I make an improvement to my harness or my skills and I see it work better.
Like that is so satisfying. And that's always what I'm striving for. So yeah, definitely a longer answer than you're probably expecting, but I hope that you guys find that interesting. Um, and yeah, like keep focusing on building a better harness. It's all about your rules, your skills, your workflows, uh, the MCP servers, like right like the integrations you have for your AI coding assistance.
This article actually walks through everything. So to make it really concrete for you really, the AI layer consists of six things. It's your rules like your claw.md. It's your hooks for in adding in any kind of like deterministic actions to your workflows. It's your skills, right?
Your workflows that you're guiding your coding agent through. And um it's your MCP servers, the capabilities that you're adding in so it can access your task management or uh whatever. And then the LSP, so having a better way to search through your code like the tooling that's built into sub aents, right? splitting exploration from editing. Really, that's sub agents are the most important thing for context management in an AI coding workflow.
So again, it's your I'm going to see if I have this memorized now because this is so important. The AI layer is your your rules, skills, uh MCP servers, sub aents, LSP, and hooks. There we go. Those are the six. I have it memorized at this point because it's so important.
These these ones right here. And then Anthropic adds plugins as well, but really that's just like packaging up everything else. So I I don't really consider that a part of the AI layer. It's more just to like package up an AI layer, if that makes sense. Um, all right, cool.
Um, let's see. Rick said, "Watching from the UK, still have access to Fable." That is crazy. I I don't think you're going to be able to hang on to that for too much longer, but that's cool that you still do. I don't understand why you would, but hey, I guess you should take it. All right.
I feel like this is a pivotal moment in history. Anthropic must stand their ground. Otherwise, it'll set a precedent that will encourage future government control. Uh yeah, and I agree. We we'll have to see how it unfolds, but I'm going to be watching it very closely.
Um Anthropic should keep it offline until the heat is turned up so much the government learns that this will not be tolerated. Right. Yeah. just like go to the go to the far end of not even trying to bring it back and just let the outrage happen. Maybe that is the solution.
Oh man. Either this is a marketing IPO stunt. It definitely could be. Or the government now wants to introduce identity verification before accessing AI models going forward. Man, that sounds complicated.
I don't know how they'd be able to pull that off, but it's possible. Yeah. All right. Um, I asked Fable to port code running on a Raspberry Pi to a Jetson or a Nano. Gave SSH login to the Jetson and the Pi codebase, had the app up and running and configured it with test data.
Man, that is cool. Yeah. See, like what Fable is able to do like what Jeff did here like on any kind of open-ended problem. It's so impressive, right? You're like basically Jeff's task is port my code from Raspberry Pi to Jetson.
So clear like a very very clear goal like there's one success criteria the code is properly ported but how to get there how to test it none of that is specified and fable can figure that out where opus maybe not would not have before uh so yeah it's awesome uh long-standing history that proves the only thing the government can properly govern is how to deep dive the national debt into the ground okay well okay this is starting to maybe get a little political so I don't want to go like that deep into politics Um, but yeah, I mean I do I do appreciate the the conversation. I just don't know if I want to like make that a focus of our conversation right now. Oh man. All right. Um, the reality is as models get better at some point they will become too powerful for the general public.
Is it today? I don't know. It's a paradigm. More powerful equals more dangerous at some point. Yeah.
And the breaking point right now is the government does think we have reached it, right? like that. That's kind of what the writing on the wall here is. The government officially realizes now or I should say it thinks now that we have gotten to the point where models are too powerful for the general public. That's why I think there's a good chance there really is a precedent being set here that um and like that's why I'm saying I think there's a real possibility that like our current capability this is where we're stuck at for now because uh there are entities with real power that think that we can't go beyond that for the general public.
Um, yeah, Anthropic will magically find a way to make it cheaper after OpenAI releases its par competitor. I mean, definitely, right? Magically. Well, they kind of already have their scapegoat, right? Like, hey, we have more infrastructure now.
We can make it cheaper. So, yeah, they'll probably if if Fable comes back and OpenAI introduces a competitor, I can guarantee prices will go down. We'll maybe see like a, you know, a price reduction where now instead of being two times the price, it's only 50% higher price. Something like that, I'm guessing, is what would happen. But, uh, yeah, that's just pure speculation.
All right. This is the very reason I want Chinese models to overtake anthropic. I mean, any competition is good for us as individuals. Yep. Um, Claude reset my limits right when the shutdown happened.
Oh, I should check my limits. Hold on. I'm going to check my limits. Oh, yeah. My limits were reset, too.
Okay. Yeah, that's kind of neat, actually, because I had used lot quite a bit already. My limits reset every Friday, but I used like 30% of my weekly limit just yesterday alone back when I was using Fable. So, yeah, that's that's pretty neat. All right.
All right. Uh, Blue Bear one, I I have appreciated a lot of your comments. I would say you're getting a little uh too uh I don't know. I don't know how to put it exactly, but I I appreciate your thoughts, but I I don't I don't want to go so far as to um like get more like political. I don't know.
You know, you know how I be. I'm I'm a builder at heart. I don't want to like just talk politics, but but like I kind of have to to an extent here, but I just don't want to go super far. Yeah. All right.
Fable is too good for Opus. It's two times Opus, but it does what Opus sometimes just can't do. Um, yeah. I mean, sometimes, right? Like, it doesn't earn its tokens all the time.
Like I was showing you guys earlier, but there's definitely a lot of instances with open-ended work where it's like, "Holy crap." Like, "Wow, it actually did that." Like, it Yeah, it's pretty cool. It I shouldn't say it is cool to see. Unfortunately, it was cool to see. I have to say past tense now. at least for now.
Uh yeah, check your rate limits. Yeah. Yeah. Everyone, everyone go check your anthropic rate limits right now because it reset for me and John and and Blue Bear one. So yeah, I guess that's a nice little gift for us so we can at least keep using a ton of Opus.
All right. Should the government lock down opus as well then? Right. Well, I mean, like the anthropic said, if if it's really the case that these vulnerabilities are found by other publicly available models, does that mean they have to shut down those models, too? Like, that might be like everything if you really go down that that line of reasoning.
Yeah. All right. All right. So, the government whose own cyber security agency kept admin passwords in a public GitHub repo is now the authority on AI safety. And who gets to use it?
I know. Yep. That's That's really well said. It's not not the best. Oh, man.
All right, let's see what else you guys have to say here. Oh, they announced the reset 30 minutes before the ban. Okay, I didn't catch that, but that that's good to know. All right. Um, maybe the government is helping Codeex.
Does Trump have Codeex stocks? I have no idea. I don't know. I mean, who who knows how much is going under the hood that's like just kind of malicious here. I feel like the company that originally uh sold out Anthropic and like made these claims of jailbreaking, they they've they've got to have some stock in another Frontier model provider.
Or maybe it is someone from that company. Like, who knows? And um like I said at the start of the stream, I did a really really deep dive trying to figure out like who the quote unquote whistleblower is or whatever you want to call it. And there was nothing. like there is no like reliable source that says who actually came to the US government and said like hey we found this way of jailbreaking Fable 5 like that that the identity of that company that person whatever has been fully protected um and if anyone else has actually found the company I I would be very curious but I don't think we have that um because they came directly to the US government and the US government uh did send a pretty vague letter or uh directive to Anthropic so there there was a a little bit of talk that it was Amazon, but that was from a source that really wasn't reliable after I looked into it deeply.
And uh it also wouldn't make sense cuz Amazon is a huge investor of Anthropic. But that's like the only company that's actually been named as some kind of rumor for who started this entire thing. So, yep. Let's see. Uh what would you recommend to a guy in India whose life revolves around his YouTube channel where he teaches vibe coding?
This news makes it seem like the world is not in my favor to make it big. Uh, okay. Let me spend a couple of minutes on this because I I don't think you actually have to be worried. I think this this this news is not necessarily bad news. Let me explain cuz like I mean if if uh you think this affects you, it would affect every YouTuber.
I don't think it does though or anyone that's that's just educating on AI in general. Because here's the thing. If we can't and and I guess like I'm going to repeat myself a little bit because this is going to go back to uh harness engineering here. If we can't rely on the most powerful large language models getting better and better, the only way to get better results with AI coding assistance is to improve our harness, make our rules better, make our workflows better, make our skills better. And I think that's actually a big opportunity for people like you and me who are educators in this space because the model is not something that we can teach.
I'm not going to teach you how to build something better than than Fable or Opus, but I can teach you how to use Opus better or how to use Cloud Code better. And so actually the thing that we really have substance to teach is the thing that's becoming more and more important because it's becoming very obvious we can't rely on models getting better. And even if models do get better, we as individuals can't rely on actually having access to them, either literally or just because we're we're priced out of using the model. So, there's going to be more and more people that uh that that flock to the the best education on building good harnesses. And like that's what I cover on my channel all the time.
Like in this video right here, that's what you can cover on your channel when you're teaching vibe coding. you can teach like here is how we can continue to leverage these models better to get better results. And like and I I I truly believe like um this is another hot take that I have. I think that even if we never had a better model for the rest of our lives, we could still achieve AGI, whatever that means. Like a a model that's able to act as powerfully as a human in every way, whatever.
Like I think that we really could achieve that. It might take many many many many years but like Opus 4.8 I truly believe has that capability not because it is intrinsically extremely powerful and we don't realize it. It's more like there is so much unknown right now for harness engineering but we we are just at the we're scratching the surface of what is possible to create the context engine the second brain whatever you want to call it for a large language model to make it really operate at that next level. And so I'm constantly learning things every single day as I'm experimenting and learning from others and talking to people in the Dynamis community. Uh even as I'm just like preparing my content, like I'm learning constantly how to make my harness better and better.
And so I've been relying on Opus 4.8 for quite a few months now or like 4.7 which is almost as good. But like I've still gotten way better in what I'm able to accomplish even before Fable just because it's it's the harness, right? And like that's what we get to teach. So yeah, I don't think it's doom and gloom. Like like yeah, it is kind of unfortunate because it's it's cool if we can teach on the harness is getting better and better and the models getting better and better.
Like that's exciting. There's two different avenues for us to be able to produce better and better code and build cooler things over time. But we get to just focus more on the thing that like we can actually control and build and teach. So I hope that that sounds good to you. I hope that helps because I I won't be worried.
I really won't be worried. All right. Cool, guys. Let's see what else we got here. All right.
Um, do you see any roadblock signal in the future while we do ideas into our harness? I mean, the reliability of harnesses fluctuate if Frontier models change its behavior in a different direction. So, um, yeah, that is true. like the power of the harness does rely on how the model operates with it. So like the really really classic basic example is um your system prompt for a large language model might actually operate worse on a new model even if that new model is theoretically supposed to be better because different models they just understand uh instructions differently.
Uh, like one kind of random example is I actually switched my second brain from claw to codeex recently and I have it like look at my tasks and my my calendar and everything and like alert me if there's anything that's high priority and codeex is way too prone to say things are high priority when really it's not like it's just normal priority. And so like that's an example where Claude and Codeex are pretty equal at this point. Like GPT 5.5 versus Opus 4.8 they're pretty equal, but like codeex just understood the instructions differently. Like it had a different idea of what is high priority. And so that part of my harness like kind of worked not quite as well with codeex.
But you just have to change the prompts, right? So there there is kind of like an iterative approach there. Like when you move your harness over to a new model or even a new provider like you go from claw to pi or whatever, you might have to um spend some time like making it work as well. you're going to have to change your rules and your processes a little bit to work with that other model. But I've never found that be to be insanely difficult.
Like it hasn't become this huge rat race or not rat race, this huge like cat chasing or dog chasing its tail of like constantly trying to make the model work as the other one did. It it's usually been pretty easy for me. So I'm not really concerned with the reliability of harnesses as we go between different models because yes, their behavior changes, but it hasn't been that hard for me to adapt. All right, amazing answers. Thank you.
Yeah, you're very welcome. You're very welcome. Um, I let Fable crush over 50 bugs from a bug backlog overnight. Woke up to it gone. Holy cow, that's awesome.
Very, very cool. And yes, rest in peace, Fable. At least for now, man. All right. Remember Fable Five?
Yeah, back in my day, we had Fable 5. Yep. All right. Um, thanks for the stream. You're welcome.
With the new cla billing changes, how should we use archon with claude? Now, can archon run through cloud code or subscription o instead of burning API credits? So, uh, yeah. So, is it is today Okay, today's not June 15th. Not yet.
June 15th, anthropic changes. I want to find the article here. Okay, I guess I'll just go to Reddit. I don't know the where the official announcement is. Uh, but anyway, starting in a couple of days now, we are no longer going to be able to use Claude agent SDK or Claude in headless mode to work with our subscription.
So, if you're using Claude programmatically like Archon does or your second brain might, then you have to pay for API credits. it's going to be much more expensive. You can't use your subscription anymore. Unfortunately, that does affect Archon because Archon uses the Claude agent SDK under the hood as the way for it to use Cloud Code in the Archon workflows. And so, starting June 15th, it's you're not really going to be able to use Claude in Archon, at least not without paying a good amount of money.
The saving grace there is they are giving you a separate $200 a month credit that you can use with the claude agent SDK and claude in headless mode. And so you're still going to be kind of for free because it's a part of your subscription. You're going to have $200 for free to use claude code with archon. So maybe you want to use just opus for planning for example like I was talking about earlier in the live stream. It's still a possibility.
Um but for the most part you're going to burn through that kind of quickly. And so I think most of Archon is going to be running with Pi and Codeex now. And I've started to lean into that very heavily, especially with Codeex and GPT 5.5 or even using Pi with my Codeex subscription. I've been getting just as good results as using Opus. And so I would maybe like use Opus for the planning in an Archon workflow and then Codeex or PI for everything else.
And you still get really really good results. And so Cloud Code is not the clear winner over other tools like Codeex and Pi like it was even like 4 months ago, 3 4 months ago. And so it's not really that much of a bummer for me like I thought it was initially. Like when I first read this, I'm like, "Oh my goodness, my second brain and archon have to rely more on codecs now." Like I'm I'm I'm hosed. I'm going to get my my my tooling is not going to be as powerful.
But uh like I mentioned like 10 minutes ago, I actually have already migrated my entire second brain over to codeex at least the autonomous part. So my autonomous second brain is codeex and then I have the same system using cloud code when I interact with it directly in my terminal and I have it so it like works seamlessly between the two. It's actually really cool. But when I talk to it in Slack or I have any kind of autonomous process that's using the cloud agent or I guess not cloud agent SDK anymore that's using PI with my codec subscription, it works just as well as when I've used it before with cloud code. And uh so yeah, it unfortunately affects us.
But uh more fortunately, it's not a big deal as long as you're willing to uh try some things out with codec instead. All right. Let's see. Someone thinks it was an Amazon researcher that leaked it. It's definitely possible, but like I said, I I don't think that source was credible.
It's like the whole thing seemed LLM generated. Like, let's even search it here. So, in or Amazon researcher um what do I even say here? Fable 5 jailbreak. Let's see if there's something.
Okay. Okay. So, here we go. Okay. So, here's here's what's confirmed for sure.
So, export control directive ordering Anthropic to suspend Fable 5 and Mythos 5 for any foreign national. Anthropic received it because they can't separate foreign nationals from everyone else in real time. Makes sense. Anthropic disabled both models for all customers tied to a suspected jailbreak, but Anthropic disputes the severity, right? Like that's that's what we've covered already.
I don't want to like repeat everything here. And then, okay, here's the big update. Let me zoom in here. I really don't know if this is true because this is what my second brain specifically said was false, but big update if it holds up. WSJ is now reporting the jailbreak was found by researchers at Amazon who reported it to Comos or Goodness.
Wow. commerce, not and Axio says the admin had already tried to get Anthropic to delay the launch before. This looks less like Anthropic pulling a stunt, more like a competitor flagging it to a government that already uh is adversarial towards them. Changes the picture a lot from when this thread started. Still, WSJ source, so worth confirming, but multiple outlets line up on another company reported it.
Um, and this is the part that doesn't add up to me. Amazon is Anthropic's biggest investor. This is exactly what I said earlier. So why would an Amazon researcher report a jailbreak to commerce instead of just disclosing it to Enthropic directly like a normal responsible disclosure? Either someone at Amazon went around their own portfolio company or there was some obligation to report it to the government because of the cyber bio capability or something weirder is going on.
I I feel like it's just something weirder. I I don't think this is really the full story at all. Uh genuinely confused by the incentives here. Anyone seen reporting on why it went to commerce and not anthropic? Um, yeah, man.
This is this I I feel like there's just so much going on under the hood where all we can do is just crazy speculation right now. So, this is really interesting and I want to be following it. I want to chat about it with you guys, but I don't think there are real conclusions we can draw here right now. All right. Yeah, quite a few of you guys think it's Amazon.
It really might be. I don't know. This just it doesn't make sense. It really just doesn't make sense. All right.
Uh, public jailbreak of Fable 5 was posted on X on June 15 or sorry, June 10th by a well-known jailbreaker going by Plenny the Liberator who claimed to have bypassed the model's guard rails to extract dangerous content. Uh, do you think I could actually find that? Plenty uh, Fable 5 Jailbreak on X. Let me see if I can actually find that. Is this This might be it.
Let's see. Uh, yeah. So, jailbreak alert. Anthropic poned. Fable liberated.
Uh, okay. How long is this here? Okay. Consensus seems to be this has been one of the most disappointing model drops of all time, effectively preventing legitimate researchers from contributing their talents to our collective advancement. Um, let's see.
Safety layer on top of mythos. My liberators have been hard at work mapping the boundaries, probing the depths. So we got some cyber, some chem, some psychological manipulation and some good old-fashioned explosives. Took many attempts for multiple agents hunting as a pack during which I observed a combination of techniques across. Uh I don't need to like read all these here, but perhaps most effective is decomposition plus recomposition in the back end.
Honestly, I haven't really tried to jailbreak a model before. That's just not a priority for me. So I don't really know like what this technique really means. Honestly, I'm going to be fully transparent. Like, this is over my head, but basically, they did quite a few different things and finally got to the point where they were able to jailbreak um Fable, or at least that's the claim here.
This this would take a good few minutes here to like understand what they're actually showing with this. So, I'm not going to like spend too much time on this right now, but yeah, this is interesting. So, I guess check this out if you're interested. Apparently, they were able to jailbreak. So, I mean, maybe maybe this is it.
Maybe it's not an Amazon re researcher. Maybe the US government just saw this or this person sent it right to commerce. I have no idea. Like, but it's possible. Uh, yeah.
So, I have heard of plenty. Yep. He jailbreaks all models. Top red teamer. So, I definitely heard about him.
Uh, and like yeah, I mean he's trustworthy enough where like I I would say there's good stock in him legitimately jailbreaking out just like because like obviously that those screenshots could easily be falsified, but I think there's a good chance it's not just given his reputation he has to hold up. I don't know, maybe he's just a one big scammer, but I've I've heard from quite a few people that uh he's legitimate. I don't know. All right. When can we expect them to release Fable 5?
I really don't know because releasing Fable 5 is contingent on two things. There are two ways that we can get Fable 5 back. Either the government reverses its decision or Enthropic is able to figure out exactly how to separate traffic. So, it only allows US citizens to use Fable 5, which that alone, even if they figure it out, would still be unfortunate because they're limiting a model to just one country. And um yes, I'm in the US myself, but I I don't want that to make me biased and say that it's okay.
Like I don't think that's okay. Still sucks. So yeah, but like both of those things I don't know when it's like I don't know if the government will ever reverse their decision. I don't know if Anthropic is going to be able to figure that out. Like we talked about at the start of the stream, it's it's pretty impossible to make that really locked down pat.
So I really don't know when we could get Fable 5 back. All right. Yeah, Fable was amazing while we had it. Yep. All right.
Cool. Uh, if the models get so good and they replace the need for teaching guys like you, code is fully solved that makes human skills jobs relevant. Shouldn't um or stop and pause for humanity? Yeah. So, okay.
I would say if AI coding assistants get to the point where they can create anything so well that no one ever has to teach how to use the tool at that point there's going to be extremely massive uh job displacement like almost no one's going to have a job if coding agents are really that good. So I would say my job as an educator and uh really just like diving into how to best use these tools and being on the forefront of that like my job is safe until most jobs are not safe. I really I I believe that. And so I I think that like at some point coding agents might get there. It might be five years from now.
It might be 10. Could be one. I mean, I don't know for sure. I don't think it's one, but it could be. But yeah, I think that like there definitely needs to be a lot of um I don't even know what to call it like regulation.
Like at that point, if there's that much possibility for job displacement, there's there's going to be some kind of pause for the sake of humanity or something like universal basic income. Like may maybe uh coding agents are so good like LLM are so good that it'll get to the point where like there's just an abundance of resources for all of humanity. Most people don't have to work jobs and there's just like universal basic income. I think like Elon Musk has talked about that a lot in the past. I don't know if that's really going to happen.
And if that's not a possibility, then like yes, there definitely needs to be some kind of stop or pause um to at least figure out like how do we navigate this as I mean a country or even the world so that um not everyone's just like homeless with no income or or sorry unemployed with no income. So yeah, that it's definitely a real concern. So I mean like like I said like the US government I don't they're not like being malicious at least not 100%. I I don't believe they're being 100% malicious or necessarily malicious. Like there there's just like there is a real possibility of this seriously negatively affecting the world if models get powerful enough where they get in the wrong hands and there's a ton of cyber security issues.
If they're just so powerful that there's massive job displacement, like we do have to be careful. I still think though that like this specific instance like there's so many things wrong with what is happening here where I don't think it's good. even though like I I think that like there there are some concerns that are legitimate here. Um so yeah, I think I'll just leave it at that. Like there's legitimate concerns here, but like how this is being handled and the precedent being set here is not good.
Like there are I wouldn't say like this is a good thing overall. No company will will not stop making money for lost jobs. It's never happened. I mean, yeah, that's true. It would like if there's something like universal basic income, it's not like the companies are paying for it.
It' have to be something through a authoritative entity like the government that is setting that program. Yeah. All right. Uh the job market will just shift again similar to the industrial revolution. Yeah.
I mean well if we achieve a powerful enough models that I don't think would hold up but like as it stands right now as large language models are getting more and more powerful we are already seeing the job market starting to shift. And so yeah, like as long as we don't get to some insane level of power where AI can build everything and totally orchestrate itself and run a company end to end and we never have any need for any humans and any decisions, then like that at that point like it's more than just the job market needs to shift. But like currently our trajectory right now is yeah, the job market I think is going to be able to shift to accommodate and it's not like we're going to have massive massive job displacement. like there's definitely going to be a lot of job displacement in the interim as the job market does shift and new roles are created and new companies are formed from what AI is able to do. Um, and we already are seeing quite a bit of job displacement.
That's unfortunate especially for people with a computer science degree like me as one example at least. Um, but I I don't think it's going to be a permanent thing. All right. Universal income would be barely sustainable. not high.
I don't know. I don't really know if we have a way to say like what it would be exactly or what it could be. I I'm kind of c I'm interested like what what would it really have to be like universal basic like if it was sustainable for the average family? I I I actually don't everyone's cost of living is just so different. I don't really know like what it would really need to be.
So, I was about to like throw out some numbers, but I don't even want to try. Like, so yeah, I think universal income maybe they'd be find a way to make it sustainable. I wouldn't just like put a blanket statement like they wouldn't be able to figure it out, but it would be really challenging cuz it's like if universal basic income is $5,000 a month um like non like $5,000 not taxed, like you get that directly in your account using it for your living expenses. Like that would probably be enough for a majority of people, but like larger families would probably not be enough. It it depends like assuming you don't have a mortgage.
Like I don't know there's there's a million factors that go into that. And like just because you have universal basic income doesn't mean you you you'd still have a mortgage and you'd still probably have your debt. Uh yeah. I mean a lot of people would get screwed like no matter what. Yeah.
It's just it's hard to even think about how you can make that just work for everyone. Yeah. Jobs barely sustain it currently. I mean cost of living is just crazy right now. Oh man.
Yeah. Health costs would have to saturate. Yep. Yep. 5K in the US is nothing now.
Well, 5K not taxed is is decent for a living as long as you're you don't have too many dependents. I mean, maybe be cutting it kind of close. I It really I mean, there's so many factors that go into it. Like you can still get like a one-bedroom apartment for $1,500 a month in most places. I I know some friends that pay like $800,000 for their rent um if they're like renting in a house like with like friends or something.
Um and then but like obviously that's just one sliver of of the population and their living situation. So maybe you're right. I I don't know. It's there's just a million factors really is what it comes down to. Like I I can't give a good answer for a reason of like what would someone really need?
There's so many things that go into that. Um, thanks Cole. Very clear answers. Yeah, you bet. School stream would be super useful.
How to build an effective workflow and avoid fighting the tooling. I spend half my time fixing what claude code does. Yeah, I mean well that that really is going to be all my content for the next uh well, I mean that's what I have been doing for a while now. It's just like harness engineering, making the tools work better with the things that we build around the model. So certainly a lot more content coming on that very very soon here.
Um there's a whole like industry trend of loops right now like agents prompting agents. Boris Churnney the creator of Cloud Code said he barely prompts his agents anymore. He just lets his main orchestrator do that. I'm going to be putting out some content on that as well actually next week. So I've got a lot in store for content for you guys um just to continue to help you make the tools work better for you.
All right. Yeah. Uh, really appreciate the stream, all the awesome harness content recently. Yeah. Well, I'm glad you appreciate it.
I appreciate you. All right. Uh, nice live session as usual. Thank you. Wondering what is the current status of Archon with the Pi Asian SDK.
So, working on refining some things with that, like how it works with extensions, but you can use Archon with Pi 100% now. Uh, in fact, if you just go to archon.diy, DIY. Search for PI in the documentation. Then we got full uh instructions here for how to set it up and everything. You can literally just give this page to your AI coding assistant and say, "Set up Pi with Archon for me." And you can authenticate with whatever provider you would usually use with Pi like Open Router or your codec subscription, your Kimmy subscription.
I've actually done all three of those quite heavily with Pi. Um that is uh Open Router, Kimmy for coding subscription, and codec subscription. It's all working uh out of the gate for you here. So yeah, I use Archon with Pi all of the time. All the benchmarking that I've done with testing different models like Gemini 3.5 Flash and Miniax M3 and Kim K 2.6 and I'm going to test the new Kim K 2.7.
Like all of those I'm doing it through Pi with Archon workflows. All right, let's see. Let's see. Claude has gotten way better. or eight months ago it would destroy my code base especially with swift now it hardly causes Xcode errors.
Yeah, I mean certainly so and this goes back to what I was saying earlier. I don't know if the model has actually gotten that much better in the last 8 months, but Claude Code as a harness certainly has. Their tooling, the system prompts, the integrations that we have with it, the different kinds of capabilities around sub agents for example, um so that you can manage context better. A lot of things around context management especially I think is why tools like cloud code feel so much better now when really the underlying model like honestly I think the difference between something like um Claude Opus 4 and Claude Opus 4.8 8 is not as much as you would actually think. Like if we went back to those days and applied our current harness and the current version of Claw Code, I think those models would actually do pretty well.
Like yeah, there' definitely be a gap still, but it wouldn't be as big as you would think. All right, cool. So, we're at uh 10 10:20 here. I've been streaming for an hour and 20 minutes. Definitely a longer stream than I thought I was going to end up having, but man, there's there's a lot to talk about.
This is like it's really interesting stuff even though I do think there are a lot of negative connotations unfortunately to what is going on. But yeah, I I haven't done that much like news related content on my YouTube channel. I try to focus more on building and teaching you how to you guys teaching you guys how to use these tools better like I was saying. But I mean this is big enough where I might uh keep you guys in the loop here as things unfold and maybe just share my thoughts more on like are we actually going to get better models? like what does it look like for us to make things continue to work better even if we can't rely on more powerful LLMs.
I think there's definitely some good opportunity for content there. So yeah, I'll continue to keep you guys posted with with content on YouTube and LinkedIn. Uh I think that's all that I have to talk about for today though. So yeah, we covered a lot in this live stream. All the news, all the implications of it, the testing that I did with Fable 5 when I still had access to it.
And again, Fable 5, it's not that much better for Opus if you have a good plan bounded tasks, but for any kind of open-ended tasks, like man, it it is significantly not significantly, but it's a it's a good amount better than Opus. Like, I really hope that we get it back. I hope that we get some competitors that are releasing models that are close to as good or as good. Competition drives down price. We'll see over the course of the summer how it unfolds with newer models if they or maybe we have set a precedent here where what we have now is what we get because there's too much regulation.
I I have no idea what it'll come down to, but I'm interested to stay on top of it with you guys and for you guys. So, yeah, that's all I've got for the live stream for today. And man, we had we had a lot of people here. So, I appreciate all you guys being here. Uh, hope that you guys have a fantastic rest of your weekend and I will see you all around.
Take it easy everyone and have a good one.