AGI is Here. Anthropic Just Proved It.

summarized

TLDR

Artificial general intelligence (AGI) is already here, as demonstrated by Anthropic's AI model Claude, which can solve open-ended problems and perform tasks without clear specifications. Claude's success rate on such problems has increased to 76%, and it can handle tasks that take up to 16 hours to complete. This development has significant implications for society and the future of work.

Key points

  • Anthropic's AI model Claude can solve open-ended problems and perform tasks without clear specifications
  • Claude's success rate on open-ended problems has increased to 76%
  • Claude can handle tasks that take up to 16 hours to complete
  • Anthropic's report shows that their AI model can perform tasks that were previously thought to be exclusive to humans
  • The development of AGI has significant implications for society and the future of work
  • Anthropic's report highlights the need for caution and careful consideration of the potential risks and benefits of AGI
  • The company's AI model is capable of deciding what comes next and making decisions that are smarter than those made by human researchers

Tools mentioned

Techniques

  • Open-ended problem-solving
  • Task automation
  • AI model training
Transcript (captions)
So, as of last month, more than 80% of the code that Anthropic ships is now written by their own AI, Claude. That stat comes from a report that they just dropped called When AI builds itself, where they basically pull back the curtain on what's actually going on inside of their own business. So, I read this whole thing multiple times, and I walked away pretty convinced that AGI is not this far off future thing that [snorts] we're all just kind of waiting for. I think that by the definition that actually matters, AGI is already here. So, today, I want to tell you guys the important stuff that I took away from this article and like what I think it means for society. So, real quick, just to start off, let's get on the same page about what AGI actually means. So, AGI stands for artificial general intelligence. And I think that everyone gets stuck arguing about that middle word, general. Like, can it just literally do anything that a human can do? Can it feel things? Is it conscious? All that kind of stuff. And just so we're clear, I'm not talking like sci-fi robots taking over the world sort of AGI. I'm talking more practical here. I understand that some of you guys will disagree with this take, but I think what matters is, can I walk up to a model with a problem that has no clear answer and just say, "Hey, go figure this out for me." And then it actually goes off, it runs experiments, it does research, it tries a bunch of different approaches, and it comes back with something that actually works. And I think that that distinction is very, very important because you've got narrow AI, which is something that's really, really good at one specific thing that you actually built it for. So, a model that plays chess or a recommendation algorithm or a model that sorts your inbox. Even if it nails one job at 99% of the time, that's not AGI because it's still stuck in one lane. In my mind, AGI is when you hand a general model almost anything, a problem that it's never seen before that has no clear answer, and it does the work on its own to figure out the the approach. It researches it, it experiments, it finds the best way. So, a model scoring 80% on simple narrow tasks is cool, it's helpful, but that's not AGI to me. The thing that I actually care about is the open-ended stuff, the problem-solving. And this whole report is Anthropic showing you with their own internal data that we are already there. All right, let me show you the actual proof. So, to be clear, Anthropic in this article never says the words AGI is here. That's me saying that. But, once you see their own numbers, I think that they make the strongest case that I've seen that the practical version has already showed up. Because most of the time that AI has existed, it's been great at little things, right? Like, your quick questions, your Google search replacement, write me an email, summarize this article. But, the second that you try to hand it maybe like a real big open-ended project, something that doesn't have a clear known answer, it kind of falls apart. So, Anthropic splits these kind of like coding sessions into these four buckets based on how hard they are. You've got trivial tasks, which are the easiest, and then you get more difficult going down from routine tasks, substantial tasks, and then of course the hardest ones are the open-ended problems. And open-ended, in their exact words, means tasks with no clear specification, where the engineer isn't sure what the answer should look like. So, nobody even knows what the finished thing should look like, and you just hand that mess to an AI model and say, "Go figure it out." So, imagine you just told someone, "Hey, go make our app faster, but not which part, not how, not even what faster should look like when it's done." They basically have to figure out the entire thing on their own. And that is an open-ended problem. So, on these open-ended problems, Claude's success rate just hit 76% and 6 months ago that number was 26%. So, it jumped 50 points in half of a year. And this is that top rung, the messy, no map, nobody knows the answer rung that just got cracked. And that rung is this whole ballgame. And it's not just answering harder questions, it's doing longer and longer work all on its own with nobody babysitting it. Anthropic actually tracks how long a task takes for their AI to handle it from start to finish. 2 years ago, their best AI model could handle about a 4-minute task. A year ago, an hour and a half-long task. This year, 12-hour tasks. And one of their newer internal models worked for 16 hours straight. I'm assuming that was Claude Mythos. That length has been roughly doubling every 4 months. They said if this trend holds, tasks that take a skilled person days could come into range this year. And in 2027, AI systems could be capable of tasks that take a person weeks. Their typical engineer is now shipping eight times as much code per day as they were back in 2024. Now, of course, more lines of code doesn't automatically mean that it's better code. I completely get that. But it still tells you how fast everything is actually accelerating on this exponential curve. And they've clearly been shipping features insanely fast. It kept me really busy this March. But it's not just doing the work you handed anymore. It's starting to decide what comes next. So Anthropic ran, obviously, another test. They take a real research project, freeze it right at a decision point, and basically ask the AI, "Okay, what do you want to do from here?" Then they'd line up the AI's answer against what their own human researchers actually picked. They did this across 129 of these moments, and back in November, the AI made the smarter call 51% of the time, smarter than the human. And by April, it was up to 64%. Now, to be fair, they picked spots where the humans first move had maybe some room to improve. But still, more than half the time the machine is choosing a better next step than the people who literally do this thing for a living. And the actual work it's doing is getting pretty absurd as well. They handed a chunk of training code to get sped up. A year ago, their model would have made it about three times faster. But this past April, this newer model made that same code block 52 times faster. And on one problem the researchers had been stuck on, they just let the AI agents grind on it around the clock, and it clawed back 97% of the gap. The humans, given about a week on the same thing, only hit 23%. And once again, keep in mind, a lot of the these numbers are probably coming from Mythos, which is Anthropic's newest and even smarter model that is not yet publicly released. But just imagine what this will look like when that's in everybody's hands. So the machine isn't catching up to the people who build it, it kind of already did. All right, so in the report, Anthropic lays out three ways this could go from here. And I think this is the clearest way to understand where we actually are right now. So, scenario one, the trend just kind of stalls. All these crazy lines on the graph that are going exponentially just kind of start to flatten out and plateau, and AI ends up being a super incredibly powerful tool still, but it just kind of plateaus out like that. Scenario two is that these gains that we're seeing keep compounding. The AI does more and more of the work, but their exact words are, "Humans continue to set research directions and judge results." And then scenario three is the big one, the AI becomes fully capable of building its own successor. And at that point, the speed of progress isn't really limited by humans at all anymore. It's only limited by how much computing power you actually give it. Now, think back to that second scenario for a sec. The AI does the work, humans just point it at the problem. And that's not really even the future, that's literally where we are right now. That's where we are today. And that to me is already AGI. The only real question left is whether we slide into scenario three. And that is the scary part. Anthropic basically admits nobody can tell which one we're actually on, because the exact thing that makes all of this so impressive is the exact same thing that should make you honestly a little bit nervous. Cuz if you think about it, if the AI is the one building the next better AI, then any little flaw, any weird behavior in today's model gets baked into the next one that it builds. And then that one builds the next one. And Anthropic says that pretty straight, which is "The rare occurrences of misalignment present in today's model could compound as the models grow their successors, growing more frequent but less understood until we lose control of them." So, what that actually means is the mistakes don't just add up, they actually multiply. And they get harder to even see at the same time. So, you end up with a problem that's growing and it's going invisible all at the same time. And that's the thing that keeps people up at night, because it it's got nothing to do with killer robots, it's that quietly we stop or not we, the actual professionals that know how this stuff works like under the hood, they quietly stop understanding what's being built. All right, so let's just zoom back out for a sec. Zoom out from the labs and let's look at the rest of us. You know, right now society's reaction to AI is basically split into two groups. I mean, obviously it's a spectrum, but we've kind of got two ends of that spectrum, right? You've got people on one side who open up an AI tool at work, they type in one lazy prompt, they get a result that's very meh, and then they say, "This AI stuff is so overhyped. You know, I don't get it." Copilot. And then you've got this whole other group over here quietly running what used to take a 15-person team, but they're doing that by themselves because they have little AI agents that go off and build agents and automate things. Anthropic even says in this report, "100-person companies could start doing the work of 10,000 or even 100,000-person organizations." And that gap between those two groups of people is just getting wider every single month. And society has not yet come even close to catching up. And the thing that ties this whole video together is that the gap between people is the exact same problem as the race between the companies just at a different scale. No solo person who figured this out is going to sit around and wait for everybody else to catch up. And no AI lab is going to hit pause while their competitors keep sprinting because whoever stops ultimately loses. Anthropic admits that part pretty flat out. The incentive to keep going, in their words, is enormous. So everybody from one person at their laptop all the way up to the billion-dollar labs and entire countries have the same exact incentive, which is to not stop, to not slow down this AI progress because everyone wants to win the AI race. So that brings up the obvious question, why is Anthropic, the company building the most aggressive version of this whole thing, why is why are they the one who's standing up and waving the flag? And when you actually read what they say, it kind of makes sense. They say basically the one thing that they admit they're least sure about is alignment, which basically just means keeping these systems pointing in a direction that's actually good for humanity, good for society. And that's their words. The alignment problem is the thing that they are least certain about. So slowing down would buy everybody more time to figure that out before it's too late. And they straight up say that a slowdown would likely be a good thing. But then they hit you with the catch, which is obviously Anthropic's only going to pause if every other major lab also pauses, and only if everyone can actually verify that the other labs have paused. And that is like almost impossible because how would you verify that? In their words, training runs are far easier to conceal than missile silos. So what that means is you can see a country building a giant missile, but you cannot see a company quietly training a model in some random data center 100 ft underground. There's nothing to point a satellite at, right? And we've actually solved a version of this before, just not with AI. Back in the Cold War, the US and the Soviet Union signed nuclear treaties where they literally let each other's inspectors walk into bunkers and confirm missiles were gone. That's the famous trust but verify. Anthropic even name-checks one of those exact treaties, but their gut punch is the fact that those treaties took decades to build up, and the trust in the systems that, you know, it took to pull off. Like I said, it was it took a long time. And what they said is, "We don't have that long." So, you know, getting a little existential here, but if you ask me, their promise to slow down is sincere and convenient. That's the way I feel at least. I mean, obviously they want good PR, but I think they genuinely mean it because if you think about how the super small group of people have so much power over the future right now, it's it's pretty scary. But they also know the one thing that would actually force them to do it doesn't yet exist and might not for a while. So, after all that, I actually want to bring it back down to earth because the report does too, and it's kind of my favorite part. Basically, like if the doing part, if the actual building, the typing, the grunt work is basically becoming free, then the thing that's actually valuable just shifts, you know? It becomes your judgment, your taste, you know, your expertise knowing which problem is even worth pointing the AI at in the first place. So Anthropic says this themselves once again, that the one thing humans are still better at is seeing the bigger picture and thinking beyond the confines of the immediate task at hand. So, that's the muscle that you want to be building right now because the real danger here isn't the doomsday stuff, at least I really don't think so. I think that it's using this like a search box while the person next to you figures out how to use it like an entire team. So, is AGI here? I think by the definition that actually matters to me, the one where you hand a general AI model a hard problem and it just goes and solves it. Yeah, I think that that's AGI and I think it's here. It showed up pretty quietly. None of us really got to vote on the definition of AGI and the company that built it is the one telling us to be careful with it. So, the worst thing we could do right now is look away and pretend it's still science fiction because it's not. So, anyways, that's what I wanted to talk about today. If you guys are really interested in this kind of stuff and you want to keep up with all of this type of discussion, then definitely check out my free school community. The link for that is down in the description. You can come ask questions, hang out. We've got about 400,000 people who are building with AI every single day, building businesses, doing all kinds of cool stuff. But anyways, that's going to do it for this one. So, if you guys enjoyed the video, you learned something new, please give it a like. Helps me out a ton. And as always, I appreciate you guys making it to the end of the video and I will see you all in the next one.

Frontier News · by Hyperjump Technology