OpenAI Astra explained..

summarized

TLDR

OpenAI's new model, Astra, reportedly solved 10 open problems in mathematics and theoretical computer science, producing a 26-page lean-certified proof for high-dimensional sphere packing. This challenges criticisms that LLMs are just stochastic parrots and highlights a shift where humans guide and verify AI-generated work, with OpenAI spending about $2,000 in compute to solve these problems. The video discusses implications for AGI, noting that while impressive, these achievements are within well-defined domains and still require human involvement in problem selection and verification.

Key points

  • OpenAI's Astra model solved 10 open problems in mathematics and theoretical computer science, including high-dimensional sphere packing, with a 26-page lean-certified proof.
  • The video contrasts Yann LeCun's criticism that LLMs lack physical world understanding with Astra's novel mathematical breakthroughs, questioning when the stochastic parrot argument ends.
  • Astra's achievements demonstrate a shift in human roles from doing work to guiding and verifying AI-generated solutions, similar to how developers now review code written by coding agents.
  • The high-dimensional sphere packing problem asks for the maximum density of identical spheres in a given space, and proving this universally across all dimensions is the hard part.
  • OpenAI reported spending about $2,000 in compute to solve all 10 open problems, though this likely doesn't include failed attempts or alternative paths.
  • The video argues that while Astra's math and CS achievements are impressive, they represent only one domain of intelligence, and AGI requires more nuanced capabilities beyond well-defined problems.
  • The cost of intelligence is dropping, allowing AI to tackle problems that previously required scarce human expertise, shifting the bottleneck to verification and guidance.
  • Websites like unsolvedmath.com list thousands of open problems that AI could target, indicating a ripe area for future AI-driven discoveries.

Tools mentioned

Techniques

  • Lean certification for formal proof verification
  • High-dimensional sphere packing mathematical proof
  • AI-assisted problem solving with human guidance and verification
Transcript (captions)
Ever since the idea of AGI became popular, one of the biggest contention was whether LLMs could meaningfully help us get there. LLMs are still what most of us use today to chat, make workflows, and develop software with. And LMS have been incredibly useful over the years. And now, OpenAI's newest model, Cename Astra, made huge advancements in math and theoretical computer science, just showing you how powerful LM really are. But LLM aren't without its criticisms, either. One of the main criticisms against LLMs is a French American computer scientist, Yen Lakun. And Yen has been long stating that LLMs just lack the ability to understand the physical world since they don't embody the world through experience. And he claimed that by 2030, nobody in their right mind would use LM anymore, at least in how they're currently configured. I mean, if you think about it, LM couldn't even count how many times the letter R appears in the word strawberry. So, how can we trust LM to do anything more critical? And we're now about 3 and 1/2 years from 2030. And yet Open's newest model Astra, which is rumored to be GBT6, just solved 10 open problems around mathematics that have been left open-ended and unresolved for a long time. And with the help of OpenAS team, Astra generated this giant manuscript showing its proof. Pretty amazing, right? Now, we typically rely on highly intelligent people or people who dedicate their entire career to specific fields to solve highly [snorts] abstract problems like this. And now AI is being used as a tool to explore solutions that we just didn't have the resources to dedicate to before. And this gives us few important things to think about. First is that LLMs truly aren't as limited as some people make out to be. LMS have largely been labeled as a stoastic parrot, meaning they don't inherently have the capacity to understand something, but rather LLM just stitch together words to make it sound plausible. But when does that argument come to an end? especially when LLMs like Astra are making what seems like truly novel breakthroughs that require deep understanding about a given topic. The second thing worth noting is the gradual shift in our role in exploration and solving complex problems where humans guide and contextualize the work that needs to be done and AI is actually the one doing the work and we just have to verify the proof afterwards. And if this looks familiar to you, this shift has already been happening largely in coding for some time with the help of coding agents. Most developers now don't actually write code but review the code that's written by coding agents like codeex and clot code and our job is to verify the work by reading through the code written by the coding agents hopefully. Let's look at the 10 open problems that OpenAI Astra model just solved. Basically all of the problems here are so abstract that only 0.1% of the world might have the ability to understand the actual proof that the model generated. Take a simple theorem like Pythagoreium theorem. This theorem that's already thousands of years old helps us find the relationship between three sides of the right triangle where in the simplest form the total area of the squares for two sides equal to the total area to the square of the hypotenuse. And based on this simple theorem alone, we built many complicated applications that assume this as the foundational truth. But how can we exactly be sure that this theorem works for all types of right triangles? We could certainly draw every kind of right triangle that's possible and prove that this theorem works in all cases, but there are literally infinite number of right triangles that we can draw. So instead, we need a universal proof. In fact, there are hundreds of different types of proofs that are out there, rearranging them to show geometrically or even algebraically. And this book that contains hundreds of different proofs. As you can see, even a simple idea like Pythagorean theorem can be supported by complicated proofs and a manuscript to ensure that the argument is logically sound, applies universally, and can be independently verified. Now, looking at the first problem that Astra just solved called highdimensional sphere packing, similar to Pythagorean theorem, the problem is simple, at least in theory. In a given space, what is the maximum density of identical spheres that can be arranged without overlapping? Another way to think about this is this. What is the theoretical minimum amount of empty space in a given space? In a 2D space, we can have a flat coin arranged in different formats. But out of all different arrangements, what is the mathematical proof that ensures that we have the maximum packing? And also, what about 3D with spheres or even higher dimensions like 40 or even 100 dimensions like the dimensional vectors that we are used to seeing in machine learning. As you might have guessed, the hard part isn't really in finding an arrangement that might work, but finding a mathematical proof that works universally across all dimensions to prove that there's no other conceivable arrangement that can reach this density. And Astra wrote this 26-page long proof that's lean certified. And the output here is what OpenAI team formalized into a manuscript. So surely this means that we've reached AGI, right? Can we at least just say that LLMs are so close to reaching AGI? But first, a quick word from Zo sponsoring this video. It's pretty obvious that more and more parts of our lives are being integrated with AI and often our interactions are scattered across claude, chat, Codex, and Menace. So, how can we have more ownership while having a 24/7 agent? Zo gives you a dedicated computer on the cloud that's yours, meaning your agent is on standby for anything that you give. And part of being 24/7 is the ability to message your agent through text. You can directly message Zo to have normal conversations through iMessage or ask about files that you have stored on your computer in the cloud and better yet build an e-commerce website for vintage watches and host directly on Zo using their AI agent natively through text. What's cool about Zo is that you can viode and launch sites that are actually useful because you can integrate all your context into Zo and run automations. Custom domains are also included at a paid plan. And all your sites are hosted on your Zoe's cloud computer. Plus, they got some neat tricks like the selector tool to edit exact areas to improve things or even add automations that plug into your website like a text anytime someone fills out a form. Try Zo today. I'll have the link in the description below. To answer the question of whether we've reached AGR or not, we have to take a few steps back. Subjects like math, computer science, and physics are only one domain of intelligence. These subjects often have clear boundaries and sometimes a known solution to work towards which tends to breed for AI to thrive. We can't really say the same for other domains beyond these subjects since other branches of intelligence are more nuanced. And as much as solving these 10 open problems is really impressive by its own merits. There still was a lot of human involvement in selecting the right problem, contextualizing them so the solution matches the problem and evaluating the entire process. But it's hard not to ignore the implications that models like Astra can help us achieve. Going back to the distribution of human intelligence, we rely on highly intelligent people and people who dedicate a large portion of their lives to solve complex problems. In open problems like the ones that OpenAI Astra solved, because human intelligence is scarce, we can only dedicate so much of our resources to the right set of problems. But now with artificial intelligence, specifically like LLMs, like Astra, the intelligence is rising and the cost of intelligence is dropping to a point where OpenAI reported to spend about $2,000 in soul's pricing to solve all 10 open problems. Really blows open the type of problems that we now have access to solve by delegating complex work to AI. And for us, just to verify the work that would have taken us huge amount of resources allocation for humans to solve. And now our bottlenecks is shifting into verification of work and guidance. Much like the shift that we have already been seeing when it comes to coding and software development. And our universe is full of more open-ended questions. Even websites like unsolved math.com that shows thousands of unsolved problems across various domains is ripe for AI to target them. One caveat to keep in mind here is that while $2,000 that OpenAI spent to solve 10 complex problems sounds impressive, they didn't separately report the costs involved in finding solutions that may have not ended up working or even problems that it attempted that it didn't lead to a solution. So $2,000 likely assumes a single path towards finding a solution. But considering how much money we have to dedicate alternatively to allocate human resources, this is definitely small in comparison. So now we come a full circle back to the topic of AGI. A lot of tech CEOs have been touting AGI for a long time now as if they're right around the corner. Even someone that I admire a lot, Deis Habis, who's the CEO of DeepMind, has been saying that we could achieve AGI by 2030 with about 50% certainty. And Demis calls our current time as a precious moment before AGI to pause and think philosophically about what a world with post scarcity would look like. And since the term AGI is so hotly debated and hated by many people, maybe the term AGI matters less than the actual impact that AI is already having in our economy of intelligence where our role in how work is done is changing across many domains.

Frontier News · by Hyperjump Technology