Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
Chinese open models like Kimiko 3 and Qwen 3.8 Max are closing the gap with closed frontier models, threatening the business models of labs like OpenAI and Anthropic as they approach IPOs. The video argues that open models create abundance and deflationary pressure, which could decelerate innovation, but a balance between open and closed models is necessary to prevent price gouging while maintaining progress.
Key points
- The gap between open and closed models has narrowed from years to months, with Kimiko 3 catching up to Anthropic's Fable 5 within a month.
- Open models pose a threat to closed labs' IPO valuations and revenue, especially from API and enterprise customers.
- Chinese labs face compute shortages due to chip regulations but still produce competitive open models like Qwen 1.5.
- OpenAI's head of strategic futures argues that open models encourage deceleration by reducing incentives for investment.
- Anthropic's Dario Amodei warned about the dangerous path of open source, pushing for regulation.
- A healthy balance of open and closed models is needed to keep costs manageable and encourage innovation.
- Mem Zero provides a dedicated memory layer for AI agents, offloading conversational memory efficiently.
Tools mentioned
Techniques
- Model architecture and training techniques
- RAG for documents
- Conversational memory with Mem Zero
- Memory offloading to context windows
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
The AI race between open and closed models has been going on for more than 6 years now, which predates the release of ChatGPT in 2022. Open models are of course what everyone likes to see since it allows people to download and run the model at home or a dedicated server. And ever since closed labs more or less branched off to privatize more capable models to offer them through API at cost, the pace of innovation started to diverge where since 2022 closed labs dominated when it comes to frontier capabilities and consumer distribution. And this very gap between open and closed models was initially about a year and a half behind to later only about a year and then later to a few months and now with the release of models like Kimiko 3 and Qwen 3.8 Max, the gap is nearly as competitive as any other frontier models. This kind of velocity in innovation led leaders in the industry calling them dangerous or even saying that open models are inherently decelerationists, which is kind of a crazy take and we'll get to exactly what that means later in the video.
Now, one of the first signals that we all saw open models potentially catching up to closed models was none other than the deep seek moment back in January 2025. Deep seek caught up to OpenAI's first reasoning model called 01 at a fraction of a cost and this unexpected event led to nearly $1 trillion sell-off in tech stocks in the stock market. But still, the time delay between OpenAI's 01 and deep seek R1 was still about 4 to 5 months, which is a comfortable lead. Except Kimiko 3 not only caught up in benchmarks, it only took them about a month since Anthropic first released their Fable 5 model, which was supposed to be a huge leap in frontier capabilities. Now, the timing is nothing short of important here because these closed labs like OpenAI and Anthropic have been eyeing for an IPO.
So, when you think about the timing here where you want all investors to believe that their own models are durable and defensible from open models that are made publicly free. And open models are making this very difficult. Open AI's IPO is targeting a $1 trillion valuation in 2027, while Anthropic recently surpassed Open AI's latest valuation at $965 billion. Anthropic is also targeting an IPO as early as this fall. And while valuation isn't just in the models capabilities alone, when we look at the entire ecosystem from agents and applications to models to infrastructure, access to chip and energy, open models from Qwen 1.5 catching up does target exactly at the model layer at a time where IPO is just around the corner.
Recently, we've seen more regulation from the US on closed models, which creates an additional pressure for the closed labs to create even more capable models that are distinguishable from open models, while also clearing the government in time to make sure that you retain your customers that directly impacts the bottom line in the revenue. A large portion of Open AI's revenue comes from subscription, whereas Anthropic gets from API pricing and enterprises. And as long as open models continue to gain traction, it poses a significant threat in user retention and revenue more immediately for API and business and enterprise revenue streams, since early adopters tend to dominate around here and they have the means to switch platforms for lower pricing. And Anthropic in response to GPT 5.6 and Qwen has been refreshing their token allowance back-to-back to incentivize their users to continue using Fable without having to shop elsewhere for intelligence. China, on the other hand, has the opposite problem, where as we all know, Chinese labs have compute shortage.
And due to a huge spike in their newly released Qwen 1.5 model, Qwen ended up restricting new subscriptions because their infrastructure simply couldn't handle the growing demand from users. Chinese labs like Qwen have shown once again that open models shouldn't be forgotten about and be remembered as a strong alternative. And speaking of remembering things, your agent also needs a strong memory, which is why we need to talk about Mem Zero sponsoring this video. When I built my custom agent, I found that one of the trickiest parts was memory. How can I get my agent to remember things about me and my conversations?
One example is an agent that keeps track of my food habits. Rag works well for documents, but it's overkill when what you really need is user-specific conversational memory. Mem Zero is a dedicated memory layer for your AI agent. This is a Python script showing you how by importing the Mem Zero library, my agent can start remembering things about my dietary needs and suggest dinner recipes around my preferences. And even my friend Justice, who has celiac disease, remember to avoid foods containing gluten.
And this memory persists even if I restart the agent or come back days later, it still knows my preferences. By offloading memory management to Mem Zero, I can use context windows more efficiently and focus on building my custom agent faster and making it more personalized without spending so much time building a memory layer from scratch. And when you're ready to go beyond a personal project, Mem Zero scales to multi-user production-grade memory without changing your code. Try Mem Zero for free in the link below. It works with your existing stack in minutes.
Even with the chip regulation that prevents advanced chips to be sold to China, a large majority of models in China are still trained using US chips that are already three to five years behind in technology even with the recent continuation of Nvidia H200 chips being sold to China. But despite all these disadvantages, Chinese open models still remain a huge threat to the US to a point where the head of strategic futures in OpenAI, Dean, who wrote this rather long post as a response to Kimi's new model Kimi Key 3. I highly recommend reading through this post since it touches on a broader philosophy around how the AI industry should be thought of, but Dean's main point here is that open models encourage deceleration by virtue. And if you're not sure what a decelerationist is, here's a quick explanation of what Dean means. If more and more frontier capable models start to become easily accessible, where people can easily download them and modify them freely, that must mean that there's now an abundance of AI models.
And having abundance has a deflationary effects in the ecosystem because it incentivizes people away from using frontier models from companies like OpenAI, Anthropic, SpaceX AI, and Google that tend to charge them at a higher premium. I mean, why would users pay for more expensive model when you have an abundance of models as alternative to choose from? So, because of that, companies will have lower expected profits, which reduces the incentive for investors to pour huge capital to into hiring top talent, building large data centers, buying high-quality GPUs. And from this point onward, everything spirals down to deceleration of AI innovation since there's no big capital flowing into a highly commoditized market. Dean's post also criticizes accelerationists since so much of AI innovation was born from commercial push.
So, wanting open models at the cost of AI acceleration could end up slowing down innovation without a commercial push. At the same time, an old video of Dario Amodei testimony has been recirculating, originally from 2023, when he said that open source is heading down a dangerous path as he talked about the potential misuse of highly capable open models. And Anthropic has since been pushing for more regulation from the government for AI innovation. And while the sentiment out there right now definitely seems to dominate for open models to dominate, which certainly is pro-consumer, training frontier models do require a lot of capital that's needed to gather trillions of high-quality tokens, building data centers with high-quality GPUs, hiring top talent in AI for model architecture and training techniques, and all these are currently spearheaded by Frontier Labs. So, we do need a healthy balance of both open and closed models to encourage closed labs from price gouging end users and be pressured to keep innovating and open labs to continue help increasing AI adoption while keeping the cost manageable.
And one company that sort of stands unique in this is Google, who doesn't really have to stretch and bend over backwards to compete. Google has custom ASIC like their TPUs, top researchers, and the infrastructure and strong ecosystem to release their products. And Google doesn't really stand to lose as much as other Frontier Labs if the model layer truly ends up becoming a commoditization at the end. Since their moat is not in a single model that they release, but their control over the entire stack that actually benefits from commoditization rather than being threatened by them.