Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
Saoud Rizwan, founder of Cline, argues that the open source community is dying due to AI-generated spam and security risks, but open weights models will become the industry standard, driving down costs and potentially displacing closed models. He urges American labs to release open weights models to maintain leadership, citing examples like Coinbase adopting GLM and the historical precedent of open compute.
Key points
- Open source communities are withering due to AI-generated PRs, issues, and security risks, leading projects like Zig to ban AI use and TL Draw to close all PRs.
- Supply chain attacks like the LightLLM compromise highlight the increased danger of depending on third-party software.
- AI labs are subsidizing subscriptions to lock in developers, but this strategy may fail as businesses switch to cheaper open weights models.
- Open weights models like GLM can match closed models in output quality at lower cost, as shown by Cline's internal tests.
- Coinbase defaulted to GLM and Kimi, cutting AI spend by nearly half, signaling industry adoption of open weights.
- The open compute project shows that standardizing on open designs commoditizes components and reduces costs for everyone.
- Inference costs are projected to drop 90% by 2030, making open weights models even more attractive.
- American labs should release open weights models to maintain competitive lead and influence over AI development.
- Cline launched an open weights subscription plan to offer discounted access to models like GLM and DeepSeek.
Tools mentioned
Techniques
- Open weights model adoption
- Inference cost optimization
- Volume-based discounts
- AI-native development infrastructure
- Model routing for cost efficiency
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
[music] Hi, I am SA um founder of Klein. Uh I started Klein as an open source project a few years ago. Um, some of you might know it as the first ever coding agent back before um the cloud Mac subscription and the codec subscriptions when people had to pay for each and every API request which uh got extremely expensive. Um this was before prompt caching became a thing and so there were people that paid hundreds of dollars a day uh using client but for a lot of people it was their first AGI moment. It was the first time they saw um LLMs be able to do their jobs end to end.
Um and they got hooked. Um and I don't think Klein would have been as successful as it as it is if it wasn't open source. Um because it allowed these developers to inspect our code and trust it and uh connect to any API so they could be comfortable with spending so much money on it. Um and uh know that they weren't getting screwed over. Um, and we were the first to add things like custom rules and and plan mode.
And a lot of that came from talking to and learning from this really incredible open source community we had around the project. Um, and so, you know, having spent most of my life building open source, it's really heartbreaking to see just like the broader open source community uh wither and and die over the last two years uh because of how AI has fundamentally changed everything about software development. Um, GitHub is effectively uh an archive of SLOP PRs and issues and security reports um where the sense of community before has turned into this like deep skepticism and distrust um of each other's responsible use of these tools. Um because AI coding can be extremely dangerous to a project and everyone's kind of had to learn that on the fly, but especially open source projects uh that rely on trusting third parties. Um, and so I wanted to share some examples of how open source has been dealing with AI.
So this is the code of conduct for Zigg, which is the language that powers bun. Um, and they essentially ban all use of AI. You can't use it on pull requests or issues or even comments. Um, and the reason for this is that to them, the core Zigg team, um, they value contributors more than they do the contributions. Um, and so the primary goal for reviewing PRs and things isn't to add new code, but it's to help grow new contributors who can become trusted um, over time.
And AI assistance completely breaks that. Um, this is an uh a post from the CEO of Kurl who says that his project is effectively being dodoed by AI generated bug reports and they're even considering um shutting down their bug bounty program for the first time in decades. Um, and this is TL Draw. Uh, they're automatically just closing all poll requests whether they're AI generated or not. Um, and it's gone so bad that GitHub added a feature to disable thirdparty pull requests altogether.
Um, which is which is really sad because pull requests were the thing that made GitHub what it is today. And we're probably going to see a lot of big open source projects um, opt into this. Um, and so when I say that open source is dead, I mean some parts of it like the community um, it's just not worth cultivating anymore, especially because building software is so cheap. Um, and uh, also the risk of supply chain attacks. I'm sure you've seen all the all the reports of of things getting compromised.
It's become more dangerous than ever to depend on third party software uh, where it takes a single compromise and a massive chain of contributors to get pawned. Um, so just as an example, light LLM um, is a Python package. It gets like three and a half million downloads a day. They were compromised for three hours where attackers um used a GitHub app that they used um to steal their Pippi publishing tokens and um publish a compromised version of the package that would install a credential harvester that would steal your API keys, your SSH keys, your crypto keys, um and also install a backdoor that lets them do remote command execution. And the only reason this was even caught as quickly as it was was just pure luck um because the malware had a bug in it where it would cause cursor to crash if you ran the light LLM MCP server.
And a security researcher noticed that um and was able to figure it out. But if this had been out any longer, it would have caused like catastrophic damage, especially because a lot of the people using Light LLM are like the enterprise customers um and developers um that have their own internal gateways. Um but despite all of this, I believe there are some parts of open source that are sticking around and becoming more important. um like allowing others to use your thing and uh freely in the public domain and build on top of and those parts about it are going to become more important than ever um particularly with open weights models because of the economic impact and so to help explain why I want to look at what's happening with inference spend right now. So, this is a report from uh an anonymous report from a CFO at an unnamed company where they accidentally spent $500 million on Claude in a single month because they didn't set the usage limits on their thousands of employees on their anthropic dashboard.
Um this is another report by Uber CTO where after they rolled Cloud out to their organization, uh 95% of their engineers were using it. 75 70 70 70% of their committed code came from claude and their monthly spend per user was up to $2,000 and they said they used their entire 2026 budget um in just four months and the crazy part is is that the AI labs are losing money too. Um this is a chart from semi analysis where they ran experiments with claude code and codec subscriptions where they would give them long u horizon coding tasks until they exhausted their weekly limits. And they found that a $200 plan for claude would give them about $8,000 worth of API usage. Um and a $200 subscription to codeex would give them about $14,000 worth of API usage.
Um so I think the strategy is pretty obvious. They're essentially going to subsidize this um until they have as many engineers dependent on their uh tooling as possible with agents in their CI and background cloud agents and looping agents and all these things where it feels like every new feature and marketing push from these labs seem to be a new workflow to standardize on to use even more tokens and to be locked in even more. Um, and then inevitably the price gouging, uh, once they've got you trapped where your developers can't work without the tools. And this isn't theoretical. We're seeing this happen live with some of the customers that we talk to.
Um, so just a quick show of hands, how many of you basically stop working whenever there's a cloud outage or a GPT outage? Yeah, same. [laughter] Um, I think that's a reason why we've seen Anthropic and OpenAI go from being API businesses to investing so much into the application layer. Um, is because they know that that's where they can set these sorts of traps and build their moat for the day that these models inevitably become a commodity. But I don't actually think the strategy is going to work.
Um, and that's the message I wanted to get across today. Uh, that this feels very shortsighted. And what we're noticing happen um in the world is that it doesn't matter how many features your CLI agent has, uh developers and businesses will just jump to whatever offers them the best value for their dollars. So if we look at current open weights models, many of which are are built in China, we'll notice that although they've lagged behind the American closed source competitors, we're at an inflection point where raw intelligence lead doesn't matter as much anymore. Um because these models are powerful enough where you don't always need the best one for all your work.
And that cost is becoming extremely important to these businesses that have kind of turned a blind eye until now. And I think we all kind of feel it that to get the best output from these models, it's more a problem of what context and tools you give the agent access to and less about its raw intelligence. with the right AI native development infrastructure with project skills and rules um systems of verification and quality gates um even a mediocre model can produce similar results as a more intelligent model it just might take more tokens um the intelligence is better placed in the system and guard rails around the model so that you don't have to be as reliant on the model or your end developer's responsible use of the model itself. So we recently shared uh an anecdotal experience where we were skeptical of the benchmark saying that GLM was better than Opus. So we tested them on a real bug from the client repo and while both models fixed the issue, GLM was the winner in terms of cost and code quality.
So GLM used twice as many tokens but only cost half as much. Opus finished faster. It used half as many tool calls. But GLM cleaned up dead code and verified that the build compiled before completing while Opus didn't. It left a bunch of type errors and it broke the production build.
Um, and so that gave us the sense that GLM was trained to spend more tokens verifying its output. Um, which is fine because the tokens are cheaper any anyways and it's really the end result that matters. And because these open weights models can deliver the same output, a little bit more tokens, we're seeing signs of the industry adopting and standardizing on these models. So this is Brian Armstrong, the CEO of of Coinbase, saying that they've uh defaulted to using GLM and Kimmy in their internal LLM gateway and that this has cut their AI spend by nearly half uh while their token usage continues to grow. And I think we'll see other businesses building their own internal tooling and routing to work with these agents in the most dollar efficient way for them.
Even if it means not having access to the latest new feature in something like cloud code. We're seeing the same thing that happened with open compute 15 years ago. So uh just a quick history lesson for those that haven't heard of this. In 2011, Facebook was just getting started on building out um stuff like distributed computing infrastructure and data centers. Uh but by the time they built it, Amazon and Google already beat them to the punch.
So it wasn't a competitive advantage. So Mark said, "All right, let's just open source it and see what happens." Uh so they took the designs for their data centers and their servers and their networking and cooling racks and all the physical hardware they spent all this energy and and money um building and they just gave it away. They published uh schematics and CAD files and everything uh and called it the open compute project. And what they saw was that the entire supply chain reorganized around it. So before open compute, every company designed its own proprietary servers.
So manufacturers were doing small production runs um of like very custom hardware which was expensive. But when Facebook's designs became this sort of shared open standard, suddenly everyone was ordering the same thing and manufacturers could do these massive standardized production runs commoditizing these components. Um so no single vendor could charge a premium and the price of everything came down uh for the whole industry including Facebook itself. Um and so what Facebook found was by giving these designs away they created the market that drove their own costs down and saved them billions of dollars down the road. And so I think the lesson um taught here is that the industry will adopt and standardize on something that they can build on top of even if it isn't the best thing.
And with how much capex we've locked in for the next 5 years for AI infrastructure buildout, open weights models are only going to get cheaper. Um there's estimates that we'll spend nearly three trillion dollars and create uh over a 100 gawatts of new data center capacity by 2030. um roughly doubling uh global capacity today. And hosting providers like base 10 and fireworks, their whole purpose is to beat each other u to beat the competition. So they'll use infrastructure efficiency gains like dedicated hardware and caching and batching volume tricks and inference specialized silicon to drive cost down even more.
And um by 2030, the estimates are that inference on a 1 trillion parameter LLM will cost um 90% less than it does today. we're seeing uh the same sort of cost cutting tricks that commoditized the cloud 10 years ago um happening in inference. Uh so in 2014 uh Google uh at their GCP live March announcement said that they would cut compute by 32% and storage by 68%. Um and AWS fired back with similar cuts within within days and this was AWS's 42nd uh price cut at that point. And so from 2015 onwards once raw compute and storage were a commodity the hyperscalers stopped competing on on it as much and started competing on other things like databases and serverless.
And I think because open weights allows these host providers to compete and cost uh optimize so aggressively, we'll see mass adoption of these foreign open weights models um because when dollars are involved, the markets are extremely efficient and the absurd API costs that these closed labs charge just won't be worth it anymore for most knowledge work. Uh and so this is me just like humbly requesting the American labs to take open weights more seriously. Uh because before we know it, all this infrastructure that we're investing in could be built on foreign models that take the world by a storm and make GPT and cla irrelevant. um mind share and adoption is incredibly important and there's a chance that if the foreign models become the standard there won't be a reason to switch back to GPT or claude or Gemini no matter what the marginal improvements are and then we lose control over the development of this technology. Um and who knows where the world is is headed if we don't have the likes of anthropic and open AI to invest so heavily into safety research in ways that perhaps these other labs wouldn't.
I think the development of a technology this transformative is deeply uh tied to the ideals of the people and the nation that's building it and to instill those values in the future of this technology we need to keep the lead. So I don't mean we need to open source our research. I think that's what gives us the lead. I think [snorts] we need to open we need to open up our models and uh start releasing more open weights models which as we know are not nearly as useful. you can use um and extract the traces and and train more you know your copycat models on them more easily but not in a way that can leaprog and this would make models uh more usable by the industry in a way that allows more competition and adoption and better price and value for customers and I think that's how we keep our lead during this very critical moment me and client believe so much in this open weights future that we launched uh an openw weightight subscription plan earlier this week that through volume uh based discounts and partnerships uh with inference host providers we can offer significant discounts compared to paying for these models um at direct API cost um and we plan on continuing to increase the usage quota for these models as they become cheaper uh you can sign up at that link client.bot/pass bot/pass.
Uh if you'd like to get a feel for how far models like GLM and DeepSeek have come and how you don't need the most expensive closed frontier model access to get work done anymore. Uh client is also uh open source and you can bring an API key and use any other provider. Um you can use it on your CLI and VS code and Jet Brains. Um and uh yeah, we continue to um add the newest models whenever they're released. So it's a good way to get a feel of like how much better the latest newest model is especially with open weights because um you can't access those with like the cloud or chat GPT subscriptions.
Uh cool that is my presentation. Thank you. [applause] >> [music]