GLM 5.2 - The Top NEW Open Weights Model

summarized

TLDR

GLM 5.2 is an open-weights model from Z.AI that achieves top-tier benchmark scores, rivaling proprietary models like Anthropic's Opus 4.8 and OpenAI's GPT 5.5, while being significantly cheaper to use. The model excels at long-horizon tasks, agentic coding, and front-end design, and includes multi-token prediction for improved speed. Weights are publicly available, enabling users to choose their own service provider.

Key points

  • GLM 5.2's open weights were released by Z.AI for both full and FP8 versions, allowing users to avoid sending data to Chinese servers.
  • The model achieves top-tier benchmark scores, beating DeepSeek Pro, Qwen 3.7 Max, and MiniMax M3, and rivaling Anthropic Opus 4.8 and OpenAI models.
  • GLM 5.2 shows a huge improvement over GLM 5.1, especially in agentic coding benchmarks like Terminal Bench and Deep Sui.
  • The model uses a long chain of thought, outputting more tokens than competitors like DeepSeek and Qwen, which boosts intelligence but increases token usage.
  • GLM 5.2 ranked number one in the Design Arena for front-end tasks, generating complex web pages with animations from simple prompts.
  • Pricing is $1.40 per million input tokens and $4.40 per million output tokens, significantly cheaper than proprietary models.
  • The model performed well in tests like the Pelican on a bike SVG test and long essay writing, producing more than 5,000 tokens on a single request.
  • Multi-token prediction contributes to faster token generation, especially for design tasks, averaging 36-40 tokens per second via OpenRouter.

Tools mentioned

Techniques

  • multi-token prediction
  • long chain of thought reasoning
  • post-training improvements for long-horizon tasks
  • agentic coding fine-tuning
Transcript (captions)
Okay, so I wasn't planning to make a video about this model. I'd seen that GLM 5.2 was coming and I kind of felt I've covered some of their stuff in the past. But I got to say that this model is very impressive. I've been playing around with it this afternoon and the thing that just tipped me over the edge is artificial analysis have just published their stats on it as well. So one of the reasons I've been reluctant to make videos for a number of the Chinese models recently is that they've been slow to actually release the weights. So they've had sort of proprietary versions and it's been not sure whether they were actually going to have the weights. Luckily to see things like MiniMax 3 actually did release the weights. Some of the Quen models we've seen weights be released but there definitely seems to be a change going on there. But I was very pleased to see that in the last 24 hours Z.AI have released the GLM 5.2 weights. They've released both the full version and the FP8 version. It generally seems that people are not releasing the base models for a lot of these things. And I got to say that kind of makes sense. We saw for example with the Cursor team basically fine-tuning the Kimmy model, they actually had a paid agreement to be able to get access to that base model and to be able to do the fine-tuning through Fireworks etc. And I got to say that that I think is totally fair. That this is one of the ways that these companies try to make a profit in here. So just jumping in, basically 5.2, a lot more improvements over how they're doing the post training on this. They're basically saying that this is built for long horizon tasks. And you can see that in a lot of the benchmarks that they've put out themselves, they're really only being beaten by Anthropic's Opus 4.8 and Fable, which is now no longer available to most people. We can see on some of the benchmarks they're definitely being beaten by OpenAI as well. Terminal bench is more of an agentic kind of benchmark there. And the one that probably is most interesting is this deep sui, right? This is a new benchmark that people have been talking about a lot to replace sui bench pro, to replace some of the ones that have been going on there. They're also showing that this not only is a lot more competitive with the Anthropic models and the OpenAI models, but this is a substantial bump from where they were for agentic coding than with the 5.1 model. So, the 5.1 model I thought was a a good model. So, already just seeing them compare themselves to that model is really kind of interesting. They've also got on the whole bandwagon of multi-token prediction. And I guess the cool thing there is that the model does seem a bit faster than some of their previous models. Now, they've got a full benchmark table in here. It's good to see them showing basically when they're being beaten and by sort of how much. Some of these things, it does seem probably the size of the model is is hindering them for things like Humanity's Last Exam without tools. You can see that Opus is doing quite a lot better there. But with tools they're catching up, and that's something that I feel is a good sign here. The ones that really pushed me over the edge into making a video though were the artificial analysis ones. So, I like their benchmarks a lot. how they check out models. They seem to be quite thorough in testing different things around the models. And we can see here that in their benchmark they're showing that GLM 5.1 the jump between that and GLM 5.2 is pretty huge here, right? So, this is a sort of mixture of a bunch of different benchmarks in here. The only ones that are beating them out are GPT 5.5, Opus 4.8, and obviously the unavailable Fable 5 with fallback though. There's been some really interesting things that before the model got taken down was that if you actually just benchmark Fable without any fallback to 4.8, it actually does pretty badly because it has so many rejections for things that are often just total sort of nonsense rejections. And so if in cases like that, the model itself ends up getting a lot of things wrong. So you have to basically have it so if it does get rejected, it falls back to 4.8 and then relies on 4.8, which is already the second top model in there. It's also interesting looking at a number of the artificial analysis benchmarks, just not only how much of a jump between 5.1 to 5.2, uh but just how this is beating out things like the DeepSeek Pro, like Qwen 3.7 Max, like MiniMax M3 that just came out last week. And of course, beating out some of the proprietary models like GPT 5.5. We're expecting at this stage the three sort of frontier labs being Anthropic, OpenAI, and Google. We're now starting to see some of these open models actually come in there. And it's going to be very interesting that the pressure is certainly on Google for the Gemini 3.5 Pro model and we'll have to see what's next there. Now one of the things that's certainly interesting in here is just how many tokens is this actually using. And if we look here, this kind of explains perhaps why the model is doing so well. It really likes long chains of thought here. You can see it's only being beaten by a few models that are actually outputting longer chains of thought. It looks like that this is basically putting out more than DeepSeek, Qwen K and even Fable in here. So, if you see here, the ideal sort of situation is this green box where you've got basically high levels of intelligence, but not a huge amount of tokens out. And we can see that the 5.2 max is definitely on the side of having a lot more tokens out than most of the models there. Interesting one of the things that I heard last month in San Francisco talking to people was just how much Open AI has been working since GPT 5.1 to keep high levels of intelligence, but reduce the number of tokens. And it does seem that this is the direction that everybody now is heading in is that they've gone for this really long token usage, but now they're pulling it back to try and get high levels of intelligence with short amounts of tokens in here. All right, the last one that I thought was really kind of interesting, too, was the design arena. So, they've got actually put this model at the front above sort of the clawed models in there. Certainly does seem really impressive in that this model can do front end, can do those kind of tasks, and be able to do them pretty well. Okay, so, I've been testing the model on Open Router today, and I'm actually using the ZAI version here. I could imagine over time you'll see Together AI, you'll see other companies basically start serving this. And it may make sense that if you would prefer not to send your data to China, or not send your data to certain data centers, etc., that you use one of them. The cool thing is because the weights are open, we actually have the ability to choose that. Now, everyone seems to be charging the same price, which is $1.40 per million tokens in and $4.40 per million tokens out. So, that's hugely cheaper than the proprietary models are charging at the moment. So, even if this does actually use more tokens to get its result, and of course you have to test this yourself, you're probably going to find that it's going to be cheaper for a lot of the use cases that you might want to use it for. And I do think for this kind of thing, I'm starting to perhaps rethink for where I was paying for monthly plans for some of the Chinese models to basically keep testing them and using them with my team to now just falling back and paying for the tokens that we actually use. Okay, so this is the tool that I use for basically testing a lot of models. We can do both local and served models in here. So, if I basically ask it for something like the Pelican test, this is kind of an interesting one that has been used for a lot of different things. You'll see that I'm actually getting pretty good end-to-end token speed coming out from this. And this basically serving from Z AI but through the open router here. You can see that we've got thinking tokens. If we look at the thinking tokens, you'll see that you get pretty good standard sort of reasoning thinking out. In this case, it's not very long. If we go to preview the SVG, we can see that hey, it's done a pretty good job at doing the Pelican on a bike test in here. Also, if we ask it to do things like long essay writing and stuff like that, it does seem to do quite nice reasoning before it gets to the actual article. But I'm finding that the reasoning is not overly long. So, I think you really need to test it for what you're going to be trying to do in here. So, this one has been able to go through. It does seem like the token speed changes from time to time. I've probably been averaging around 36 to 40 tokens per second using the OpenRouter API. But, the cool thing here is that this has been able to go through and actually write a long article. So, when I ask for 5,000 words, often I'm getting at best maybe sort of with some models only 500 words or only a very small amount of tokens back. Now, I don't expect necessarily that it's going to hit the 5,000 words, but it is interesting that this is going on quite a long way compared to other models out there. Okay, getting to the end, you can see that the word count for tokens did actually get above 5,000 tokens there. So, it certainly done a lot better than a lot of the other models for that particular kind of task. Giving it tasks like logic puzzles and stuff like that, it seems to be doing reasonably well. I do like that the reasoning tokens or thinking tokens seems to go up a lot when it actually makes sense for it to go up. So, that's a good sign for me. Okay, and then the other thing is you have the design, which is what this is supposed to be very good at is basically ranking number one for this. You can see here we've got it designing a homepage for the Dario Wellness Retreat with in the Tuscan hills, perhaps even with a horse named Calypso. And we can see that, okay, it's flying through this kind of thing, being able to generate out tokens quite easily. And my guess is that this is where their multi-token prediction is coming in really beneficial, that it's just able to go through and get this quite nicely. Okay, so I had to end up rerunning it because the first time it went over the 8,000 tokens that I set as max tokens, but the funny thing is giving it the second time of 32K, it actually finished just nicely over 8K. So, you can see here is Dario's Wellness Retreat. It's done a pretty nice job, right? It's got a anthropic look to it, which I guess kind of makes sense. It's gotten images. It's basically been able to pull all of that. It's even got different sorts of animations for things coming in and out, which is pretty amazing from this simple prompt up here that hasn't really got a lot of stuff going on in there. So, it's been able to do that. Unfortunately, I don't think it's actually got all the different pages and stuff yet. This is just basically just sort of render out and see what it actually sort of made out. But overall, I got to say that this model is definitely a model you want to check out. This could be an alternative to a lot of the proprietary models that you're using. I really don't see the point in using some of the models, perhaps like a Sonnet, perhaps like a Gemini Flash, if this thing can replace it and do it for a much cheaper cost. And like I said, at the moment, you've got a limited number of service providers here. My guess is that over the next week, we're going to see this go up quite a lot. A lot of people are going to be serving this model. And of course, as always, if you are using Open Router, come in here and see what is actually going on with your prompts and tokens that you're sending in. Are they actually retaining them and perhaps using them for training? Some of the companies are retaining them, but not using for the them for training. But that's something that you definitely want to be looking into. So anyway, the video is probably a bit longer than I expected it to be. I just wanted to say that this looks like a really cool model and I really think it's definitely worth you trying out for your particular use case. Anyway, as always, if you like the video, please click like and subscribe and I will talk to you in the next video. Bye for now.

Frontier News · by Hyperjump Technology