New Models
2026-08-17
Zhipu did not build a new model. It kept training the old one — and says coding got 50 percent better.
Zhipu released GLM-5.3 on 14 August. The unusual part is what it is not: the base model is unchanged from GLM-5.2 — the same roughly 744-billion-parameter mixture-of-experts design with about 40 billion parameters active per token. Every claimed gain comes from post-training: a longer training run, tens of times more long-context task environments, and more varied kinds of them. Zhipu says coding ability improved about 50 percent over GLM-5.2, ranks first among open-weight models on Terminal Bench 3.0 and Agents' Last Exam, and scores 84.5 percent on the CyberGym security benchmark. On the company's own Code Bench it puts GLM-5.3 at 31.4 percent against Claude Opus 4.8 at 29.5 percent, using about 50,000 output tokens where Opus used about 120,000. Most of these numbers are Zhipu's own evaluations. The weights are not out yet: the company says it will publish them roughly two weeks after release, once security hardening is finished.
Why it mattersFor two years the recipe for a better model has been a bigger, freshly trained base. This is a claim that you can leave the base alone and get a large jump from what you train it on afterwards — which is far cheaper and is something a small lab could copy. Treat the size of the jump with caution until the weights are public: almost every number here was produced and reported by the company that sells the model, and nobody outside it has checked them yet.
✓ Verified · 3 sources
Read in the app — free, in 9 languages
Related stories
Rumour: the anonymous model that just topped a coding benchmark, for free, is said to be Zhipu's unreleased flagship
2026-08-21DeepSeek's cheap workhorse can now see — and on agent tasks that need eyes it says it is close to Anthropic's best
2026-08-21Nvidia is paying $6 billion for the machine that builds a rival's models — and hiring 109 of the people who ran it
2026-08-21Given four hours and a GPU to improve the way AI is trained, the best agent scored 0.25 out of 1 — and most never tried
2026-08-21Show the robot once — three to twelve seconds — and it gets the job right 59 times out of 100 with no training at all
2026-08-21