aiminute. ← All AI news
New Models AI Minute Newsroom 2026-08-27

The stealth model developers spent weeks probing is now MIT-licensed and downloadable — and Z.ai says the entire preview ran on Chinese chips

The stealth model developers spent weeks probing is now MIT-licensed and downloadable — and Z.ai says the entire preview ran on Chinese chips

Z.ai released GLM-5.3-Flash on 26 August and confirmed it is the model that had been circulating anonymously as Ox Alpha. It is a mixture-of-experts model with 320 billion total parameters and 18 billion active per token, a one-million-token context window, weights on Hugging Face under the MIT licence, and — a first for the GLM-5 line — native multimodality, meaning image and video are handled by the base model rather than bolted on. The architecture pairs linear attention for local dependencies with sparse attention for long-range context, and at full context uses a component called IndexPool to compress indexer key vectors so latency and memory do not blow up. Z.ai reports training on a roughly 30-trillion-token multimodal corpus. On its own numbers the model scores 63.4 on DeepSWE v1.1 against GLM-5.2's 46.2, 48.8 on AutomationBench against 26.2, and 84.3 on Terminal-Bench 2.1 next to Claude Opus 4.8's 85.0. Pricing is $0.15 per million input tokens, $0.03 cached and $0.50 output — roughly a tenth of GLM-5.2 — with a 50 percent promotion running to 9 September. The claim that will be argued over hardest: Z.ai says the whole Ox Alpha preview was served on domestically produced Chinese accelerators using its own serving engine, reporting a threefold end-to-end improvement across tens of thousands of them. All benchmark figures are self-reported and have not been independently replicated.

Why it mattersTwo separate things happened here and it is worth keeping them apart. The first is a pricing event: a 320B model that answers coding and agent benchmarks within a point or two of a closed frontier model, released under MIT — the most permissive licence there is, meaning you can ship it in a commercial product without asking anyone — at roughly a tenth of what its predecessor cost. That collapses the gap between what you rent and what you can run yourself, and it does so in exactly the categories companies pay for. The second is the hardware claim, and it deserves more scepticism than the benchmarks. Export controls rest on the premise that advanced compute is the chokepoint. If a lab can serve a frontier-adjacent model to a global audience of testers entirely on domestic silicon, that premise narrows — but only to serving. Inference is far less demanding than training, we do not know which chips or how efficiently, and the only source is the company that benefits from the story. Note it, do not bank on it.
#Coding

✓ Verified · 5 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

IBM's new open models were not taught to describe using a terminal. They were trained by using one.
2026-08-27
The nameless model that burned through 42 trillion tokens in six days has an owner: Z.ai says Ox Alpha is a GLM, and the weights are opening
2026-08-26
The Qwen4 preview is open: six billion parameters active per token, and on Alibaba's own table it fixes more real bugs than Opus 4.6
2026-08-26
An agent that keeps thinking after you stop typing — and the whole harness is under 10,000 lines of Bash
2026-08-26
Show the robot one video of a ten-minute job, and it does the ten-minute job — no retraining, nothing typed in
2026-08-26