aiminute. ← All AI news
New Models AI Minute Newsroom 2026-08-26

The Qwen4 preview is open: six billion parameters active per token, and on Alibaba's own table it fixes more real bugs than Opus 4.6

The Qwen4 preview is open: six billion parameters active per token, and on Alibaba's own table it fixes more real bugs than Opus 4.6

Alibaba's Qwen team has opened the weights of Qwen3.8-Flash-Next, the model it put a countdown clock on this morning. It is a multimodal mixture-of-experts system: 125 billion parameters in the main network, plus a 51-billion-parameter n-gram embedding table and a 4-billion-parameter multi-token prediction layer — about 180 billion in total in BF16, of which roughly 6 billion are active for any given token. The architecture is new: a hybrid of Gated DeltaNet and Qwen Sparse Attention, 512 experts with ten routed and one shared, a gated residual, and a training recipe mixing the Muon and AdamW optimisers. Context is 262,144 tokens natively and extends to a million. Qwen's own published table puts it at 62.5 on SWE-bench Pro against 53.4 for Claude Opus 4.6 Max, 81.0 on SWE-bench Multilingual against 77.5, and 58.7 on DeepSWE 1.1 — the lab's numbers, on the lab's harness, not yet independently reproduced. The team says training cost about a ninth of what Qwen3.7-Plus cost. The weights are on Hugging Face and ModelScope under a qwen-community-1.0 licence: 131 files, roughly 360GB, with an FP8 build that halves the memory. Qwen calls it an early preview of the architecture Qwen4 will be built on, and says a hosted version follows at $0.16 per million input tokens and $0.47 per million output.

Why it mattersThe number that matters is not 62.5, it is six billion. A model that lights up only six billion parameters per token runs at roughly the speed of a small model while carrying the knowledge of a large one, which is why the FP8 build is within reach of one well-equipped workstation instead of a rack. If the coding scores survive independent testing, the gap between what you rent from a frontier lab and what you can run yourself narrows again — and it narrows on the axis that costs money, which is tokens per second per watt. The caveat is plain: these are the maker's own benchmarks, published the same day as the weights, and the first thing anyone should do is check them.
#Coding

✓ Verified · 4 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

The nameless model that burned through 42 trillion tokens in six days has an owner: Z.ai says Ox Alpha is a GLM, and the weights are opening
2026-08-26
An agent that keeps thinking after you stop typing — and the whole harness is under 10,000 lines of Bash
2026-08-26
Show the robot one video of a ten-minute job, and it does the ten-minute job — no retraining, nothing typed in
2026-08-26
Alibaba has put a countdown clock on the first piece of Qwen4 — the weights open tonight at 23:00 in Beijing
2026-08-26
Alibaba has posted a page for a model it calls a preview of Qwen4, due 26 August — the specs went up, then came down
2026-08-25