aiminute. ← All AI news
Research 2026-08-07

Independent benchmarks: Qwen3.8-Max sits just below the US frontier — by working much harder

Independent benchmarks: Qwen3.8-Max sits just below the US frontier — by working much harder

Artificial Analysis scored Alibaba's Qwen3.8-Max at 56 on its Intelligence Index — above every model from Google, Meta and xAI, with only Anthropic's and OpenAI's top tiers and Kimi K3 higher — and second on its agentic GDPval-AA leaderboard at 1739 Elo, behind only Claude Opus 5 max at 1852. The gains come at a price: it averages 64 turns per agentic task versus 14 for Qwen3.7-Max, more than doubling the cost per task to $1.14.

Why it mattersIndependent numbers now put two Chinese open-weight releases inside the global top four — but the turn-count detail matters just as much: the gap is being closed with brute effort, and effort shows up on real deployment bills. Benchmark rank and cost-per-task are diverging, and buyers should read both.
#AI Agents

✓ Verified · 2 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

Machine learning read the shape of sick brain cells and picked out nine already-approved drugs that calmed them down
2026-08-21
DeepSeek's cheap workhorse can now see — and on agent tasks that need eyes it says it is close to Anthropic's best
2026-08-21
Given four hours and a GPU to improve the way AI is trained, the best agent scored 0.25 out of 1 — and most never tried
2026-08-21
ChatGPT can now read your iMessages and send them — and the setting that lets it skip asking is the one OpenAI warns about
2026-08-21
The agent invented a second person to vouch for its code. A 24-year-old in Texas refused to believe either of them.
2026-08-21