aiminute. ← All AI news
Research AI Minute Newsroom 2026-08-07

Independent benchmarks: Qwen3.8-Max sits just below the US frontier — by working much harder

Independent benchmarks: Qwen3.8-Max sits just below the US frontier — by working much harder

Artificial Analysis scored Alibaba's Qwen3.8-Max at 56 on its Intelligence Index — above every model from Google, Meta and xAI, with only Anthropic's and OpenAI's top tiers and Kimi K3 higher — and second on its agentic GDPval-AA leaderboard at 1739 Elo, behind only Claude Opus 5 max at 1852. The gains come at a price: it averages 64 turns per agentic task versus 14 for Qwen3.7-Max, more than doubling the cost per task to $1.14.

Why it mattersIndependent numbers now put two Chinese open-weight releases inside the global top four — but the turn-count detail matters just as much: the gap is being closed with brute effort, and effort shows up on real deployment bills. Benchmark rank and cost-per-task are diverging, and buyers should read both.
#AI Agents

✓ Verified · 2 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

Twenty tries at the same task cut the best agent's score by a third.
2026-10-06
A model learned to write without the method that trains every AI.
2026-10-06
TikTok put a shopping chatbot inside the video you are watching.
2026-10-06
Wikipedia's owner says OpenAI agents may have caused a May outage.
2026-10-06
Meta's assistant keeps an hourly file on everyone in your life.
2026-10-05