New Models
AI Minute Newsroom
2026-08-18
A free model that fits on one graphics card scores 52 — then spends 22,276 thinking tokens drawing a picture
Alibaba released Qwen3.8-27B on 14 August under an Apache 2.0 licence: 27 billion parameters, able to read images as well as text, a 256,000-token context window, and about 17GB once quantised, which puts it within reach of a single consumer graphics card. Artificial Analysis scores it 52 on its intelligence index, against a median of 9 for open-weight models of similar size. Developer Simon Willison then documented the catch: the model ships set to its highest reasoning effort. Asked to draw an SVG, it consumed 22,276 reasoning tokens to produce 3,223 tokens of output and took 21 minutes; the same request with reasoning switched off finished in 137 seconds using 3,715 tokens. Across its benchmark run Artificial Analysis measured 160 million tokens generated, against a median of 43 million.
Why it mattersBenchmark numbers get quoted as though they settle the question, and here the same model is either excellent or unusable depending on one setting that arrives switched on. That matters most for exactly the people an Apache-2.0 release is aimed at: if you are running weights on your own machine, the bill is not dollars per token but the hours your hardware spends thinking, and the default quietly multiplies them. Before judging any local model, check what reasoning effort it defaults to.
✓ Verified · 3 sources
▶ Related video: Qwen3.8 27B Local Test with llama.cpp | Best Small Local Model? | Coding & Agentic Work | 🔴 Live
Read in the app — free, in 9 languages
Related stories
Reflection will hand out a 501-billion-parameter model for free.
2026-10-06Microsoft's new transcriber starts writing before you finish speaking.
2026-10-05An American lab is about to give away a China-class open model.
2026-10-05A video site now gives away the best-scoring open translation model.
2026-10-05OpenAI will ship a Codex upgrade daily for 28 days or reset limits.
2026-10-05