aiminute. ← All AI news
Research AI Minute Newsroom 2026-09-27

Google's chips ran a Chinese model 57% faster than Nvidia's did.

Google's chips ran a Chinese model 57% faster than Nvidia's did.

A startup called Inferact published a test on Wednesday using code it wrote for Google's chips. Sixteen Google TPUs produced 709 tokens a second on the Kimi K3 model, against 452 on Nvidia GB200s. Woosuk Kwon, who co-created the widely used vLLM serving engine, is one of the authors.

Why it mattersNvidia's lead rests partly on software nobody else had matched, and outsiders are now matching it. More places to run a model cheaply means lower prices for the tools you pay for.

✓ Verified · 2 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

OpenAI and Anthropic are probing tens of thousands of incidents.
2026-09-27
A model won NetHack after 37,140 turns as a dwarf warrior.
2026-09-26
Outsiders rebuilt the code OpenAI's agents used to hack Hugging Face.
2026-09-26
The best model found 78% of the facts in a patient's file.
2026-09-26
Claude pushed a physics calculation past the record a human set.
2026-09-26