A startup called Inferact published a test on Wednesday using code it wrote for Google's chips. Sixteen Google TPUs produced 709 tokens a second on the Kimi K3 model, against 452 on Nvidia GB200s. Woosuk Kwon, who co-created the widely used vLLM serving engine, is one of the authors.