New Models
2026-08-14
OpenAI's fastest model is not a smaller model — it is the same one, running on a rival's chips
OpenAI opened a limited preview of Ultrafast mode on 13 August: GPT-5.6 Sol served at up to 750 output tokens per second in the API, which the company puts at up to 14 times its standard speed. The speed does not come from a distilled or trimmed model but from Cerebras hardware, so the output is meant to be the same Sol customers already use. Access is restricted to a small group of API customers while capacity is expanded, and OpenAI has not published pricing. It points to incident response, customer support, financial analysis and e-commerce as the workloads where latency, rather than raw capability, is the constraint.
Why it mattersUntil now, getting a fast answer meant accepting a weaker one: a smaller model, a shorter context, a specialised fine-tune. Separating speed from capability changes which products are buildable — an agent that has to make twenty tool calls before it says anything is unusable at standard speed and unremarkable at 750 tokens a second. There is also a quieter admission in the announcement. OpenAI's fastest frontier serving runs on a competitor's silicon, not on Nvidia's and not on the infrastructure the company is spending hundreds of billions of dollars to build.
✓ Verified · 3 sources
▶ Related video: Why OpenAI Chose Cerebras Over NVIDIA for GPT-5.6 Sol
Read in the app — free, in 9 languages
Related stories
DeepSeek's cheap workhorse can now see — and on agent tasks that need eyes it says it is close to Anthropic's best
2026-08-21Show the robot once — three to twelve seconds — and it gets the job right 59 times out of 100 with no training at all
2026-08-21The model invents its own tasks, builds the rig to test them, then trains on the results — and DeepReinforce gave the weights away
2026-08-20A free model that fits on one graphics card scores 52 — then spends 22,276 thinking tokens drawing a picture
2026-08-18Zhipu did not build a new model. It kept training the old one — and says coding got 50 percent better.
2026-08-17