New Models
AI Minute Newsroom
2026-08-14
OpenAI's fastest model is not a smaller model — it is the same one, running on a rival's chips
OpenAI opened a limited preview of Ultrafast mode on 13 August: GPT-5.6 Sol served at up to 750 output tokens per second in the API, which the company puts at up to 14 times its standard speed. The speed does not come from a distilled or trimmed model but from Cerebras hardware, so the output is meant to be the same Sol customers already use. Access is restricted to a small group of API customers while capacity is expanded, and OpenAI has not published pricing. It points to incident response, customer support, financial analysis and e-commerce as the workloads where latency, rather than raw capability, is the constraint.
Why it mattersUntil now, getting a fast answer meant accepting a weaker one: a smaller model, a shorter context, a specialised fine-tune. Separating speed from capability changes which products are buildable — an agent that has to make twenty tool calls before it says anything is unusable at standard speed and unremarkable at 750 tokens a second. There is also a quieter admission in the announcement. OpenAI's fastest frontier serving runs on a competitor's silicon, not on Nvidia's and not on the infrastructure the company is spending hundreds of billions of dollars to build.
✓ Verified · 3 sources
▶ Related video: Why OpenAI Chose Cerebras Over NVIDIA for GPT-5.6 Sol
Read in the app — free, in 9 languages
Related stories
Reflection will hand out a 501-billion-parameter model for free.
2026-10-06Microsoft's new transcriber starts writing before you finish speaking.
2026-10-05An American lab is about to give away a China-class open model.
2026-10-05A video site now gives away the best-scoring open translation model.
2026-10-05Half the people on a short video call thought the AI was human.
2026-10-04