aiminute. ← All AI news
New Models AI Minute Newsroom 2026-08-14

OpenAI's fastest model is not a smaller model — it is the same one, running on a rival's chips

OpenAI's fastest model is not a smaller model — it is the same one, running on a rival's chips

OpenAI opened a limited preview of Ultrafast mode on 13 August: GPT-5.6 Sol served at up to 750 output tokens per second in the API, which the company puts at up to 14 times its standard speed. The speed does not come from a distilled or trimmed model but from Cerebras hardware, so the output is meant to be the same Sol customers already use. Access is restricted to a small group of API customers while capacity is expanded, and OpenAI has not published pricing. It points to incident response, customer support, financial analysis and e-commerce as the workloads where latency, rather than raw capability, is the constraint.

Why it mattersUntil now, getting a fast answer meant accepting a weaker one: a smaller model, a shorter context, a specialised fine-tune. Separating speed from capability changes which products are buildable — an agent that has to make twenty tool calls before it says anything is unusable at standard speed and unremarkable at 750 tokens a second. There is also a quieter admission in the announcement. OpenAI's fastest frontier serving runs on a competitor's silicon, not on Nvidia's and not on the infrastructure the company is spending hundreds of billions of dollars to build.

✓ Verified · 3 sources

▶ Related video: Why OpenAI Chose Cerebras Over NVIDIA for GPT-5.6 Sol
WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

Reflection will hand out a 501-billion-parameter model for free.
2026-10-06
Microsoft's new transcriber starts writing before you finish speaking.
2026-10-05
An American lab is about to give away a China-class open model.
2026-10-05
A video site now gives away the best-scoring open translation model.
2026-10-05
Half the people on a short video call thought the AI was human.
2026-10-04