Summary
OpenAI announced a preview of Ultrafast mode, serving GPT-5.6 Sol at up to 750 output tokens per second, about 14x the standard mode. Inference runs on Cerebras wafer-scale hardware instead of GPUs. Access is limited to a small preview group, and pricing is undisclosed.
Why it matters
This is the first time OpenAI serves a frontier model on third-party, non-GPU inference silicon. It validates alternative inference hardware at frontier quality, and it signals that speed-tiered API products are becoming the next competitive axis.
Technical details
| Speed Output Toks | 750 |
|---|---|
| Speedup | 14x |
| Hardware | Cerebras wafer-scale |
| Access | limited preview |
| Pricing | undisclosed |
Tags
openaicerebrasinferencespeedpreviewgpt-5.6