Summary

OpenAI announced a preview of Ultrafast mode, serving GPT-5.6 Sol at up to 750 output tokens per second, about 14x the standard mode. Inference runs on Cerebras wafer-scale hardware instead of GPUs. Access is limited to a small preview group, and pricing is undisclosed.

Why it matters
This is the first time OpenAI serves a frontier model on third-party, non-GPU inference silicon. It validates alternative inference hardware at frontier quality, and it signals that speed-tiered API products are becoming the next competitive axis.
Technical details
Speed Output Toks 750
Speedup 14x
Hardware Cerebras wafer-scale
Access limited preview
Pricing undisclosed
Tags
openaicerebrasinferencespeedpreviewgpt-5.6