Morning Edition · Friday, August 14, 2026Published at 2:27 AM EDT · New York
The Ultrafast service tier claims up to 14 times the throughput of standard serving on the same model, drawing on a supply agreement Cerebras values at more than $20 billion.

OpenAI opened a limited preview of Ultrafast, an application programming interface (API) service tier that runs GPT-5.6 Sol at up to 750 output tokens per second, which the company puts at up to 14 times the throughput of the same model on…
Track frontier labs, chips, export controls, model releases, regulation, and AI infrastructure.
The Global Intelligence Brief stays free.
Part of a tracked trend
The Inference-Cost Efficiency Race
Techniques that cut tokens generated and KV-cache memory per query will keep compressing the marginal cost of serving reasoning models, making inference efficiency a recurring competitive axis alongside raw capability.
Start a discussion in Townsquare.
More from this edition
Comments
0No comments yet.