14x Speed Boost via Cerebras Compute
OpenAI officially introduced a new inference mode dubbed 'Ultrafast' for its GPT-5.6 Sol model in mid-August 2026, capable of delivering output speeds of up to 750 tokens per second. According to OpenAI announcements reported by The Decoder, this new mode accelerates GPT-5.6 Sol's response rate by up to 14 times compared to standard inference speeds. The breakthrough is powered by specialized hardware from Cerebras, stemming from a $10 billion commercial partnership between the two companies.
Shifting to a Three-Tiered API Pricing Model
The launch of Ultrafast signals a major shift in OpenAI's API commercialization strategy. The Decoder reports that API pricing is now split into three distinct tiers:
* Standard: Default baseline inference latency and standard pricing. * Fast: Optimized low-latency tier for interactive workloads. * Ultrafast: High-throughput 750 tokens/second tier powered by specialized wafer-scale engines.
Instead of billing solely on input and output token volumes, OpenAI has effectively productized inference speed and latency as an independent pricing dimension. This gives developers and enterprise customers greater flexibility to optimize costs based on the specific latency tolerances of their applications.
Targeted Enterprise Rollout and Real-Time AI Workloads
OpenAI noted that Ultrafast mode is currently accessible in a limited preview through the OpenAI API for select enterprise partners. Broader access will roll out as Cerebras expands and deploys additional compute infrastructure. This initial rollout focuses on compute-heavy tasks requiring real-time responsiveness, sub-second latency, and autonomous AI agents that rely on rapid multi-step reasoning loops.
OpenAI has not yet disclosed detailed per-token pricing across the Standard, Fast, and Ultrafast tiers, nor has it provided a specific timeline for rolling out Ultrafast mode to general users beyond the API interface.