On August 13, 2026, OpenAI previewed Ultrafast, a new API service tier that runs GPT-5.6 Sol up to 14 times faster than Standard processing. The company says the tier is powered by Cerebras and can generate up to 750 output tokens per second.
What OpenAI actually said
Until now, real-time speed usually meant picking a smaller or more specialized model. OpenAI's pitch is that Ultrafast keeps Sol's intelligence and raises the speed. The company lists time-sensitive work: incident response, financial research and security, customer support and voice, commerce while a shopper is still deciding, and live research loops that used to wait overnight.
Cerebras says Ultrafast runs with the same intelligence as GPT-5.6 Sol Standard. Treat that as the vendor claim. It is not an Alberta business result.
Who can use it
Not everyone. OpenAI says GPT-5.6 Sol on Ultrafast is available today to a select group of API customers, with access expanding as capacity grows. There is a waitlist form. If a salesperson tells you it is already on every ChatGPT seat, ask for the API preview confirmation in writing.
What a local operator should do
Speed is useful for after-hours intake, voice, and support. Faster answers still need a review step when the work is a quote, a safety call, a payment, or an angry customer. Do not put Ultrafast on a live phone line just because the demo is snappy. Start with one recorded internal path: log review, a draft callback note, or a checkout help script.
