OpenAI Just Made Its Smartest Model Run 14 Times Faster. It Didn't Make It Dumber.
OpenAI's limited Ultrafast preview keeps GPT-5.6 Sol's full reasoning capability while using Cerebras hardware to reach up to 750 output tokens per second.
For most of the last three years, AI companies have offered developers a blunt trade-off: pick a smart model and wait, or pick a fast model and settle. OpenAI's newest API tier is a bet that you shouldn't have to choose. On August 13, the company began a limited preview of Ultrafast mode for GPT-5.6 Sol — its flagship reasoning model — running up to 14 times faster than standard processing, at speeds up to 750 output tokens per second, without swapping in a smaller, weaker model to get there.
The trick: don't shrink the model, speed up the chip
Every previous "fast" option from a major AI lab has involved some kind of compromise. Distill a smaller model. Cache aggressively. Trim reasoning steps. Ultrafast takes a different approach: it keeps GPT-5.6 Sol's full capabilities intact and instead changes where the model runs.
The acceleration comes from Cerebras Systems, the wafer-scale chip company OpenAI struck a landmark compute deal with back in January 2026 — a multi-year agreement reported at more than $10 billion for roughly 750 megawatts of Cerebras compute delivered through 2028. Cerebras' Wafer-Scale Engine chips are built for exactly this kind of low-latency inference workload, trading the flexibility of GPU clusters for raw, dedicated speed on a single, enormous piece of silicon.
OpenAI's own framing is telling: Ultrafast "puts the company's most intelligent model into low-latency settings, seeking to deliver more useful work per second without making that tradeoff." In other words, this isn't a new, cheaper GPT-5.6 variant — it's the same model, just running on hardware built to answer faster.
What 14x actually looks like in practice
Raw throughput numbers are easy to gloss over, so OpenAI's early testers have been showing their work. One example making the rounds: a financial analysis dashboard that took 12 minutes and 20 seconds to generate on standard GPT-5.6 Sol came back in 1 minute and 50 seconds on Ultrafast — roughly a 6.7x wall-clock improvement on a real, multi-step task, even if the peak theoretical throughput is higher.
That gap between "peak tokens per second" and "how much faster did my actual job finish" matters, because it points to where this tier is genuinely useful: not chatbot small talk, but workflows where a model has to think in a loop and a human — or another system — is waiting on the other end. OpenAI says it's already using Ultrafast internally for incident response (reading logs, tracing failures, synthesizing what happened across systems) and for research workflows that involve chaining together multiple searches and tool calls.
Early access customers echo that framing. Jane Street's John Crepezzi, describing the AI Assistants team's experience with the preview, put it this way: "The increase in speed brought by Cerebras is impressive. It enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them." Podium, Basis, and Rogo are also named as early testers, spanning customer support, financial research, and commerce use cases.
Not new, but a new tier of "new"
This isn't OpenAI's first attempt to sell speed as a feature. The company already offers a "Fast Mode" for GPT-5.6 Sol that runs up to 2.5x faster than standard, at roughly double the price. Ultrafast sits well above that — both in speed and, presumably, in cost, though OpenAI hasn't published pricing, quotas, region availability, or a general-availability timeline. For now, access is limited to a small group of customers, with a sign-up form for everyone else waiting to get in line as capacity expands.
That capacity question is really the whole story. Wafer-scale chips are extraordinary at raw inference speed, but they're also expensive, hard to manufacture at volume, and nothing like the commodity GPU fleets that power most cloud AI today. Whether Ultrafast becomes a mainstream API option or stays a premium niche for latency-obsessed customers will depend entirely on how fast Cerebras can build out that 750-megawatt commitment — and how many other model providers decide dedicated silicon is worth the bet too.
The bigger pattern
Ultrafast lands in the middle of a broader shift in how AI labs compete. OpenAI and Anthropic have spent recent months cutting prices on standard inference even as new entrants like DeepSeek raise theirs, while simultaneously investing in premium speed tiers for customers who'll pay more to skip the wait. Increasingly, the real battleground isn't just "whose model scores higher on a benchmark" — it's "whose infrastructure can deliver that intelligence fast enough to be useful in a live workflow." For an industry that spent 2023 and 2024 chasing raw capability, 2026's fight is shaping up to be about latency, orchestration, and the unglamorous plumbing that turns a smart model into a fast one.
Sources
Cerebras Systems: Cerebras Powers Ultrafast Mode for OpenAI's GPT-5.6 Sol — https://investors.cerebras.ai/news-releases/news-release-details/cerebras-powers-ultrafast-mode-openais-gpt-56-sol
The Decoder: GPT-5.6 Sol goes 14x faster as OpenAI launches Ultrafast mode powered by Cerebras — https://the-decoder.com/gpt-5-6-sol-goes-14x-faster-as-openai-launches-ultrafast-mode-powered-by-cerebras/
Help Net Security: OpenAI's GPT-5.6 Sol runs up to 14x faster with Ultrafast mode — https://www.helpnetsecurity.com/2026/08/14/openais-gpt-5-6-sol-runs-up-to-14x-faster-with-ultrafast-mode/
TestingCatalog: OpenAI previews Ultrafast API tier for GPT-5.6 Sol — https://www.testingcatalog.com/openai-previews-ultrafast-api-tier-for-gpt-5-6-sol/
Data Center Dynamics: OpenAI signs $10 billion deal with Cerebras — https://www.datacenterdynamics.com/en/news/openai-signs-10-billion-deal-with-cerebras-with-750mw-of-big-chip-compute/