OpenAI Releases ‘Ultrafast’ Mode to Make GPT-5.6 Sol 14x Faster

chatgpt

OpenAI Releases ‘Ultrafast’ Mode to Make GPT-5.6 Sol 14x Faster with Cerebras hardware and up to 750 output tokens per second.

OpenAI releases ‘Ultrafast’ mode to make GPT-5.6 Sol 14x faster

OpenAI releases ‘Ultrafast’ mode to make GPT-5.6 Sol 14x faster, marking a significant shift in how frontier AI models can be deployed for real-time applications. The new service tier, announced on August 13, 2026, leverages Cerebras wafer-scale hardware to deliver up to 750 output tokens per second while maintaining the full intelligence of GPT-5.6 Sol.

Ultrafast is designed for use cases where every second matters, such as customer service, financial analysis, incident response, and AI agent workflows. OpenAI says the mode brings its most intelligent model into products and workflows that previously required smaller, faster models due to latency constraints.

Speed and performance

In standard mode, GPT-5.6 Sol generates around 53 output tokens per second. Ultrafast increases that to up to 750 tokens per second, a 14x improvement. Compared to OpenAI’s existing Fast mode, which was about 2.5x faster than Standard, Ultrafast is roughly 5.6x faster still.

OpenAI highlights several benchmark results:

  • On Humanity’s Last Exam, a 2,500-question AI benchmark, GPT-5.6 Sol in Ultrafast mode completes the test in about 11 hours 11 minutes, compared to 78 hours 27 minutes for Claude Fable 5.
  • On GDP-Val, a measure of real-world task performance, Ultrafast mode completes work about 5.6x faster than Standard mode.

These figures suggest that Ultrafast is not just about raw token speed, but also about reducing end-to-end task time for complex workflows.

Powered by Cerebras

Ultrafast runs on Cerebras wafer-scale inference hardware, which is designed to keep large models in memory and reduce data movement bottlenecks. This allows GPT-5.6 Sol to run at high speed without splitting the model across multiple chips or sacrificing quality.

The collaboration with Cerebras is a notable development because it shows OpenAI diversifying its inference infrastructure beyond traditional GPU-based setups. For enterprises, this could mean more predictable latency and better performance for time-sensitive applications.

Availability and use cases

Ultrafast is currently available as a limited preview to a select group of OpenAI API customers. OpenAI says it is already being tested in areas such as:

  • Real-time customer support and chatbots.
  • Financial market analysis and trading assistance.
  • Incident response and security operations.
  • E-commerce recommendations and personalisation.
  • Multi-step AI agent workflows.

Developers interested in the preview can join a waitlist through OpenAI’s API portal. The company has not yet announced a broader rollout timeline or pricing details.

Why this matters

Until now, there has been a trade-off between intelligence and speed: the most capable models were also the slowest, forcing developers to choose between quality and latency. Ultrafast challenges that assumption by showing that a frontier model can run at near-real-time speeds for certain workloads.

If OpenAI can scale Ultrafast reliably, it could change how AI is integrated into products. Real-time voice assistants, live coding copilots, interactive tutoring, and high-frequency decision support become more feasible when the underlying model can respond in fractions of a second without dropping to a weaker tier.

Summary: OpenAI has released Ultrafast mode for GPT-5.6 Sol, using Cerebras hardware to achieve up to 750 output tokens per second and 14x faster inference than Standard. The mode is in limited preview for API customers and targets real-time enterprise and agent workflows.

Read Previous

Facebook Launches Standalone Creator Studio App With AI Tools

Read Next

WhatsApp’s New AI Tool Can Warn Users About Scam Messages