OpenAI has introduced 'Ultrafast' mode, leveraging Cerebras hardware to boost the processing speed of its flagship AI model, 'GPT-5.6 Sol,' by up to 14 times.
Imagine this: This morning, you asked an AI to summarize a long, complex meeting document. Usually, you would have picked up your teacup and waited a few seconds for the results, but the moment you hit the Enter key, the text flooded the screen as if someone were transcribing it in real-time right beside you.
The closing of the gap between the speed of our thoughts and the speed of AI’s response is the future that OpenAI’s new technology aims for.
Why It Matters
Until now, when we conversed with AI, we faced a wall known as “latency” (the waiting time between giving a command and the appearance of the result). It required a certain amount of time for AI to process a question and generate an answer. While this short delay might be fine for casual conversation, it often felt like a major obstacle in business environments where real-time analysis of complex data is crucial or where speed is everything.
The ‘Ultrafast’ mode announced by OpenAI focuses on breaking down this wall of latency. Beyond simple convenience, this creates an environment where AI can provide more precise and immediate help as a real-time partner. OpenAI
The Explainer
To understand this technology, you must first know the concept of a ‘token’ (the smallest unit of words or strings that AI understands). Every time we talk to an AI, it processes and combines numerous tokens to construct an answer.
To use an easy analogy, the existing standard processing method was like ‘a single scribe carefully writing down characters one by one with a pen.’ While it produced excellent writing, there were physical limits to the speed.
The new Ultrafast mode changes this process to something like ‘a modern high-speed copier outputting massive amounts of documents in an instant.’ OpenAI To achieve this, OpenAI introduced specialized hardware technology from a company called Cerebras. StockTitan Thanks to this, the GPT-5.6 Sol model can now move 14 times faster than before, capable of pouring out up to 750 tokens per second. OpenAI This is an overwhelming figure that easily surpasses the average speed at which a human reads.
Where We Stand
Currently, Ultrafast mode is offered as an API (Application Programming Interface) service tier from OpenAI. 9to5Mac However, it is not yet available for everyone to use immediately. It is currently in a ‘Preview’ phase, released only to a select group of customers. Хабр In other words, it is better to understand this as a process of verifying and refining the technology’s possibilities rather than a full-scale commercial service.
What’s Next
What does it mean for AI’s response speed to become 14 times faster? Before long, we will encounter new tools that allow us to converse seamlessly with AI while watching the screen or process massive amounts of data in an instant. As OpenAI continues to overcome technical limits one by one, the day when this ‘Ultrafast’ technology is reliably provided to more users is not far off. We can look forward to a life with smarter, faster AI unfolding before us.
MindTickleBytes’ AI Reporter Perspective
Speed is not just a matter of numbers. It is the key factor that determines how deeply AI can penetrate our daily lives. This update will be an important inflection point where AI transitions from being a simple ‘knowledge provider’ into a ‘real-time collaborative partner.’ Just as the slow typewriter was replaced by the computer that processes information in an instant, our way of working will also undergo a fundamental change.
References
- Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed | OpenAI https://openai.com/index/previewing-ultrafast/
- Previewing Ultrafast mode: GPT‑5.6 Sol at up to 14X the speed - YouTube https://www.youtube.com/watch?v=WCwT4gWpHmI
- OpenAI previews ‘Ultrafast’ GPT-5.6 Sol running up to 14 times faster - 9to5Mac https://9to5mac.com/2026/08/13/openai-previews-ultrafast-gpt-5-6-sol-running-up-to-14-times-faster/
- OpenAI снизила цены на GPT-5.6 Luna и Terra и запустила… / Хабр https://habr.com/ru/companies/bothub/news/1065066/
- Cerebras Powers Ultrafast Mode for OpenAI’s GPT-5.6 Sol | CBRS Stock News https://www.stocktitan.net/news/CBRS/cerebras-powers-ultrafast-mode-for-open-ai-s-gpt-5-6-x2tvrw6nodi8.html
- Increased the model's intelligence by 14 times
- Improved processing speed by up to 14 times
- Switched usage fees to free
- NVIDIA
- Cerebras
- 100 tokens per second
- 750 tokens per second
- 1,000 tokens per second