OpenAI and Cerebras have launched a new service tier called Ultrafast Mode, now available in a limited preview for select API customers. This collaboration marks a significant shift in how frontier AI models handle high-stakes tasks by prioritizing output speed without sacrificing intelligence or accuracy.

The service powers the GPT-5.6 Sol model, achieving output speeds of up to 750 tokens per second. To demonstrate this capability, researchers compared the model against alternatives using Humanity's Last Exam, a benchmark featuring 2,500 complex questions. GPT-5.6 Sol on Ultrafast mode completed the entire set in 11 hours and 11 minutes, a process that took competing models more than 78 hours to finish. This represents a significant gain in efficiency for time-sensitive tasks.

Technical performance is driven by the Cerebras Wafer-Scale Engine. Traditional hardware often struggles with memory bandwidth, as model weights must frequently move between chip memory and off-chip storage. The Cerebras architecture bypasses this by packing 44 GB of SRAM directly on the wafer. This keeps model weights on the chip, allowing tokens to process through layers without the delays caused by off-chip data movement.

The implications for real-world application are substantial. For industries like cybersecurity, engineering, and finance, this speed allows for near real-time responses to critical incidents. Security teams can detect threats faster, and engineers can debug production outages in moments rather than waiting for extended processing times. Researchers also report that the speed allows them to maintain their focus on deep problem-solving rather than context-switching while waiting for outputs.

OpenAI is currently expanding access to this service to additional customers over time. This technology changes the threshold for what organizations can achieve when AI speed aligns with human thought patterns, making frontier models practical for direct integration into active workflows where every second matters.