Via gizmodo.com
OpenAI previews Ultrafast mode for GPT-5.6 Sol at up to 14 times faster speeds
The Cerebras powered service can generate up to 750 output tokens per second and will launch first through the OpenAI API.
OpenAI introduced an early preview of Ultrafast, a new service tier that runs GPT-5.6 Sol at speeds up to 14 times faster than Standard processing.
The service is powered by Cerebras and can generate as many as 750 output tokens per second. OpenAI said Ultrafast will initially launch through its API.
The company is positioning the tier for applications where response time is critical, including financial research, incident response, customer support, voice applications, commerce and live experimentation.
OpenAI said the goal is to bring its most capable model into real time workflows without requiring customers to switch to smaller models for faster responses.
An initial group of companies including Jane Street, Podium, Basis and Rogo has been testing the service across areas such as coding, financial research, commerce and customer support.
The news moving money, markets, and the world—before your day starts.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
OpenAI is also using Ultrafast internally for incident response and research. Engineers are testing it to analyze logs, traces and other information during active outages, while researchers are using the faster inference to shorten experimentation cycles.
The launch expands OpenAI’s partnership with Cerebras, which provides the infrastructure behind the new speed tier.
Ultrafast remains in preview with limited customer access, with OpenAI planning to expand availability as capacity increases.