OpenAI rolls out GPT-6.1 Sol Ultrafast access
OpenAI is adding an Ultrafast mode for GPT-6.1 Sol across its API, Codex and ChatGPT Work, charging a premium for faster agent responses.
OpenAI has begun rolling out an Ultrafast mode for GPT-6.1 Sol across the API, Codex and ChatGPT Work. The company says the mode can generate tokens at up to eight times the speed of the standard Sol configuration while targeting comparable intelligence for demanding work.
Pricing and availability
The API price is $12 per million input tokens and $60 per million output tokens. OpenAI is positioning the mode for latency-sensitive workloads such as debugging live incidents, agents navigating software interfaces and interactive applications where waiting for a long response changes the user experience.
In Codex and ChatGPT Work, access is limited to Pro 500, eligible usage-based Enterprise and credit-based Edu plans, with enterprise administrators required to enable it. OpenAI says the rollout covers all supported regions and includes data-residency support in the United States and European Union. The company also says it has added EU residency support for GPT-6.1 Sol Fast and GPT-6 Luna Fast.
The announcement is a deployment and pricing change rather than a new model release. It follows OpenAI’s earlier introduction of GPT-6.1 Sol and highlights a product strategy in which the same model family is sold in distinct latency tiers. That matters for agents: a faster model can reduce the time spent waiting between tool calls, but the economics depend on whether the additional output price is offset by shorter sessions or higher task completion rates.
Why it matters
Inference speed is becoming a product feature in its own right. For coding agents and computer-use systems, every extra pause can interrupt a workflow or force developers to reduce the number of reasoning steps. OpenAI is therefore selling responsiveness as part of model capability, not merely as an infrastructure metric.
The unresolved issue is value at task level. An eightfold speed claim does not establish eightfold productivity, and the premium output price will matter for long-running agents. Buyers will need measurements that combine latency, tool-call success and total cost per completed task.