⚡ Uncle Cat AI Radar
ModelsIndustryAgents

OpenAI Tests GPT-5.6 Sol at 750 Tokens per Second

A Cerebras-powered Ultrafast mode accelerates OpenAI’s top model by up to 14 times, initially for a limited customer group.

Frontier intelligence, interactive latency

OpenAI has introduced Ultrafast, a new inference mode that can run GPT-5.6 Sol at up to 750 generated tokens per second. The company says this represents as much as a fourteenfold speed increase, using Cerebras wafer-scale inference systems rather than the conventional GPU infrastructure serving the model’s standard mode.

The rollout is deliberately limited. OpenAI is working with an initial group of customers to identify where the additional speed produces enough practical value to justify a separate service tier. It has not announced general availability or public pricing, so Ultrafast should be understood as an active deployment trial rather than a default upgrade for every GPT-5.6 Sol user.

The performance target is significant because Sol is OpenAI’s most capable GPT-5.6 model and is normally used for complex technical work and long-running agent tasks. Higher intelligence often comes with longer responses and added reasoning latency. At 750 tokens per second, multi-step reasoning traces, code generation and iterative agent loops can complete quickly enough for use cases that previously required a smaller model or an asynchronous workflow.

Cerebras achieves the acceleration with wafer-scale processors designed to keep model weights and computation close together, reducing the communication bottlenecks common in large GPU clusters. The deployment also follows a multiyear OpenAI–Cerebras agreement covering a large build-out of high-speed inference capacity, making Ultrafast an early product-level expression of that infrastructure partnership.

Why it matters

Model quality is no longer the only frontier. For agents that repeatedly reason, call tools, inspect results and try again, latency compounds across every step. Making a top-tier model an order of magnitude faster can change which workflows are economically and ergonomically viable, from live coding assistance to real-time decision support. The unanswered question is price: without public economics or broad access, Ultrafast proves technical feasibility but not yet mass-market viability. Even so, it increases pressure on inference providers to compete on response time as aggressively as model labs compete on intelligence.

Sources