⚡ Uncle Cat AI Radar
IndustryAgents

Cerebras CS-4 Pairs Three WSE-3 Turbo Processors for Faster Inference

Cerebras says its fourth-generation system delivers up to twice the inference speed of CS-3 through three WSE-3 Turbo processors and a redesigned rack-scale platform.

A system-level upgrade around WSE-3 Turbo

Cerebras unveiled the CS-4 at its Supernova event, presenting a fourth-generation system built around three new WSE-3 Turbo processors. It is not simply a CS-3 running the same wafer at a higher frequency: Cerebras describes CS-4 as a new rack-scale platform combining more powerful processors with redesigned wafer-to-wafer communication, power delivery, cooling and input-output systems.

In a company demonstration using OpenAI’s GPT-OSS model, the CS-4 generated 4,465 tokens per second, compared with 2,308 tokens per second on the CS-3. Cerebras contrasted that result with 131 tokens per second for a conventional GPU system. These are vendor-reported measurements rather than independently reproduced benchmarks, and performance will vary with the model, workload and serving configuration.

Cerebras says CS-4 can deliver up to twice the inference performance of CS-3 and up to ten times more token throughput per watt. Those figures refer to internal benchmarks and projections across selected workloads and should not be treated as universal production results.

A modular rack for split inference

CS-4 is the first system based on Cerebras’s Nexus rack-scale architecture. Its rear-mounted Wafer-Scale Backpack combines power conversion, direct liquid cooling, high-speed input-output and control electronics around each wafer. Cerebras says the module has 50% fewer components, uses 60% more automated manufacturing and reduces deployment time from days to hours.

The system also provides native support for disaggregated inference, in which different hardware handles prompt processing and token generation. Cerebras identifies AMD Helios and AWS Trainium among the platforms that can perform prefill before transferring model state to CS-4 for latency-sensitive decoding. This is broader than an AMD-only configuration and allows operators to combine Cerebras hardware with GPU- or ASIC-based prefill systems.

Cerebras says CS-4 can reduce wafer-to-wafer communication latency to as little as two microseconds and support more than 1,000 tokens per second for models exceeding ten trillion parameters. The latter claim is based on company extrapolation from internal testing, not a demonstrated independent benchmark.

Why it matters

CS-4 shows how inference gains can come from coordinated changes to processors, interconnects, power, cooling and rack design rather than from a semiconductor-process transition alone. The architecture also reflects growing interest in assigning prefill and decoding to different types of hardware.

Cerebras says the first CS-4 shipments will begin in the third quarter of 2026. Independent production testing will still be needed to establish how its speed, aggregate throughput, energy use and economics compare across real workloads.

Sources