Cerebras has unveiled the CS-4, the latest generation of its wafer-scale machines, claiming up to 30x faster inference than production GPU systems and a record for the fastest inference commercially available.
The system is built around the company's new WSE-Turbo processor and shifts what Cerebras calls the inference Pareto frontier: up to 10x more throughput per watt than its predecessor CS-3, while generating tokens up to 30x faster than GPU-based systems. By cutting wafer-to-wafer interconnect latency to 2 microseconds, the CS-4 sustains more than 1,000 tokens per second even on models exceeding 10 trillion parameters — interactivity at a scale usually associated with batch processing.
The CS-4 is also the first machine built on the new Cerebras Nexus Platform Architecture, a modular design organized around three elements: compute, power, and I/O. Each "Wafer-Scale Backpack" folds the wafer, power conversion, direct liquid cooling, high-speed I/O, and control electronics into a single 3D package with 50% fewer components, cutting deployment time from days to hours. A new programmable I/O subsystem doubles bandwidth and lets wafers link within and across racks without a switch.
Power delivery has been moved to within 0.5 millimeters of the processor — roughly 100x closer than typical GPU boards — reducing losses and enabling higher operating frequencies. Cerebras says first shipments begin this quarter, with the PowerRack infrastructure installable before compute arrives.