Launch · Hacker News ·

Cerebras CS-4 promises AI inference 30x faster than GPUs

Cerebras unveiled the CS-4, a wafer-scale system built on its new WSE-Turbo chip claiming up to 30x faster AI inference than production GPU systems.

Based on reporting by Hacker News — analysis by dalili

Cerebras has unveiled the CS-4, the latest generation of its wafer-scale machines, claiming up to 30x faster inference than production GPU systems and a record for the fastest inference commercially available.

The system is built around the company's new WSE-Turbo processor and shifts what Cerebras calls the inference Pareto frontier: up to 10x more throughput per watt than its predecessor CS-3, while generating tokens up to 30x faster than GPU-based systems. By cutting wafer-to-wafer interconnect latency to 2 microseconds, the CS-4 sustains more than 1,000 tokens per second even on models exceeding 10 trillion parameters — interactivity at a scale usually associated with batch processing.

The CS-4 is also the first machine built on the new Cerebras Nexus Platform Architecture, a modular design organized around three elements: compute, power, and I/O. Each "Wafer-Scale Backpack" folds the wafer, power conversion, direct liquid cooling, high-speed I/O, and control electronics into a single 3D package with 50% fewer components, cutting deployment time from days to hours. A new programmable I/O subsystem doubles bandwidth and lets wafers link within and across racks without a switch.

Power delivery has been moved to within 0.5 millimeters of the processor — roughly 100x closer than typical GPU boards — reducing losses and enabling higher operating frequencies. Cerebras says first shipments begin this quarter, with the PowerRack infrastructure installable before compute arrives.

Key takeaways

  • CS-4 claims up to 30x faster inference than production GPU systems
  • New WSE-Turbo chip and modular Nexus platform debut with the system
  • Over 1,000 tokens per second on models past 10 trillion parameters
  • Wafer-Scale Backpack design cuts deployment from days to hours

Why it matters

Inference cost and speed are becoming the battleground of the AI economy as workloads shift from training to serving. Cerebras is betting that wafer-scale integration can outpace GPU clusters on both tokens per second and efficiency. If CS-4 numbers hold in independent benchmarks, hyperscalers and sovereign AI programs gain a credible second option for serving trillion-parameter and agentic workloads — pressure on GPU pricing that the whole market feels.

Related

  1. The Verge ·

    Roland Launches Melody Flip, Its First Generative AI Music Tool

  2. The Verge ·

    Anker launches MindBase, an on-device AI hub for smart home security