
Cerebras doubles CS-4 performance without a new chip
Cerebras doubled CS-3 performance in the CS-4 on the same 5nm WSE-3 wafer and the same 44 GB of on-wafer memory, buying the gain from higher clock speeds, better power delivery and cooling, and three wafers per rack instead of two. Andrew Feldman cites up to 4,400 tokens per second per user and up to 30x Nvidia GPU setups, with OpenAI running Codex Spark on Cerebras hardware — inference speed leadership is now a packaging and cooling contest, not a node-shrink one. SemiAnalysis reads the networking gains as small, crediting other optimizations for most of the jump.
Source: the-decoder.com ↗
A single rack now holds three wafers instead of two and delivers up to 4,400 tokens per second per user.
Cerebras
Why this matters
- → Performance gains now come from packaging and cooling, not chip shrinks—shifting the AI hardware bottleneck.
- → 30x speedup over Nvidia GPUs on the same 5nm process signals a fundamental architectural advantage.
- → Inference speed leadership is increasingly determined by system design, not process node.
Cooling, not silicon