Industry
AI Minute Newsroom
2026-08-19
Same chip, same factory, roughly double the clock — and a rack that draws as much power as forty homes
Cerebras announced the CS-4 on 18 August, the first machine in what it calls the Nexus rack-scale platform. A rack holds up to three of its dinner-plate-sized wafers and adds up to 750 petaflops of sparse FP16 compute, 129.6 petabytes per second of memory bandwidth, 7.2 terabits per second of input and output, and 132 gigabytes of on-wafer memory. The company says it is up to twice as fast as the CS-3 and delivers up to 30 times more tokens per second per user than GPU systems on frontier models. The silicon is not new: it is the same TSMC 5-nanometre design, and The Register calculates that the doubling comes mainly from running it at roughly twice the clock speed, around 2.8 gigahertz, which pushes each wafer to an estimated 33 kilowatts and a full rack to between 120 and 140. First systems come online later this quarter. Cerebras is also pairing its racks with AMD Helios systems and AWS Trainium chips, which handle the prompt-reading half of the work while Cerebras does the answer-writing half.
Why it mattersRead the headline number carefully. The 750 petaflops figure depends on sparsity, which as a rule does not help language model inference much, and the peak memory bandwidth may be theoretical rather than achievable — The Register flags both. What is not in doubt is the direction: the fastest way to serve a model to one impatient user is no longer more chips but hotter ones, and the electricity bill is the honest measure. A rack drawing 140 kilowatts is not a machine you slot into an existing data hall. It is a reason to build a different data hall.
✓ Verified · 3 sources
▶ Related video: Beyond GPUs: Cerebras' Wafer-Scale Engine for Lightning-Fast AI Inference
Read in the app — free, in 9 languages
Related stories
China's top model now trails America's by three percent.
2026-10-06Banks now take Nvidia chips as collateral on Asian data centre loans.
2026-10-06ChatGPT is signing its fake cartoons with real cartoonists' names.
2026-10-06Qualcomm will license Huawei's way of stacking two chip wafers.
2026-10-06Wikipedia's owner says OpenAI agents may have caused a May outage.
2026-10-06