Cerebras Wafer-Scale Engine (WSE) chips are processors built from an entire silicon wafer rather than from a small die cut out of one. They combine AI compute cores, on-chip SRAM and a communication fabric on that wafer, aiming to keep computation and data movement close together. That design differs from conventional GPU systems, which use packaged GPU dies and may split large AI models across multiple processors.
The WSE is the chip; CS-3 and CS-4 are complete AI systems built around WSE processors. The architecture can suit particular AI workloads, but its design and headline specifications do not establish that it is faster or better than every GPU system.
What a Cerebras Wafer-Scale Engine does
A WSE provides the compute, memory and on-chip communication used to run AI workloads, including model training and inference. Cerebras introduced WSE-3 in 2024 as the processor inside its CS-3 system. The company’s current product information describes WSE-3 Turbo (WSE-3T) as powering the rack-scale CS-4 system. These are different product levels: a processor is not the same thing as the complete computer system that houses, powers, cools and connects it.
Cerebras says its 2024 WSE-3 has 4 trillion transistors, 900,000 AI-optimized compute cores, 125 petaflops of peak AI performance, 44 GB of on-chip SRAM and a 5 nm process. These are manufacturer-published specifications, not independent measures of performance on a particular model or task. Cerebras’ WSE-3 announcement gives its launch details.
#1 Best Overall
- Axial-tech fans now feature a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
- 2.5-slot design allows for greater build compatibility while maintaining cooling performance
- 0dB technology lets you enjoy light gaming in relative silence
- Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
- Dual ball fan bearings last up to twice as long as sleeve bearing designs
What “wafer-scale” means
Processors are typically fabricated on a silicon wafer, then the wafer is cut into smaller dies that are packaged as individual chips. Cerebras instead retains a processed wafer as one large processor. Sandia’s explanation of the approach describes WSE-3 cores positioned close to on-wafer SRAM, with communication across the processor handled by an on-wafer fabric. The aim is to reduce the distance data must travel and the need to coordinate work across separate chips.
Keeping a wafer intact creates manufacturing challenges: imperfections can occur across a large piece of silicon. Cerebras says its design uses redundant compute cores and routing, with a fail-in-place approach that disables flaws and routes around them. This is the company’s account of how it addresses defects in a wafer-scale processor. Sandia’s deployment announcement describes the wafer-versus-die distinction and a CS-3 research deployment; Cerebras’ chip page describes its current chip lineup and defect-handling approach.
Rank #2
- Powered by the NVIDIA Blackwell architecture and DLSS 4
- Powered by GeForce RTX 5070 Ti
- Integrated with 16GB GDDR7 256bit memory interface
- PCIe 5.0
- WINDFORCE cooling system
How a WSE differs from a GPU system
A conventional GPU is a packaged processor built from a die cut from a wafer. In large AI systems, several GPUs may share a workload: a model can be divided across processors, and those processors must exchange data. Cerebras’ approach puts many cores, SRAM and the communication fabric on one wafer-scale processor. The practical distinction is not simply chip size; it is also where memory sits and how a system organizes computation and communication.
| Comparison | Cerebras WSE-3 | NVIDIA H100 |
|---|---|---|
| Processor area | 46,225 mm², according to Cerebras’ 2024 filing | 814 mm², according to Cerebras’ 2024 filing |
| Memory location and amount | 44 GB on-chip SRAM, as reported by Cerebras in 2024 | 0.05 GB on-chip memory, as reported by Cerebras in 2024; H100 also uses off-chip high-bandwidth memory (HBM) |
| Memory bandwidth | 21 PB/s, as reported by Cerebras in 2024 | 0.003 PB/s, as reported by Cerebras in 2024 |
| Work distribution | Cerebras says a model can be kept on one WSE. For multi-WSE training, its described approach uses data parallelism, with systems working on separate training data rather than splitting the model across WSEs. | GPU systems can distribute a large model across several processors, coordinating their work. |
The figures above are a company-published comparison with the H100 specifically. Cerebras characterizes the differences as 57 times the chip area, 880 times the on-chip memory and 7,000 times the memory bandwidth. The figures describe different processor and memory arrangements; they should not be treated as a universal GPU comparison or as proof of a corresponding speed advantage. Cerebras’ filing provides the comparison and its architectural rationale.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesRank #3
- Powered by the NVIDIA Blackwell architecture and DLSS 4
- Powered by GeForce RTX 5060
- Integrated with 8GB GDDR7 128bit memory interface
- PCIe 5.0
- WINDFORCE cooling system
Why keeping compute and memory close can matter
AI processors spend time moving data as well as performing calculations. Putting a large amount of SRAM close to compute cores is intended to reduce some of that movement within the processor. A wafer-scale fabric also gives the WSE a way to connect its many cores without relying on links between separate GPU packages for every communication step.
That design goal does not by itself determine end-to-end results. A workload’s model, precision, batch size, software support, system configuration and communication needs all affect performance. GPU platforms also serve uses beyond AI, and their software ecosystems and available system configurations may be a better fit for some users. Compare the actual workload and supported software rather than relying on core counts, area or bandwidth figures alone. Cerebras’ developer documentation describes supported models and cluster details.
Rank #4
- Powered by Radeon RX 9070 XT
- WINDFORCE Cooling System
- Hawk Fan
- Server-grade Thermal Conductive Gel
- RGB Lighting
What performance claims do—and do not—show
In an August 2024 inference announcement, Cerebras reported 1,800 tokens per second for Llama 3.1 8B and 450 tokens per second for Llama 3.1 70B, describing the results as 20 times faster than NVIDIA GPU-based solutions in hyperscale clouds. The announcement also quoted Artificial Analysis benchmark results of above 1,800 output tokens per second on the 8B model and above 446 on the 70B model. These are dated results for named models and a particular service comparison, not current service guarantees or evidence that WSE systems outperform GPUs on other models or configurations. Cerebras’ inference launch announcement contains the claim and its attribution.
The figures in Cerebras’ filing and announcements are company claims; the Artificial Analysis result is quoted by Cerebras. A fair comparison of speed, latency, power use or cost needs measurements for the same model, precision, batch size, software versions and system configuration. The evidence cited here does not establish a universal ranking between WSE and GPU systems.
Best Value
- Axial-tech fans now feature a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
- Phase-change GPU thermal pad helps ensure optimal heat transfer, lowering GPU temperatures for enhanced performance and reliability
- 2.5-slot design allows for greater build compatibility while maintaining cooling performance
- Dual-ball fan bearings last up to twice as long as standard conventional sleeve bearings designs
- 0dB technology lets you enjoy light gaming in relative silence
Where WSE systems are used
Cerebras presents WSE-3 and CS-3 for AI training and inference. Sandia National Laboratories announced a CS-3 cluster deployment for research into large AI models and potential modeling and simulation work. That is an example of intended and investigated use, not evidence that every scientific workload will benefit. Sandia’s Justin Newcomer, senior manager of its ASC program, said the system could support development of large-scale trusted AI models on secure internal Tri-lab data while addressing memory and power challenges faced by GPU systems. Sandia’s announcement provides the deployment context.
Cerebras also offers an inference service powered by CS-3/WSE-3. Its 2024 launch announcement described an API compatible with the OpenAI Chat Completions API; service features and pricing can change. Separately, Cerebras describes an AWS disaggregated inference arrangement in which Trainium handles prefill and CS-3 handles decode, connected through AWS networking and made available via Amazon Bedrock. That is a deployment account from Cerebras, not an independent evaluation of its performance. Cerebras’ account of disaggregated inference explains the arrangement.
How to decide whether a WSE or GPU system fits
For a real deployment decision, compare complete systems on the work you need to run. Chip specifications alone do not establish system speed, power requirements, availability or total cost.
Quick Recap
- Workload and model: Confirm that the system supports the models, operations and precision you need.
- Software: Check framework and model support, the process for adapting or deploying your code, and any cluster requirements in the vendor documentation.
- Performance: Seek results for the same model, precision, batch size and measurement method, using clearly identified software versions and system configurations.
- System constraints: Compare the complete system’s power, facility, networking and capacity requirements—not only processor-level specifications.
- Access and cost: Compare the system or service you can actually obtain, including its availability and current pricing.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Recommended Free Tools

