

Cerebras · v4 · 5× · vu le 19 août 2026
The CS-4 is the fourth-generation AI Inference rack system from Cerebras Systems, announced on August 18-19, 2026. It is based on three new WSE-3-Turbo Wafer-Scale processors (5nm) and the new modular "Nexus" platform architecture, which decouples compute, power, and I/O subsystems. Cerebras positions the system as the fastest AI accelerator for Inference to date, claiming up to 30x higher Tokens per second per user compared to GPU systems and up to 10x higher throughput per watt compared to the predecessor CS-3. No specific end-customer price for the hardware is disclosed; CS-4 systems are distributed through individual enterprise contracts.
Fonctionnalités
| Manufacturing Process (nm) | 5 nm (TSMC), WSE-3 Turbo, 4 trillion transistors, 900,000 AI cores |
| License | Proprietary hardware/enterprise contract; also usable as a cloud inference service via API |
| Platform | Nexus Platform Architecture (modular: Compute/Backpack, Power, I/O) |
| Price | No public unit price; sold via custom enterprise contracts |
| Compute Performance (FLOPS/TOPS) | 750 PFLOPS AI compute (system, 3x WSE-3 Turbo at 250 PFLOPS each) |
| Release Date | Announced August 18, 2026; first shipments in Q3 2026 |
| Memory | 129.6 PB/s aggregate memory bandwidth (system); 44 GB on-wafer SRAM per WSE-3 Turbo |
| Availability | Early access; general availability starting this quarter (Q3 2026) |