Vergleichsskizze: links eine runde Siliziumscheibe, die in viele kleine rechteckige Chips zersägt wird, rechts dieselbe Scheibe als ein einziger quadratischer Riesenchip, daneben zum Größenvergleich ein handelsüblicher Prozessor.

Wafer Scale Engine

The Wafer Scale Engine is a computer chip made by the company Cerebras that is as large as an entire dinner plate. Instead of sawing a round silicon disc into hundreds of small chips, it remains a single giant piece – which makes it especially fast at training AI models.

Computer chips are made from round discs of silicon called wafers. Normally, at the end of the process, such a disc is sawed into hundreds of small rectangular chips. The company Cerebras does exactly the opposite. It leaves the disc almost entirely intact and sells it as a single, enormous chip. This component is called the Wafer Scale Engine, and it is about as big as a dinner plate. For comparison: a typical processor from a data center is barely bigger than a matchbox.

Why one plate is better than 80 postcards

Modern AI systems don’t compute on a single chip. Training a large language model often involves thousands of chips working simultaneously. These chips must constantly exchange intermediate results. This exchange is precisely the real bottleneck.

Moving data within a chip is extremely fast and uses little power. Sending data from one chip to the next, on the other hand, costs many times more time and energy. Think of it like a conversation: two people in the same room talk to each other effortlessly. If they’re sitting in different buildings, every question requires a phone call.

The Wafer Scale Engine drives this overhead down by handling, on a single piece of silicon, a task that would otherwise require dozens of chips. The current generation carries around four trillion transistors and roughly 900,000 small compute cores. A powerful graphics chip from a data center has a few tens of thousands of cores. So the difference in size isn’t a minor detail — it’s the whole point.

The trick with the broken spots

Manufacturing chips always produces defects. Dust particles and tiny irregularities render individual spots on the wafer unusable. For ordinary chips, this isn’t a big deal: you simply discard the affected small chips and sell the rest. For a chip that occupies the entire disc, a single defect would destroy the whole product.

That’s why Cerebras builds in redundancy from the start. The chip contains more compute cores than it needs. If a core fails, it is disabled, and the connections are rerouted around it. This works similarly to a detour in road traffic: the route is closed off, but traffic keeps flowing anyway.

A second problem is heat. A chip this large consumes several tens of kilowatts, roughly as much as a few dozen electric kettles. That’s why the Wafer Scale Engine sits inside its own enclosure with liquid cooling. The memory for the model data sits partly directly on the chip, and partly in external units connected via very wide lines.

Cerebras versus Nvidia in the headlines

In business news, the term almost always appears in the same context: as a counter-design to Nvidia’s market power. Nvidia supplies the graphics chips on which most of the AI world computes. Cerebras is one of the few challengers with a genuinely different technical idea. The company’s IPO and major contracts from the Middle East were therefore covered extensively.

In practice, you can only buy such a system if you are a research institution, a corporation, or a data center operator. The price runs into the millions. Anyone who still wants to try out the technology rents computing time through Cerebras’s cloud services. Some chatbots and coding assistants respond there noticeably fast because the models run on wafer-scale hardware.

A common misconception is that the Wafer Scale Engine is simply an especially fast processor for everything. That’s not true. It is tailored to a specific kind of computation, the kind found in neural networks. For ordinary programs on an office computer it would be unsuitable — and with its liquid cooling and plate-sized form factor, impractical anyway.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.