NVL72

NVL72

NVL72 is a complete server rack from Nvidia in which 72 graphics processors are so tightly connected that they behave like a single giant computer. It was built to run very large AI models and is regarded as a benchmark for modern data centers.

NVL72 is a fully assembled server rack from the company Nvidia. A server rack is a person-height metal frame into which computer parts are slid. Inside this rack sit 72 particularly powerful computing chips, so-called graphics processors or GPUs for short. These chips were originally intended for computer games, because they can perform a great many simple calculations simultaneously. That is exactly what artificial intelligence also needs. What is special about the NVL72 is not the number of chips, but the way they are connected: they are coupled so tightly that programs can address them as a single, very large processor. The number 72 in the name simply stands for the number of GPUs installed.

Why an entire rack counts as one chip

Large language models have long since stopped fitting into the memory of a single chip. They must therefore be broken into pieces and distributed across many chips. This creates a problem: the chips must constantly exchange intermediate results. If this connection is slow, the expensive processors spend most of their time waiting for data. The benefit of additional chips then largely evaporates.

The NVL72 solves exactly this problem. All 72 chips are connected to a shared, extremely fast interconnect network. To the software, this looks like a processor with enormous memory. Nvidia therefore speaks of a compute rack that behaves like a single GPU. That is marketing, but it captures the technical reality fairly well.

Economically, the thing is important for another reason as well. Nvidia no longer sells just individual chips, but entire racks priced in the millions. Anyone reading about AI investments in the news is often, indirectly, reading about such systems. Operators like Microsoft, Amazon, or Meta order them in large quantities.

What’s inside the rack

Inside the NVL72 sit 72 GPUs of the Blackwell generation together with 36 main processors of the Grace type. Main processors handle control tasks and everything that doesn’t run massively in parallel. Two GPUs each share one such control chip. They are installed in flat trays that are stacked one above the other in the rack.

The connection is called NVLink. You can picture it as an extremely wide data highway between the chips. Data flows between all the chips in the rack at a rate on the order of several terabytes per second. For comparison: a very fast home internet connection manages about one thousandth of a gigabyte per second. The difference thus amounts to several million times.

Such a rack consumes around 120 kilowatts of electricity. That corresponds to the consumption of roughly thirty single-family homes. Air cooling is no longer sufficient for this, so cooling fluid flows directly through the trays. Many older data centers cannot install such racks at all, because they lack the power supply and cooling circuit needed. This is one of the reasons why new data centers are being built everywhere right now.

NVL72 in quarterly earnings and chat windows

You will almost never see an NVL72 directly. The racks stand in secured data centers, often in halls with hundreds of identical units. But if you use an AI chatbot, your request very likely ends up on such hardware. When training new models as well, many of these racks are connected together into a single cluster.

The name regularly appears in business news. Analysts estimate how many NVL72 systems Nvidia ships per quarter and derive revenue figures from that. A common mistake here is comparing these systems to ordinary servers. An NVL72 costs roughly as much as a condominium and replaces a small server hall.

Successors have already been announced, such as variants with newer chips under names like GB300 NVL72. The basic idea remains the same: connect as many chips as possible as tightly as possible. The term NVL72 therefore now stands less for a single product than for a design approach of modern AI data centers.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.