
RDMA
RDMA (Remote Direct Memory Access) is a networking technique that allows one computer to write to or read from another computer's memory directly – without requiring any intervention from its processor. This makes data transfer in data centers and AI clusters extremely fast and efficient.
When two computers exchange data over a network, the process usually works like this: the sending machine transmits a packet, the receiver interrupts its work, its processor accepts the data, stores it, and confirms receipt. This costs time – and above all, computing capacity. RDMA, short for “Remote Direct Memory Access,” takes a different approach. One computer writes directly into the memory of another machine, without that machine’s processor having to become active at all. The receiver, so to speak, only notices once the data has already arrived. The principle is comparable to placing a book on someone’s desk while they’re busy with something else – they don’t need to open the door or accept anything. This is made possible by specially adapted network cards that handle the data transfer independently.
RDMA as the foundation of modern AI clusters
Why does this matter so much? AI models such as the large language models behind chatbots aren’t trained on a single chip. Instead, dozens or hundreds of specialized computing chips – so-called GPUs – work together simultaneously. They must constantly exchange intermediate results with one another. For models with billions of parameters, i.e. the numerical values that make up the model, this happens thousands of times per second.
If every one of these transfers kept the processor busy, it would be permanently occupied managing data packets instead of accelerating the actual training. RDMA offloads this task entirely to the network card. The processor remains free. Latency – the delay between sending and receiving – drops to a fraction of a millisecond. For AI training, this isn’t a nice-to-have, it’s a prerequisite: without RDMA, the chips would spend much of their time waiting for the next round of data to arrive.
How RDMA bypasses the processor
The decisive component is the so-called RDMA-capable network card (also called a NIC, for Network Interface Card). It knows the memory addresses of both machines involved and is permitted to access them directly – a procedure known as DMA, or “Direct Memory Access.” In ordinary DMA, this happens within a single computer, for example when a hard drive writes data into RAM without burdening the processor. RDMA extends this principle to two machines connected via a network.
For this to work, both sides must first register and release a shared memory region. The network cards then coordinate the transfer independently. In practice, there are two widely used technologies for this: InfiniBand, a special high-speed network for data centers, and RoCE (RDMA over Converged Ethernet), which implements the same principle over conventional Ethernet cables. Both achieve transfer rates of several hundred gigabits per second – that is, tens of thousands of times faster than a typical home connection.
A common misconception: RDMA doesn’t make the network itself faster; rather, it eliminates the detour through the processor. The line speed stays the same; what changes is how efficiently that speed is utilized.
RDMA in products and headlines
Anyone reading about Nvidia GPUs, Meta's AI clusters, or the Microsoft data center that runs ChatGPT is implicitly reading about RDMA as well. Nvidia’s NVLink and the associated InfiniBand networks rely entirely on RDMA to couple chips in a cluster so tightly that they behave as if they were sitting on a single circuit board. Cloud providers such as Amazon Web Services and Google Cloud also offer RDMA-capable instances that customers can rent for demanding AI training.
Outside of AI, RDMA is primarily found in high-performance databases and the financial sector. There, microsecond-level latency is critical – for instance in automated stock trading, where a fraction of a millisecond can determine profit or loss. Supercomputers used in climate research or physics simulation have also relied on RDMA for decades; however, the AI boom has given the technology a new level of recognition and significantly grown the market for the corresponding hardware.