DeepGEMM

DeepGEMM

DeepGEMM is a library developed by DeepSeek that performs matrix multiplications on AI accelerators such as GPUs extremely fast. It is a central building block for operating large language models efficiently, and was released by DeepSeek as an open-source project.

Large AI models are, at their core, computing machines. The most common computational operation they perform is matrix multiplication: two large tables of numbers are processed against each other according to fixed rules. This operation — called General Matrix Multiplication, or GEMM for short — occurs thousands of times per second when operating a language model. DeepGEMM is a software library from the Chinese AI company DeepSeek that executes exactly this computational operation on specialized hardware as fast as possible. The library was released as open source in early 2025, meaning it was made freely accessible to everyone.

Matrix multiplication as a bottleneck for AI models

Anyone operating a large language model pays for every answer it generates. The most expensive line item is the computing time on specialized processors, so-called GPUs or AI accelerators. The more efficiently matrix multiplication runs, the more requests a GPU can answer per second — and the lower the costs.

This relationship makes a library like DeepGEMM strategically significant. Major providers like Nvidia do offer their own optimized solutions for their hardware, such as the cuBLAS library. However, DeepSeek has shown that with targeted adjustments, official solutions can be significantly outperformed on certain tasks. In its own benchmarks — standardized speed tests — DeepGEMM achieved up to 1.3 times the speedup compared to comparable implementations.

How DeepGEMM pushes hardware to its limits

Modern GPUs have specialized computing units built for matrix multiplication — in Nvidia’s current models, these are called Tensor Cores. The problem: these units are only fully utilized when data is delivered in the right format and in the right order. If that is not the case, the cores sit idle waiting for input, and computing capacity is wasted unused.

DeepGEMM is designed for a particularly compact number format called FP8. FP8 stores each number in just 8 bits instead of the usual 16 or 32 bits. This sounds like a simple trick, but it is technically demanding: smaller number formats are less precise, and the software must actively compensate for this imprecision. DeepGEMM does this using so-called scaling factors, which are calculated tile by tile — that is, for small blocks of the matrix — rather than using a single value for the entire table. The result is a good compromise between speed and computational accuracy.

In addition, DeepGEMM does not write its machine code by hand, but generates it automatically at runtime — custom-tailored to the exact size of the respective matrices. This approach, called just-in-time compilation, is more demanding to develop, but delivers better results on many real-world inputs than code fixed in advance.

DeepGEMM in practice and in the news

DeepGEMM is the computational foundation of DeepSeek’s models, including DeepSeek-V3 and DeepSeek-R1. These models made headlines in early 2025 because, despite a comparatively small training budget, they were able to keep pace with significantly more expensive Western competitor models. DeepGEMM was one of the reasons for this: whoever uses available hardware more efficiently gets further with the same budget.

By releasing it as open source, researchers and companies worldwide can incorporate the library into their own projects or use it as a template for their own optimizations. This is unusual: optimizations that reach so deeply into the hardware are normally considered trade secrets. DeepSeek’s decision to disclose them has fueled the discussion about how much competitive advantage really lies in proprietary tools — and how much can arise from open collaboration.

For users of AI products, DeepGEMM is invisible. It works in the background, much like an optimized engine you don’t see but whose performance you feel — in the form of faster answers and lower prices for AI services.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.