
FlashMLA
FlashMLA is a method developed by DeepSeek that executes a specific computational step in large language models — the so-called attention — significantly faster and with less memory usage. It is specifically optimized for Nvidia's H800 GPUs and was released as an open-source project in early 2025.
When a language model processes a text, it has to calculate for each word which other words in the sentence are relevant. This step is called attention. It is one of the most computationally intensive parts of all. FlashMLA is a technique from the Chinese AI company DeepSeek that speeds up exactly this step. Specifically, it concerns a variant called Multi-head Latent Attention, or MLA for short, in which the model stores information about the previous conversation history in a compressed cache. FlashMLA makes this compressed computation as efficient as possible on modern graphics processors.
The memory problem with long texts
The longer a conversation with an AI model becomes, the more context the model has to keep track of simultaneously. To do this, it creates a cache known as the KV cache — “KV” stands for “Key” and “Value,” two data tables that attention works with. This cache grows with every new message and quickly consumes large amounts of GPU memory.
Multi-head Latent Attention solves this problem by heavily compressing the cache. Instead of storing all the tables at full size, MLA condenses them into a much smaller dataset. This saves memory — but at the same time creates a new computational challenge, because the data has to be unpacked again before the actual computation can take place. This is exactly where FlashMLA comes in: it makes this unpacking and computing so fast that the memory advantage is not bought at the cost of lost time.
How FlashMLA pushes GPU hardware to its limits
Modern graphics processors have two types of memory: a small, extremely fast buffer right on the chip and a large, slower memory next to it. The bottleneck in many AI computations is not the computing power itself, but the constant shuffling of data back and forth between these two memory types. This problem is called “memory-bandwidth-bound” — the bandwidth, meaning how much data can move per second, is the limiting factor.
FlashMLA is written in such a way that it loads data in large blocks into the fast on-chip buffer and performs as many computational steps as possible there at once before writing back to the slower memory. This strategy is called “tiling.” In addition, FlashMLA uses a special feature of the Nvidia H800 GPU called WGMMA and a fast asynchronous memory mechanism called TMA — both hardware features that the chip explicitly provides for exactly such operations. On an H800 GPU, FlashMLA thereby achieves up to 580 teraflops according to DeepSeek, meaning 580 trillion computing operations per second — a large portion of the chip’s theoretically possible maximum performance.
It is important to note the difference from the older FlashAttention method, which also accelerates attention but works with the classic, uncompressed KV cache. FlashMLA, on the other hand, is specifically tailored to the compressed MLA cache. Both techniques follow a similar philosophy but are intended for different model architectures.
FlashMLA in practice and in the news
FlashMLA is a core component of DeepSeek’s own model DeepSeek-V3 and the DeepSeek-R1 model based on it. These models caused a stir in early 2025 because, despite significantly cheaper hardware, they were able to keep pace with the leading models from the US. FlashMLA was one of the technical building blocks that made this possible.
In February 2025, DeepSeek released FlashMLA as an open-source project on the platform GitHub. Within a few days, the repository was starred thousands of times — a sign of how great the interest was within the developer community. Since then, other companies and researchers have been free to use and adapt the code.
In tech news, FlashMLA usually comes up in connection with the question of how much a powerful AI really has to cost. With optimizations like this, DeepSeek has shown that it is possible to achieve a lot with less expensive hardware — which has increased pressure on Western providers and reignited the discussion about AI export restrictions.