NVFP4

NVFP4

NVFP4 is a number format developed by NVIDIA that stores weights and activations in AI models using just 4 bits, drastically boosting compute speed and energy efficiency. It was introduced with the Blackwell GPU generation and is specifically designed for the demands of modern large language models.

At its core, an AI model consists of millions or billions of numerical values called weights. The more bits used per number, the more precise it is — but the more memory and compute time it also requires. NVFP4 is a number format from NVIDIA that represents each of these values using only 4 bits. For comparison: older standard formats use 16 or even 32 bits per number. With NVFP4, four to eight times as many values fit into the same memory — and the graphics card's compute units can correspondingly perform more operations per second. The format was introduced with NVIDIA’s Blackwell chip generation in 2025.

Why 4 bits make such a difference

Modern language models like GPT-4 or Llama have hundreds of billions of weights. Even if each number is shortened by just a few bits, this adds up to many gigabytes that no longer need to be transferred or stored. The bottleneck during inference — that is, when actually computing an answer — is often not the compute power itself, but the speed at which data flows between memory and the compute unit. Narrower formats like NVFP4 directly ease this exact bottleneck.

NVIDIA states that Blackwell GPUs with NVFP4 can perform up to four times as many calculations per second as with the previously standard 16-bit format FP16. This is not a cosmetic improvement, but a factor that determines whether a model runs on a single GPU or whether many of them need to be pooled together. For companies, this directly translates into lower costs per request.

How NVFP4 preserves precision despite fewer bits

4 bits sounds like very little — and it is. With 4 bits, only 16 different numerical values can be represented. That’s not enough to precisely encode arbitrary decimal numbers. NVFP4 is therefore what’s called a floating-point format: instead of using a fixed scale, it splits the 4 bits into sign, exponent, and mantissa — similar to scientific notation, where one says “3.1 × 10²” instead of “310”. This makes it possible to represent very small and very large values using the same number of bits.

Because 16 levels are still coarse, NVFP4 requires what’s called scaling: groups of weights are given a shared factor, which is stored in a somewhat wider format. This factor stretches or compresses the 16 levels so that they adequately cover the actual value range of the group. The result is a trade-off: minimally less accuracy in exchange for massively more speed. In practice, the loss of quality is barely noticeable in well-trained models.

An important difference from older quantization methods — that is, older techniques for reducing bit width after the fact — is that NVFP4 is meant to be factored in already during training. NVIDIA recommends training models directly with this format in mind, or at least fine-tuning them for it. This increases the upfront effort but significantly improves the quality of the compressed version.

NVFP4 in products and news

NVFP4 debuted with NVIDIA’s Blackwell architecture, specifically with the B100 and B200 GPUs as well as the GeForce RTX 5000 series. These chips contain specialized compute units called Tensor Cores, which are explicitly designed for 4-bit operations. Without this hardware support, the format would provide no speed advantage.

In the tech press, NVFP4 mainly appears in reports about AI infrastructure and operating costs. Anyone offering large language models as a cloud service — for example via platforms like Amazon Web Services or Google Cloud — is interested in NVFP4, because every answer a user receives costs money. Faster computation on the same hardware means more answers per second and lower cost per request.

For end users, the format is invisible. Someone using ChatGPT or another service doesn’t choose a number format. But the effect is noticeable: answers arrive faster, servers can serve more users simultaneously, and providers can offer cheaper plans. NVFP4 is therefore less a user-facing feature than an infrastructure decision — one that nonetheless determines how fast and how affordable AI services will become in the coming years.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.