
FP64 Emulation
FP64 emulation refers to a technique in which a computer chip mimics highly precise calculations with 64-bit floating-point numbers, even though it cannot handle them directly in hardware. It is an important cost factor when running AI models and scientific simulations.
Computers internally calculate with numbers stored as long chains of zeros and ones. The more digits such a chain has, the more precise the result. FP64 stands for “64-bit floating-point number” — a number with a great many decimal places, encoded in 64 such digits. Some chips can execute FP64 calculations directly in their circuitry. Other chips, especially those built for AI, are less well equipped for this. FP64 emulation means: the chip fakes this capability by executing several simpler computational operations in sequence, which together deliver the same result as a genuine FP64 calculation.
Precision as a scarce resource
Whether a chip computes with 64 or fewer digits sounds like a technical detail. In practice, it determines whether certain calculations can be carried out meaningfully at all. Physics simulations, weather models, and financial calculations rely on FP64 because small computational errors accumulate over many steps and would otherwise render the result useless.
AI training, on the other hand, often gets by with less precise numbers — 16 or 32 bits are usually sufficient there. That’s why manufacturers like Nvidia build their AI chips to be very fast at lower precision, but deliberately slower at FP64. Anyone who still needs FP64 on such chips must emulate it — and pays for that with speed.
How a chip mimics higher precision
A 64-bit number can be split into two 32-bit numbers that together carry the same information. The chip then adds, multiplies, or divides these two halves separately and reassembles the result correctly afterward. This is called “double-single arithmetic” or, more generally, “software emulation”.
The price for this is considerable. Instead of a single computational operation, the chip needs, depending on the method, four to eight. A chip that might lose perhaps ten percent of its theoretical maximum performance with genuine FP64 hardware can drop to a fifth or less of its speed under emulation. For scientific data centers that need thousands of such operations per second, this makes a massive difference in operating costs and runtime.
Important: the result is computationally identical to a native FP64 calculation — it is not an approximation. Emulation means “mimicking,” not “estimating.” The difference lies exclusively in speed, not in correctness.
FP64 emulation in the chip industry and in the news
The term comes up primarily when new AI chips are unveiled. Analysts then check the datasheet to see how high a chip’s native FP64 performance is — and whether it achieves that figure only through emulation. Nvidia’s H100 chip, for example, is significantly more capable at FP64 than older AI chips because it supports FP64 natively. Its successor, the B200, by contrast, primarily advertises lower precision levels, which shows that the AI training market sets different priorities than science does.
Cloud providers such as Amazon Web Services or Google Cloud sell access to various chip types. For customers from research — such as universities or pharmaceutical companies — FP64 performance is a key purchasing criterion. Anyone who accidentally books an AI-optimized chip for scientific simulations unintentionally ends up in emulation and is then puzzled by unexpectedly long computation times.
A common misconception is the assumption that FP64 emulation is a flaw or a lower-quality stopgap solution. In reality, it is a deliberate design decision: chip designers sacrifice FP64 speed in favor of more transistors for AI-specific operations. For the right use case, this makes sense — one just needs to know which chip is needed for what.