
Matrix Multiplication
Matrix multiplication is a mathematical operation in which two rectangular tables of numbers are combined with each other. It is the central computational operation behind nearly all AI models and largely determines how much computing power a model requires.
A matrix is a rectangular table of numbers — for example, four rows and four columns. In matrix multiplication, two such tables are combined according to a fixed rule, producing a new table. The rule is: each cell in the result is created by multiplying a row of the first matrix with a column of the second matrix step by step and adding up all the products. This sounds like school material — and it is. In AI, however, it is the most frequent and computationally intensive operation of all.
Matrix multiplication as the heart of AI models
A neural network — the basic principle behind language models like ChatGPT or image generators — consists of many layers. In each of these layers, incoming data is multiplied by a matrix full of learned weights. This happens hundreds or thousands of times per request.
The size of these matrices determines how powerful a model is. Large models have weight matrices with billions of entries. The sheer number of computational operations required for this is the main reason why training and running such models consumes so much electricity and hardware. According to estimates, over 90 percent of the computational work of a large language model consists of matrix multiplications.
Anyone who understands matrix multiplication thus gets a direct view into why AI is so expensive — and why researchers put so much effort into making it more efficient.
Why graphics chips are built for matrix multiplication
A normal CPU, like the one found in every laptop, processes tasks sequentially — one after another. Matrix multiplication, however, can be broken down into thousands of independent partial calculations. All of these partial calculations can take place simultaneously.
This is exactly what GPUs (graphics processing units) are built for: they have thousands of small computing units that work in parallel. What would take minutes on a CPU, a GPU accomplishes in seconds. Companies like Nvidia make their money precisely from this — their chips are, at their core, highly optimized machines for matrix multiplication. Newer specialized processors such as Google's TPU (Tensor Processing Unit) go even further and are geared exclusively toward this one operation.
A typical matrix product for a medium-sized model involves matrices with thousands of rows and columns. The number of individual multiplications involved grows with the cube of the matrix size — which explains why merely doubling the model size can multiply the computational costs many times over.
Matmul in news, products, and research
In tech media, matrix multiplication usually appears under the abbreviation “matmul.” When Nvidia unveils a new chip, the focus is almost always on one key figure: how many trillions of matrix multiplications the chip can perform per second, measured in FLOPS (Floating Point Operations Per Second).
Efficiency research, too, often revolves directly around this operation. Methods such as quantization — where the numbers in the matrices are stored less precisely but more compactly — or low-rank approximation, where large matrices are approximated by two smaller ones, all aim to speed up or shrink matrix multiplications.
In 2022, Google’s AlphaTensor caused a stir: an AI had discovered a new algorithm that solves matrix multiplication for certain sizes in fewer steps than all previously known methods. This was the first time in decades that this fundamental problem had seen any improvement — a sign of how central the operation is to the entire field.