Schichtdiagramm des KI-Stacks: unten die Hardware-Schicht mit Chips und Stromversorgung, darüber die Infrastruktur- und Cloud-Schicht mit Rechenzentren, darüber die Modell-Schicht mit trainierten KI-Modellen und Schnittstellen, ganz oben die Anwendungsschicht mit Chatbot und Apps; Pfeile zeigen den Weg einer Nutzeranfrage von oben nach unten und der Antwort zurück.

AI Stack

The AI stack is the totality of layers needed for an AI product to work: from the chips through the data centers and the models to the app you use yourself. The term helps to structure the market, because different companies make money in each layer.

When you use an AI app on your phone, many layers work together behind the scenes. At the very bottom are special computing chips that handle the many calculation steps. Above that are large data centers, meaning halls full of such computers, which can be rented over the internet. Above that lies the actual program that has learned from example data to generate text or images. And right at the top is the app with the input field that you see. This structure of stacked layers is called the AI stack, from the English word for a pile.

Why the stack structures the market

The term is primarily an organizing framework. In news about AI companies, the question of who actually profits from the boom keeps coming up. The answer depends on which layer a company occupies. Chip manufacturers like Nvidia sell to everyone else and earn money regardless of which app ends up being successful.

Whoever sits further up depends on the layers below. A start-up building a learning app pays for every request made to the model it uses. If those prices rise, its own profit shrinks immediately. That’s why some companies try to occupy several layers themselves. Google, for example, develops its own chips, operates its own data centers, trains its own models, and builds its own apps.

For investors and journalists, the stack is therefore a standard tool. Phrases like “value creation is shifting upward” simply mean: the money is moving from the chips to the applications. Whether that’s true is disputed, but the debate is conducted in exactly these terms.

The four layers from bottom to top

The bottom layer is the hardware. This includes graphics processing units, or GPUs for short, i.e. chips that carry out very many simple calculations simultaneously. That’s exactly what AI needs. Added to this are storage, network cables, and the power supply. A modern AI data center consumes as much electricity as a small city.

Above that lies the infrastructure layer. Companies like Amazon, Microsoft, or Google rent out computing power by the hour. This is called the cloud. A start-up thus doesn’t have to build its own hall but books computing time like a streaming service. This also includes programs that monitor training and manage data.

The third layer is the models themselves, such as GPT, Gemini, or Llama. They are trained once, at great expense, and then used millions of times. Many providers make them available via an interface, a so-called API. Other programs can address the model through this without owning it. Right at the top is the application layer: the chatbot, the coding assistant, the translation feature in the browser. Only this layer is seen by the user.

The stack in headlines and products

Whenever a company presents quarterly results, it gets assigned to a layer. Nvidia is considered hardware, OpenAI a model provider, Notion or Duolingo an application. Reports about export restrictions on chips affect the bottom layer, but have upward effects on all the others.

A common misconception is to take the AI stack for a fixed technical standard. It is only a mental model. Some depictions count five layers, others three, and the boundaries shift. Nor is the order always clean: a model provider can simultaneously operate its own app and thereby become a competitor to its own customers.

In everyday life, the stack is invisible to you. When you type a question into a chatbot, it travels through the app, via an interface to the model, from there into a data center and onto a chip. The answer takes the same path back, usually within one to two seconds. Anyone who knows the stack better understands why such answers cost money and why some providers set limits on free usage.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.