
Trainium
Trainium is a specialized chip that Amazon developed in-house to train AI models. It is the company's own alternative to Nvidia's expensive computing chips and runs exclusively in Amazon's data centers.
Trainium is a computer chip made by Amazon. It was built for a single task: training artificial intelligence. Training here means that a program derives rules on its own from vast amounts of sample text or images. That requires weeks of computation with enormous quantities of numbers, and ordinary processors are far too slow for that. Amazon doesn’t build the chip to sell in stores, but only installs it in its own data centers. Anyone who wants to use it rents computing time there through Amazon’s cloud, AWS. The first version arrived in 2020, and since 2024 there has been a significantly more powerful second generation, Trainium2.
Amazon’s way out of dependence on Nvidia
Almost all major AI models of recent years have been computed on chips made by the company Nvidia. This gives Nvidia a market position that is otherwise rarely seen: demand exceeds supply, and prices are correspondingly high. A single Nvidia compute unit for data centers costs tens of thousands of dollars, depending on the model. Anyone who needs hundreds of thousands of them is vulnerable to pressure.
Amazon is therefore trying to cover part of this demand itself. A chip of its own lowers costs, because no middleman takes a cut. It also secures delivery, because Amazon isn’t standing in the same queue as everyone else. Google does the same thing with its TPU chips, and Microsoft and Meta are also working on their own designs.
For investors, Trainium is therefore above all a signal. Every chip that comes from Amazon’s own production is a chip that Nvidia doesn’t sell. Whether this trend truly becomes dangerous for Nvidia remains an open question. So far, the overall market is growing so fast that both sides are gaining.
Specialized instead of versatile
A normal processor in a laptop is an all-rounder. It can play music, calculate spreadsheets, and manage an operating system. Trainium can do almost none of that. The chip is designed to perform very many multiplications and additions simultaneously. That is exactly what AI training ultimately comes down to.
You can picture the difference as with vehicles. A compact car will get you anywhere, a Formula 1 car only onto the racetrack — but there it’s unbeatably fast. Trainium is the race car. It forgoes versatility and thereby gains speed per watt of electricity consumed.
A single chip is never enough for large models. Amazon therefore connects many Trainium chips into servers, and these servers in turn into large clusters. What matters here is the connection between them, since all chips must constantly exchange intermediate results. Amazon calls such large clusters UltraClusters, some with over a hundred thousand chips. There is also a sibling chip called Inferentia, which doesn’t train but lets the finished model respond during everyday operation.
Trainium in the news and in products
You will never hold a Trainium chip directly in your hand. You only notice it indirectly when a voice assistant or a chatbot responds whose model was created on Amazon hardware. The best-known customer is Anthropic, the company behind the chatbot Claude. Amazon has invested billions there and, with Project Rainier, is building a huge Trainium facility for this company.
In business news, the chips usually come up in Amazon’s quarterly earnings. Then the discussion centers on the AWS cloud business and on the question of how much money is flowing into new data centers. Nvidia’s stock price also sometimes reacts to news about new Trainium generations.
A common misconception is that Trainium is a competing product to ChatGPT or similar services. The chip is not an AI model, but merely the machine underneath it. A second misconception: better chips don’t automatically make a model smarter. They make training cheaper and faster, which in turn is what makes larger models affordable in the first place.