Cost per Token

Cost per Token

Cost per token indicates how much money a single unit of text costs that an AI language program reads or writes. This metric is the standard yardstick for how expensive it is to operate such programs.

Programs like ChatGPT break down every text into small building blocks. Such a building block is called a token and corresponds roughly to a syllable or a short word. The word “sunshine” consists of two or three such building blocks depending on the program, and a period or a comma also counts as its own building block. The cost per token states how much money is charged for a single one of these building blocks. Providers bill both for the text you feed in and for the text the program writes back. Because a single building block is extremely cheap, prices are almost always quoted per one million tokens.

Why entire business models hinge on this number

Anyone building an app with an AI feature typically buys computing power from a large provider. Payment is not made per month and not per user, but according to tokens consumed. Every single response therefore generates costs. A start-up running a free app with millions of users can be brought down by this if usage grows faster than revenue.

That is why this metric constantly appears in business news. When a provider halves its price per million tokens, that is news worth reporting. Between 2023 and 2025, these prices fell by more than a factor of ten for comparable performance. Applications that were previously unaffordable suddenly became possible as a result.

For providers, the price is also a weapon in competition. Whoever offers it cheaper wins customers, but earns less per request. Some providers even sell temporarily below their own computing costs in order to secure market share. This is reminiscent of streaming services that, in their early phase, spend more on content than they take in from subscriptions.

What drives the price up or down

Behind every token lies real computational work on specialized graphics chips. The larger the model, the more computing steps are needed per token. On top of that come electricity, cooling, and the purchase of the chips. The provider allocates these costs across the number of tokens sold and adds a margin.

One important detail: input and output cost different amounts. The model can process input text in one go. Output text, by contrast, is generated token by token, and each new building block requires its own computation pass. That is why output is typically three to five times as expensive as input.

The price is lowered above all through technology. Smaller models for simple tasks, more coarsely stored numbers within the model, and better chips reduce the cost per token. Another trick is called caching: text that is repeatedly sent along is temporarily stored and billed more cheaply the second time around. One common misconception, by the way, is that a lower token price is automatically cheaper. Some models think through long intermediate steps and thereby consume a multiple of the tokens.

Where you encounter token-based billing

You encounter it most directly on the pricing pages of AI providers. There you find tables with amounts like “$0.50 per million input tokens.” Anyone who programs themselves and uses such an interface sees at the end of the month exactly how many tokens were consumed. Many providers even display usage live in an overview window.

Indirectly, you encounter this metric in almost every AI product. When a chatbot in the free version only gives short answers or has a daily limit, a cost calculation is often behind it. Even the question of whether a company offers an AI feature at all hinges on this number.

In analyst reports, the token price is now watched in a manner similar to how the oil price is watched in industry. It is regarded as the base price for machine text work. It should be distinguished from training costs: those are incurred once when the model is built and run into the millions. Cost per token, by contrast, concerns ongoing operations, and that is paid anew every single day.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.