Token Arbitrage

Token arbitrage refers to exploiting price differences between AI services that bill by processed text unit. Whoever buys requests cheaply and resells them at a higher price or redirects them to a cheaper model earns the difference.

Language models like ChatGPT don’t bill text by the word, but in small text building blocks. These building blocks are called tokens, and an English word often consists of one or two of them. Providers charge a fixed price for each million of these building blocks. Prices vary widely: between the cheapest and the most expensive provider, the factor can sometimes be as high as a hundred. Token arbitrage means turning exactly this price difference into a business. You buy computing power where it costs little, and resell it where customers pay more.

Why the price range is so large

The market for AI computing power is young and confusing. There is no central exchange where a uniform price is established. Every provider calculates its own prices, depending on its data centers, its electricity costs, and its utilization. Some companies even sell below their own costs in order to gain market share.

On top of that comes a second difference: not every task needs the most expensive model. A spelling correction can be handled by a small model for a fraction of the price. Deploying a large flagship model for that is wasteful. Whoever separates these tasks lowers their costs significantly without the user noticing any difference in quality.

For investors and observers, this matters because it explains the profit margins of entire companies. Many AI startups don’t build their own model. They buy computing power from others and sell an interface wrapped around it. Their business lives and dies by this difference.

The path of a request from the customer to the cheapest provider

In the simplest case, a provider charges the customer a monthly fee but only pays behind the scenes based on actual usage. Whoever pays twenty euros a month but only makes five euros' worth of requests generates fifteen euros of profit. Heavy users, on the other hand, cost money. Such offerings only work because both groups balance each other out.

Technically more elaborate is what’s called routing. A small additional program checks every incoming request and estimates how difficult it is. Simple questions are sent to a cheap model, complicated ones to an expensive one. The customer pays a flat rate, while the actual costs fluctuate strongly depending on the request.

A third lever is caching, as it’s known in technical jargon. If many users ask the same question, the model only needs to compute the answer once. Some providers offer discounts of up to ninety percent for such reused pieces of text. Whoever structures their requests cleverly pays only a fraction for the same content.

Where this business model appears in the news

It is most visible in coding assistants and writing tools. These services often cost twenty to thirty euros a month and access models from OpenAI, Anthropic, or Google. When a major provider lowers its prices, the margins of these companies rise overnight. If it raises them, the entire business model is shaken.

A second arena is platforms that offer many models through a single interface. They advertise that they automatically select the cheapest suitable model. Their value to the customer lies exactly in managing these price differences.

A common misconception is confusing token arbitrage with cryptocurrencies. There, the term token means something completely different, namely a digital coin. Here, it is exclusively about text building blocks and their billing. And unlike classic arbitrage on the stock exchange, the advantage doesn’t disappear immediately, because prices aren’t publicly compared.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.