Off-Peak Pricing

Off-Peak Pricing

Off-peak pricing refers to rates that apply outside peak hours and are therefore lower. In the AI industry, providers offer such discounts when customers shift compute jobs to hours of low utilization.

Off-peak literally means “outside the peak.” It refers to the time when little is going on. Off-peak pricing, then, is pricing that applies precisely then and is lower than usual. The principle is familiar from rail travel: a ticket for Tuesday morning costs less than one for Friday afternoon. The provider thereby rewards everyone who is flexible and can adapt. With providers of computing power for artificial intelligence, it works the same way. Anyone who runs their compute jobs at night or on weekends pays noticeably less.

Why data centers sell cheaper at night

A data center is a huge hall full of specialized chips. These chips cost a lot of money, whether they’re working or standing idle. For the operator, an unused chip is therefore pure loss. That’s why they’d rather sell the time cheaply than not at all.

Demand fluctuates sharply throughout the day. During the day, millions of people type questions into chatbots and wait for instant answers. At night, this rush collapses because most users are asleep. A provider who wants to fill the night must make it attractive. That is exactly what off-peak pricing accomplishes.

For customers, this is more than a nice discount. Running large AI models is one of the biggest cost blocks for many tech companies. When part of the work runs at half the rate, it changes the entire calculation. That’s why off-peak pricing now regularly appears in quarterly results and analyst reports.

The deal: discount in exchange for waiting time

At the core of every off-peak offer is a trade. The customer gives up control over the timing and gets a lower price in return. So they no longer tell the provider “right now, immediately,” but rather “sometime in the next few hours.” The provider then slots the job into a gap where capacity is free.

Technically, the request lands in a queue, a kind of digital waiting line. A management program sorts all pending jobs and starts them as soon as chips become available. Such jobs are called batch processing: many requests are collected and processed in bundles instead of individually and immediately. Typical commitments promise processing within 24 hours.

The discounts involved are not pocket change. With several major providers, batch processing costs only half the normal price. Some models additionally tie the discount to fixed time windows, for instance at night between half past midnight and half past eight. A common misconception is that answer quality suffers as a result. The model computes identically, just later.

Who has computing done at night

Typical candidates are tasks that nobody is waiting on. An online shop generates tens of thousands of product descriptions overnight. A media company analyzes its archive and automatically assigns tags. A bank checks a batch of contracts for risky clauses. In all these cases, all that matters is that the result is ready by morning.

Not suitable is anything where a human is sitting in front of a screen. A customer service chatbot must not stay silent for twelve hours. Nor can an assistant that helps with coding in real time. Many companies therefore run a dual track, separating urgent tasks from deferrable ones.

In the news, the term also appears outside of AI. Electricity providers have used off-peak pricing for decades so that heat pumps and electric cars charge at night. Both worlds are now growing together: data centers are deliberately shifting work to hours with cheap electricity or especially abundant wind and solar power. When analysts discuss the energy costs of AI, this shift is a central argument.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.