MTBF

MTBF

MTBF is a metric for how long a technical device runs on average before it fails. It is usually given in hours and comes from measurements taken across many devices simultaneously, not from observing a single unit.

MTBF is the abbreviation for “Mean Time Between Failures,” i.e. the average time between two failures. The metric indicates how long a device typically functions before something breaks. It is almost always given in operating hours. A hard drive with an MTBF of 1.2 million hours sounds like a lifespan of 137 years. That is a misunderstanding: the value is derived by testing a very large number of devices for a short period and extrapolating the results. It therefore describes a failure rate within a large group, not the durability of a single unit.

Why data centers watch this number

A large data center houses tens of thousands of hard drives, power supplies, and fans. With that many components, something fails practically every day. MTBF makes it possible to estimate these failures in advance. Anyone operating 10,000 drives with an MTBF of one million hours must statistically expect a failure every 100 hours. From this follows how many spare parts need to be kept in stock and how many technicians need to be on duty.

This becomes a real problem especially when training large AI models. Such computing runs occupy thousands of graphics cards for weeks at a time. If even a single card fails, the entire job often comes to a halt. Operators report interruptions occurring just hours apart. That is why they regularly save intermediate states, so as not to have to start over from scratch after a failure.

For customers, MTBF is also a selling point. Manufacturers of servers, industrial controllers, or medical devices list it in their data sheets. It thus also serves as a basis for contracts guaranteeing availability.

How the value is derived

The calculation is fundamentally simple. You add up the total operating time of all tested devices and divide it by the number of failures. If 1,000 devices run for 1,000 hours and two of them break, that yields an MTBF of 500,000 hours. Yet the test itself only took about six weeks.

This extrapolation comes with an important caveat. It only applies to the normal usage phase, during which the failure rate remains roughly constant. Right at the start, more devices fail because manufacturing defects surface. Toward the end of the lifespan, the rate rises again as materials wear out. Experts call this curve the bathtub curve because of its shape. MTBF only describes the flat bottom of that tub.

MTBF must be distinguished from two similar metrics. MTTF, Mean Time To Failure, applies to components that are not repaired but replaced. MTTR, Mean Time To Repair, on the other hand, measures how long a repair takes. Only both values together yield the actual availability of a system.

From the hard drive to the cloud guarantee

Most commonly, the figure appears in the data sheets of storage media. Server hard drives are advertised with 1 to 2.5 million hours MTBF, while simple consumer models have significantly lower values. Power supplies, fans, and industrial routers also carry such figures. Some manufacturers now prefer to state the annual failure rate as a percentage instead, since it appears less misleading.

MTBF also appears indirectly in every cloud contract. When a provider guarantees 99.99 percent availability, that promise is based on a calculation involving failure frequency and repair duration. Such commitments are called a Service Level Agreement, or SLA for short. If the commitment is not met, money is usually refunded.

In news from the chip and AI industry, the term comes up when discussing the reliability of new hardware. Operators then compare how stably different generations of graphics cards run under continuous operation. For investors, this is relevant because frequent failures noticeably increase operating costs.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.