Kreislaufdiagramm mit vier Stationen im Ring: mehr Nutzer, mehr Nutzungsdaten, besseres Produkt, höhere Attraktivität; Pfeile führen im Kreis zurück zum Ausgangspunkt und zeigen die Selbstverstärkung.

Data Flywheel

A data flywheel describes a self-reinforcing cycle: a product gets used, this usage generates data, that data makes the product better, and this attracts even more users. The term explains why large providers in the AI market often extend their lead rather than lose it.

A flywheel is a heavy disc that is hard to get moving. Once it’s spinning, it keeps its motion almost by itself. This is exactly the image behind the data flywheel. It refers to a cycle in four steps: people use a product, this generates data about their usage, this data makes the product better, and the better product attracts more people. Each round makes the next one easier. The advantage therefore doesn’t grow evenly, but accelerates over time.

Why the runner-up has it so hard

Whoever has many users early on collects a lot of data early on. This can be used to improve a product that the competition still lacks. The competitor would therefore have to build a better product without having the data needed to do so. This is a chicken-and-egg problem, and it explains why market leaders online often remain very stable.

For investors, this is a central argument. A company with a functioning flywheel is allowed to run losses for years, as long as the cycle keeps turning. That’s why companies often release AI products for free or very cheaply. The calculation goes: usage first, then data, then quality, and money eventually. Whether this calculation works out is the real point of contention in many valuations of AI startups.

It’s important to distinguish this from the network effect. In a network effect, a service becomes more valuable because other users are there — a messenger app without friends is useless. In a data flywheel, the service gets better because it has learned from earlier usage. Both can occur together, but they are not the same thing.

What happens in each round of the cycle

The first round is the hardest. Without users there is no data, so the product initially has to convince through its basic functionality alone. Some companies buy in data for this purpose or generate it artificially. Only after that does the actual cycle begin.

The crucial question is what data actually gets generated. In the case of a chatbot, this is not just the users' questions. What’s most valuable are signals about whether an answer was good: a user clicks the thumbs-up, asks again in annoyance, or copies the text and keeps working with it. Such feedback flows into further training rounds and sharpens the model. In navigation apps, it’s speed data from thousands of phones, from which traffic jams are calculated.

However, a flywheel only turns if the feedback is fast. If there’s a year between usage and improvement, nobody feels the effect. That’s why companies build automatic evaluation systems that continuously detect errors. And the cycle can also run in the wrong direction: if a model mainly receives its own outputs as new data, errors reinforce themselves instead of disappearing.

Where the effect shows up in real products

The best-known examples predate the current AI boom. A search engine learns from every click which result fits which search query. A video platform learns from drop-offs which recommendation was bad. An online retailer learns from orders which products are bought together.

In the news today, the term mainly appears in connection with language models and autonomous driving. Manufacturers of driver assistance systems collect data from millions of driven kilometers, especially from rare and tricky situations. Whoever has more vehicles on the road sees more such cases. This is precisely how companies like Tesla justify their claimed lead.

A common misconception is that data volume automatically brings an advantage. Millions of trivial logs are of little use. What matters is whether the data fits the specific task and whether it contains usable feedback. In addition, data protection rules set limits: not everything that can be technically collected may also be used for training.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.