Ablaufskizze: Ein KI-Crawler fragt eine Website an. Ein Prüfdienst dazwischen erkennt den Crawler und antwortet ohne Bezahlvereinbarung mit dem Statuscode „Zahlung erforderlich". Bei akzeptiertem Preis wird abgerechnet und der Inhalt ausgeliefert; ein menschlicher Besucher im Browser passiert die Prüfung ohne Gebühr.

Pay-per-Crawl

Pay-per-Crawl means: whoever automatically collects text and images from a website has to pay for it. Website operators want to use this to charge AI companies that use their content to train programs.

Large parts of the internet are constantly being scanned by programs that autonomously call up page after page and store the content. Such programs are called crawlers. Until now, this scanning was almost always free of charge: a website stood open on the net, and anyone was allowed to read it, even a machine. Pay-per-Crawl reverses this. The website operator demands money for each of these automated retrievals, or blocks access if payment is not made. People who read the page normally in a browser are not affected by this.

The dispute over paid content on the net

The background is a tangible conflict. Newspapers, forums, and encyclopedias have created and paid for their own texts over the years. AI companies collect these texts and use them as training material for chatbots. The chatbot then answers the question directly, and nobody clicks through to the original page anymore.

For many media outlets, this is an existential problem. Their business model relies on people visiting their page and seeing advertising there or taking out a subscription. If visitors stay away, the revenue source disappears, even though the content continues to be used. Pay-per-Crawl is meant to turn data collection back into a business in which both sides get something.

Conversely, AI companies argue that a fee-based net would become expensive and confusing. Small providers might not be able to afford the fees. This would give the largest corporations yet another advantage, because only they can pay across the board.

Bouncer between crawler and server

Technically, detection is needed first. Every retrieval of a website leaves traces: a network address and an identifier of the program making the request. Known crawlers of major AI companies can be identified fairly reliably from this. A service positioned between the visitor and the actual server checks every request before it goes through.

If this service detects a crawler without a payment agreement, it does not deliver the content. Instead, it responds with a status code that essentially means: payment required. The crawler can then accept a price, and the amount is billed. The provider Cloudflare introduced such a system in 2025 and blocks AI crawlers by default for new customers.

It is important to distinguish this from an older method. In the robots.txt file, a website has been able to note for decades which crawlers it does not want. However, this is only a polite request without technical effect. Pay-per-Crawl, by contrast, enforces the rule, because the content simply is not delivered without payment.

What this means for users and markets

Pay-per-Crawl is barely visible directly. Anyone who visits a news site notices nothing of it. It becomes noticeable indirectly, for example when a chatbot gives an evasive answer on current topics because it is no longer allowed to access certain sources.

In business news, however, the term comes up frequently. Publishers such as the New York Times or Axel Springer are negotiating licensing agreements with AI providers, sometimes worth several million euros per year. Pay-per-Crawl is the automated version of this: instead of one large contract, many small transactions. Reddit, Stack Overflow, and Getty Images have taken similar paths.

A common misconception is that this settles all copyright questions. Pay-per-Crawl only regulates technical access, not the legal question of whether training with other people’s texts is permitted at all. This question continues to occupy courts in the US and Europe. Both processes run in parallel and influence each other.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.