Crawl Budget

Crawl Budget

The crawl budget is the number of pages a search engine fetches from a website within a given period of time. It plays a role in determining how quickly new or changed content even appears in search at all.

Search engines like Google send out programs that roam the internet, independently calling up page after page and storing their content. These programs are called crawlers, and the process is called crawling. No provider can constantly fetch every single page in the world, because that costs computing power and puts a load on other people’s servers. That’s why every website receives a kind of quota: a limited number of fetches per time period. This exact quota is called the crawl budget. It is not a fixed number published anywhere, but a value that the search engine continuously adjusts itself.

Why large shops and news sites pay attention to it

For a private website with thirty pages, the crawl budget plays practically no role at all. The crawler manages to cover such sites completely anyway. It becomes interesting starting at several tens of thousands of URLs. A large online shop with millions of product pages can never be fetched completely, let alone every day.

The consequence is unpleasantly concrete: what isn’t crawled cannot appear in search results or be updated there. A product has long been sold out, but Google still shows the old price. A news article from this morning only shows up in the evening. For news, such delays determine whether a piece still finds readers at all.

A common misconception is that the crawl budget is a quality judgment. That’s not quite true. It only describes how often a page is fetched, not how well it is rated. However, the two are related: pages that frequently offer new content and have many links from other websites are visited more often.

What the quota is made up of

Google describes the crawl budget as the result of two factors. The first is the server’s load limit. If a website responds quickly and without errors, the crawler increases its pace. If it becomes slow or returns error messages, the crawler immediately throttles back. This is a protective measure so the crawler doesn’t effectively bring a website to a standstill.

The second factor is the search engine’s level of interest. Popular, frequently updated, and well-linked pages are checked more often than URLs that haven’t changed in years. You can picture it like a newspaper editorial team with a limited number of reporters. It sends someone where new developments are expected, not repeatedly to the same empty archive.

Budget is wasted above all through superfluous URLs. Typical examples are filter and sorting links in online shops, which turn a single product list into thousands of nearly identical variants. Added to this are redirect chains, broken links, and pages with duplicate content. Operators counter this with a file called robots.txt, which blocks certain areas from crawlers, and with a sitemap, i.e., a directory of the truly important URLs.

Where the term shows up in practice

The most direct encounter with this topic is in Google Search Console, a free tool for website operators. There is a report there with crawl statistics: it shows how many requests the crawler made per day and how quickly the server responded. From such curves, experts can tell whether a problem lies with the server or with the structure of the website.

In the professional field of search engine optimization, SEO for short, the crawl budget belongs to what’s known as technical optimization. This isn’t about text content, but about load times, linking, and server responses. Large publishers and retail companies employ their own specialists for this.

The topic has become newly relevant since providers of AI systems also began operating their own crawlers to collect training data. As a result, some websites report significantly more automated visits than before. Operators now have to decide which programs they grant access to. The dispute over this has since become a regular topic in business news.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.