SEO

Crawl Budget

Also called crawl allocation

The number of URLs a search engine is able and willing to crawl on a site in a given period, set by server capacity and demand.

Quick facts: Crawl Budget

Category
SEO
Also called
crawl allocation
Level
Advanced
Affects
Indexing speed, discovery of new pages, freshness of indexed content on large sites
Where to see it
Google Search Console Crawl stats and Page indexing reports, server log files, a site crawler such as Screaming Frog
In this article4
  1. How crawl budget works
  2. Why it matters
  3. Common mistakes
  4. How to act on it

Crawl budget is the number of URLs a search engine is able and willing to fetch from your site within a given period. For Google, it is the combination of how much crawling your server can handle and how much Google wants to crawl.

How crawl budget works

Google describes crawl budget as the combination of two things:

  • Crawl capacity limit. How many connections Googlebot will open at once and how fast it will request pages without overloading the server. If responses are quick and error-free, the limit rises; if the server slows down or returns errors, it falls. This is closely tied to crawl rate.
  • Crawl demand. How much Google wants to crawl, based on how popular your URLs are, how often they change and how many new URLs it discovers.

Every URL Googlebot fetches uses part of that allowance, whether it is a valuable product page, a redirect, a 404 or the thousandth filtered version of the same category. Crawling is a prerequisite for indexing, so if the budget is spent on low-value URLs, important pages are discovered or refreshed later than they should be.

Why it matters

For most small business sites, it does not matter much. Google’s own documentation says crawl budget is mainly a concern for very large sites, or for sites with many thousands of pages that change daily. A 60-page site for a solicitors’ firm in Manchester will be crawled comfortably however it is set up.

It becomes a real issue on sites that generate URLs faster than Google wants to fetch them. A UK fashion retailer with filters for size, colour, brand, price and fit can produce hundreds of thousands of combinations from a few thousand products. A property portal, a jobs board or a large news archive can do the same. On these sites, new products may take weeks to appear in search and price changes may not be picked up, because Googlebot is busy elsewhere.

One possible symptom is a long list of “Discovered, currently not indexed” URLs in Search Console’s Page indexing report, meaning Google knows the pages exist but has not yet fetched them. That status has other causes too, such as low perceived quality, so read it alongside the Crawl stats report rather than on its own.

Common mistakes

  • Worrying about it on a small site. If you have a few hundred pages and they are indexed, the problem lies elsewhere.
  • Leaving faceted navigation fully crawlable. Every filter combination becomes a URL for Googlebot to try.
  • Using noindex to save budget. A noindexed page still has to be crawled for the tag to be seen. Blocking in robots.txt stops the crawl; noindex does not.
  • Long redirect chains and soft errors. Each hop and each broken page is a request that delivers nothing.
  • Slow server responses. A sluggish server lowers the crawl capacity limit for the whole site.

How to act on it

Start by establishing whether you have a problem. Compare the number of pages you want indexed with the number Google has indexed, and look at the Crawl stats report under Settings in Search Console to see how many requests Googlebot makes each day and what it is fetching. On larger sites, log file analysis shows exactly which URLs Googlebot spends its time on.

If budget is being wasted, block worthless parameter and filter combinations in robots.txt, keep internal links pointing at canonical URLs, fix redirect chains, keep the XML sitemap limited to indexable pages and improve server response times. This is the kind of work I carry out as part of technical SEO for larger and ecommerce sites.

Do and do not

Do

  • Check Crawl stats before assuming a problem
  • Block worthless filter and parameter URLs in robots.txt
  • Keep the XML sitemap to indexable pages only

Do not

  • Rely on noindex to save crawl budget
  • Leave every filter combination crawlable
  • Spend time on crawl budget for a small site

Questions people ask about this

Does my small business website need to worry about crawl budget?

Usually not. Google says crawl budget is mainly a concern for very large sites or sites with many thousands of frequently changing pages. If your important pages are indexed and new ones appear within a few days, crawl budget is not your bottleneck.

Does blocking pages in robots.txt save crawl budget?

Yes. A URL disallowed in robots.txt is not fetched, so the request can go elsewhere. Bear in mind that blocked pages can still be indexed without their content if other pages link to them, so robots.txt is for controlling crawling, not for removing pages from search.

Can I ask Google to crawl my site more?

Not directly. Crawl demand rises when pages are popular, linked to and updated, and capacity rises when your server responds quickly and reliably. You can request indexing of individual URLs in Search Console, but that does not raise the overall budget.

Related terms

Found this useful?

Share it, or ask an AI to summarise it

Back to the glossary

Knowing the term is the easy part

Applying it to your own site and budget is the work. Book a call and I will tell you what actually applies to you.