Crawl budget is the number of URLs a search engine is able and willing to fetch from your site within a given period. For Google, it is the combination of how much crawling your server can handle and how much Google wants to crawl.
How crawl budget works
Google describes crawl budget as the combination of two things:
- Crawl capacity limit. How many connections Googlebot will open at once and how fast it will request pages without overloading the server. If responses are quick and error-free, the limit rises; if the server slows down or returns errors, it falls. This is closely tied to crawl rate.
- Crawl demand. How much Google wants to crawl, based on how popular your URLs are, how often they change and how many new URLs it discovers.
Every URL Googlebot fetches uses part of that allowance, whether it is a valuable product page, a redirect, a 404 or the thousandth filtered version of the same category. Crawling is a prerequisite for indexing, so if the budget is spent on low-value URLs, important pages are discovered or refreshed later than they should be.
Why it matters
For most small business sites, it does not matter much. Google’s own documentation says crawl budget is mainly a concern for very large sites, or for sites with many thousands of pages that change daily. A 60-page site for a solicitors’ firm in Manchester will be crawled comfortably however it is set up.
It becomes a real issue on sites that generate URLs faster than Google wants to fetch them. A UK fashion retailer with filters for size, colour, brand, price and fit can produce hundreds of thousands of combinations from a few thousand products. A property portal, a jobs board or a large news archive can do the same. On these sites, new products may take weeks to appear in search and price changes may not be picked up, because Googlebot is busy elsewhere.
One possible symptom is a long list of “Discovered, currently not indexed” URLs in Search Console’s Page indexing report, meaning Google knows the pages exist but has not yet fetched them. That status has other causes too, such as low perceived quality, so read it alongside the Crawl stats report rather than on its own.
Common mistakes
- Worrying about it on a small site. If you have a few hundred pages and they are indexed, the problem lies elsewhere.
- Leaving faceted navigation fully crawlable. Every filter combination becomes a URL for Googlebot to try.
- Using noindex to save budget. A noindexed page still has to be crawled for the tag to be seen. Blocking in robots.txt stops the crawl; noindex does not.
- Long redirect chains and soft errors. Each hop and each broken page is a request that delivers nothing.
- Slow server responses. A sluggish server lowers the crawl capacity limit for the whole site.
How to act on it
Start by establishing whether you have a problem. Compare the number of pages you want indexed with the number Google has indexed, and look at the Crawl stats report under Settings in Search Console to see how many requests Googlebot makes each day and what it is fetching. On larger sites, log file analysis shows exactly which URLs Googlebot spends its time on.
If budget is being wasted, block worthless parameter and filter combinations in robots.txt, keep internal links pointing at canonical URLs, fix redirect chains, keep the XML sitemap limited to indexable pages and improve server response times. This is the kind of work I carry out as part of technical SEO for larger and ecommerce sites.
