SEO

Crawling

Also called crawl, spidering

The process by which search engine bots discover and fetch pages on the web, the first step before a page can be indexed and ranked.

Quick facts: Crawling

Category
SEO
Also called
crawl, spidering
Level
Beginner
Affects
Discovery of pages, indexing, how quickly changes appear in search
Where to see it
Google Search Console URL Inspection and Crawl stats, a site crawler such as Screaming Frog, server logs, robots.txt tester
In this article4
  1. How crawling works
  2. Why it matters
  3. Common mistakes
  4. How to act on it

Crawling is the process by which search engines send automated programs, called bots or spiders, to discover pages on the web and download their content. It is the first step in search: a page that has not been crawled cannot be indexed, and a page that is not indexed cannot rank.

How crawling works

Google’s crawler, Googlebot, keeps a vast list of known URLs and a queue of ones to fetch. The cycle runs roughly like this:

  1. Discovery. New URLs are found mainly by following links from pages already known, and from XML sitemaps submitted by site owners.
  2. Permission check. Before fetching, the bot reads the site’s robots.txt file to see which paths it may visit.
  3. Fetching. The bot requests the page and records the server’s response: the HTML, the status code (200, 301, 404 and so on) and headers.
  4. Rendering. For pages that rely on JavaScript, Google runs the scripts in a headless version of Chrome to see the finished page. Rendering can happen some time after the first fetch.
  5. Extraction. Links found on the page join the queue, and the content is passed on for indexing.

Since the move to mobile-first indexing, Google crawls sites primarily with its smartphone crawler, so what a phone visitor would see is what counts. Bing, AI companies and SEO tools run their own crawlers in much the same way.

Crawling is ongoing. Pages are revisited at intervals that depend on how important they seem and how often they change, so a busy homepage may be fetched daily and an old blog post every few weeks.

Why it matters

Every SEO effort assumes the page has been crawled. If a new service page for a firm of chartered surveyors in Leeds is not linked from anywhere on the site and is missing from the sitemap, Googlebot may never find it. If it sits behind a login, a form or a robots.txt block, the bot cannot read it. However good the content is, it will not appear.

Crawling problems are also common after a redesign. A developer may leave a staging-site robots.txt rule in place, move content into JavaScript that renders slowly, or change URLs without redirects. The first sign is often a steady fall in indexed pages a few weeks after launch.

Common mistakes

  • Confusing crawling with indexing. Blocking a page in robots.txt stops crawling, but the URL can still be indexed from links elsewhere. To keep a page out of results, allow crawling and use noindex.
  • Orphan pages. Pages with no internal links pointing to them are hard for bots to find and send no internal signal that they matter.
  • Links that bots cannot follow. Navigation built with JavaScript click events instead of real anchor links with href attributes may not be followed.
  • Blocking CSS and JavaScript. If Google cannot load the files that build the page, it may not see the content or layout properly.

How to act on it

Check that your robots.txt blocks only what it should, that your XML sitemap lists the pages you want found, and that every important page is linked from somewhere sensible within a few clicks of the homepage. Use the URL Inspection tool in Search Console to see when Google last crawled a page and how it rendered it, and the Crawl stats report to see overall activity.

A crawl of your own site with a tool such as Screaming Frog mimics what a bot does and shows broken links, redirects and pages it could not reach. Fixing what stops search engines crawling a site properly is the foundation of my technical SEO service.

Do and do not

Do

  • Link every important page from relevant pages on the site
  • Keep the XML sitemap limited to indexable URLs
  • Check robots.txt after every launch or redesign

Do not

  • Use robots.txt to keep pages out of search results
  • Build navigation without real href links
  • Block the CSS and JavaScript Google needs to render pages

Questions people ask about this

What is the difference between crawling and indexing?

Crawling is fetching the page; indexing is analysing it and storing it so it can appear in results. A page can be crawled but not indexed, for example if it has a noindex tag or Google judges it a duplicate. It can even be indexed without being crawled, if robots.txt blocks it but other sites link to it.

How often does Google crawl my website?

It varies by page. Google revisits pages it considers important or frequently updated more often, and older, rarely changing pages less often. The Crawl stats report in Search Console shows how many requests Googlebot makes to your site each day, and URL Inspection shows the last crawl date for a single page.

How do I get Google to crawl a new page?

Link to it from relevant existing pages, add it to your XML sitemap, and request indexing through the URL Inspection tool in Search Console. Internal links from well-visited pages are the most reliable long-term signal. Requesting indexing speeds things up but does not guarantee the page will be indexed.

Related terms

Found this useful?

Share it, or ask an AI to summarise it

Back to the glossary

Knowing the term is the easy part

Applying it to your own site and budget is the work. Book a call and I will tell you what actually applies to you.