Crawling is the process by which search engines send automated programs, called bots or spiders, to discover pages on the web and download their content. It is the first step in search: a page that has not been crawled cannot be indexed, and a page that is not indexed cannot rank.
How crawling works
Google’s crawler, Googlebot, keeps a vast list of known URLs and a queue of ones to fetch. The cycle runs roughly like this:
- Discovery. New URLs are found mainly by following links from pages already known, and from XML sitemaps submitted by site owners.
- Permission check. Before fetching, the bot reads the site’s robots.txt file to see which paths it may visit.
- Fetching. The bot requests the page and records the server’s response: the HTML, the status code (200, 301, 404 and so on) and headers.
- Rendering. For pages that rely on JavaScript, Google runs the scripts in a headless version of Chrome to see the finished page. Rendering can happen some time after the first fetch.
- Extraction. Links found on the page join the queue, and the content is passed on for indexing.
Since the move to mobile-first indexing, Google crawls sites primarily with its smartphone crawler, so what a phone visitor would see is what counts. Bing, AI companies and SEO tools run their own crawlers in much the same way.
Crawling is ongoing. Pages are revisited at intervals that depend on how important they seem and how often they change, so a busy homepage may be fetched daily and an old blog post every few weeks.
Why it matters
Every SEO effort assumes the page has been crawled. If a new service page for a firm of chartered surveyors in Leeds is not linked from anywhere on the site and is missing from the sitemap, Googlebot may never find it. If it sits behind a login, a form or a robots.txt block, the bot cannot read it. However good the content is, it will not appear.
Crawling problems are also common after a redesign. A developer may leave a staging-site robots.txt rule in place, move content into JavaScript that renders slowly, or change URLs without redirects. The first sign is often a steady fall in indexed pages a few weeks after launch.
Common mistakes
- Confusing crawling with indexing. Blocking a page in robots.txt stops crawling, but the URL can still be indexed from links elsewhere. To keep a page out of results, allow crawling and use noindex.
- Orphan pages. Pages with no internal links pointing to them are hard for bots to find and send no internal signal that they matter.
- Links that bots cannot follow. Navigation built with JavaScript click events instead of real anchor links with href attributes may not be followed.
- Blocking CSS and JavaScript. If Google cannot load the files that build the page, it may not see the content or layout properly.
How to act on it
Check that your robots.txt blocks only what it should, that your XML sitemap lists the pages you want found, and that every important page is linked from somewhere sensible within a few clicks of the homepage. Use the URL Inspection tool in Search Console to see when Google last crawled a page and how it rendered it, and the Crawl stats report to see overall activity.
A crawl of your own site with a tool such as Screaming Frog mimics what a bot does and shows broken links, redirects and pages it could not reach. Fixing what stops search engines crawling a site properly is the foundation of my technical SEO service.
