SEO

Googlebot

Also called Google crawler, Google spider

Google's web crawler, which fetches, renders and follows links on pages so they can be added to the Google Search index.

Quick facts: Googlebot

Category
SEO
Also called
Google crawler, Google spider
Level
Beginner
Affects
Discovery, crawling, rendering and indexing of every page
Where to see it
Search Console URL Inspection and Crawl stats, server logs, robots.txt report, reverse DNS lookup
In this article4
  1. How Googlebot works
  2. Why it matters
  3. Common mistakes
  4. How to act on it

Googlebot is the web crawler Google uses to find and fetch pages for its search index: an automated program that requests your pages, reads them and follows their links to discover more. If Googlebot cannot reach a page, or sees something different from what visitors see, that page cannot rank properly in Google Search.

How Googlebot works

There are two main versions, Googlebot Smartphone and Googlebot Desktop. Since Google moved to mobile-first indexing, almost every site is crawled mainly by the smartphone version, so the mobile version of your page is the one that counts.

Crawling starts from URLs Google already knows about, from links on other pages and from sitemaps you submit. For each URL, Googlebot checks your robots.txt file first, then requests the page and records the status code and the HTML. Pages are then queued for rendering in an up-to-date version of Chromium, which runs the JavaScript and produces the version Google indexes. Links found on the rendered page join the queue.

Googlebot decides for itself how fast and how often to crawl. It slows down if your server responds slowly or returns errors, and visits important, frequently changing pages more often than stable ones. It ignores the crawl-delay directive in robots.txt. Google also runs other crawlers with their own names, such as Googlebot-Image and Googlebot-Video, plus special-purpose fetchers for products like ads and feeds.

Why it matters

Everything else in SEO depends on Googlebot getting through. A UK online shop with thousands of filter combinations can waste much of its crawl budget on near-duplicate URLs, so new products take weeks to appear. A site whose main content loads only after a script runs may show Googlebot an empty shell. A firewall or security plugin that treats Googlebot as an attacker can quietly block a whole site.

It also matters for security and data. Many bots claim to be Googlebot in their user agent string to slip past defences or scrape content. Treating every request labelled “Googlebot” as genuine can let scrapers through, while blocking real Googlebot by mistake can take a site out of search.

Common mistakes

  • Blocking CSS and JavaScript. Old robots.txt rules that block theme or script folders stop Googlebot rendering the page as visitors see it.
  • Using robots.txt to remove pages. Disallowing a URL stops crawling but not necessarily indexing. Use noindex on a crawlable page instead.
  • Serving Googlebot different content. Showing it a special version is cloaking, which breaks Google’s spam policies.
  • Trusting the user agent alone. Verify suspected Googlebot traffic before allowing or blocking it.
  • Ignoring server errors. Repeated 5xx responses make Googlebot crawl less, slowing down every update.

How to act on it

Use the URL Inspection tool in Google Search Console on your most important pages, then click “Test live URL” and “View tested page” to see the rendered HTML and screenshot Googlebot produced. If the content is missing there, Google is not seeing it either.

Check the Crawl stats report under Settings in Search Console for response times, error rates and which file types are being fetched. To confirm real Googlebot visits in your server logs, run a reverse DNS lookup on the IP address: genuine requests resolve to googlebot.com, google.com or googleusercontent.com, and a forward lookup should return the same IP. Google also publishes its crawler IP ranges.

When Googlebot is struggling with a site, the cause is usually in templates, server configuration or JavaScript rather than in individual pages, which is why it sits at the centre of technical SEO work.

Do and do not

Do

  • Check rendered pages with URL Inspection
  • Verify suspicious Googlebot traffic with reverse DNS
  • Keep server errors and slow responses to a minimum

Do not

  • Block CSS or JavaScript folders in robots.txt
  • Show Googlebot different content from visitors
  • Rely on robots.txt to keep pages out of the index

Questions people ask about this

How often does Googlebot crawl my site?

There is no fixed schedule. Googlebot visits popular, frequently updated pages more often and stable or little-linked pages less often, and it adjusts to how quickly your server responds. The Crawl stats report in Google Search Console shows how many requests it has made each day over the last 90 days.

How can I tell if a visit is really from Googlebot?

Run a reverse DNS lookup on the IP address in your server log. A real Googlebot request resolves to a hostname ending in googlebot.com, google.com or googleusercontent.com, and a forward lookup on that hostname returns the original IP. You can also compare the IP with the ranges Google publishes.

Can I ask Googlebot to crawl a page now?

Yes, within limits. In Google Search Console, inspect the URL and click Request indexing, which adds it to the crawl queue, though there is a daily quota and no guarantee of timing. For many pages at once, submit or update your XML sitemap and make sure the pages are linked from elsewhere on the site.

Related terms

Found this useful?

Share it, or ask an AI to summarise it

Back to the glossary

Knowing the term is the easy part

Applying it to your own site and budget is the work. Book a call and I will tell you what actually applies to you.