Robots.txt is a plain text file at the root of a website that tells search engine crawlers and other bots which parts of the site they may and may not fetch. It lives at an address such as example.co.uk/robots.txt, and a well-behaved crawler reads it before requesting anything else.
How robots.txt works
The file is made of groups of rules. Each group starts with a User-agent line naming the crawler it applies to (an asterisk means every crawler), followed by Disallow and Allow lines listing URL paths. A short file for a WordPress site might block the admin area for everyone, allow the single admin file that front-end features rely on, and point to the XML sitemap with a Sitemap line.
A few rules govern how the file is read:
- It applies only to the exact host and protocol it sits on. A shop on shop.example.co.uk needs its own file.
- When rules conflict, Google follows the most specific match (the longest path), and where two are equally specific, the less restrictive one.
- Paths are case-sensitive. Google and Bing support the wildcard * for any characters and $ for the end of a URL.
- Google ignores the crawl-delay line; Bing honours it.
- Google reads only the first 500 KiB of the file.
The standard was formalised as RFC 9309 in 2022. It is a request, not a lock: Googlebot and other reputable crawlers obey it, but scrapers and bad actors ignore it.
Why it matters
Robots.txt controls crawling, not indexing, and most problems come from confusing the two. A URL blocked by robots.txt can still appear in Google if other pages link to it, shown without a description because Google could not read it. And because Google cannot fetch a blocked page, it cannot see a noindex tag on it either. To keep a page out of the index, allow crawling and use noindex; to stop crawlers wasting time on low-value URLs, use robots.txt.
Used well, the file keeps crawlers away from internal search results, filtered and sorted URL variations and basket pages, so large sites spend Google’s attention on pages that matter. Used badly, it can remove a site from search. A single “Disallow: /” line carried over from a staging build is one of the most common reasons a newly launched UK business site gets no organic traffic at all.
It is also where site owners now set their position on AI crawlers. At the time of writing (October 2026), user agents such as GPTBot, ClaudeBot and Google-Extended can each be allowed or blocked. Google-Extended governs whether Google may use your content for its Gemini models; it does not affect crawling or ranking in Google Search.
Common mistakes
- Blocking the whole site after launch because the staging rules were copied across.
- Using Disallow to remove pages from Google Which leaves them indexed but unreadable.
- Blocking CSS and JavaScript files So Google cannot render pages properly.
- Listing private areas to hide them. The file is public, so it advertises exactly where they are. Protect anything private with a login.
- Letting the file return a server error. If robots.txt returns a 5xx error, Google may pause crawling the site; a 404 is treated as no restrictions.
How to act on it
Open yoursite.co.uk/robots.txt in a browser and read it line by line. For every Disallow rule, ask what it blocks and why. Check the robots.txt report in Search Console (under Settings) to see the version Google last fetched and any parsing problems, and use the URL Inspection tool to confirm important pages are not blocked. Keep the file short; most small business sites need only a few lines.
Make reading robots.txt the last check before and the first check after any launch or migration. If the rules have grown complicated, or you are not sure what a line does, reviewing them is a standard part of my technical SEO service.
