Disallow is the rule in a robots.txt file that tells search engine crawlers not to request certain URLs on your site. Its counterpart, Allow, makes an exception to a Disallow rule, so you can block a folder while still permitting one path inside it.
How Disallow works
Robots.txt sits at the root of a domain, for example example.co.uk/robots.txt. It is made of groups: each group names one or more crawlers in a User-agent line, followed by the rules for them. A simple group for an online shop might contain:
- User-agent: * applies the group to every crawler without its own group.
- Disallow: /basket/ blocks the basket and everything under it.
- Disallow: /search blocks internal search results.
- Allow: /search/help/ reopens one path inside the blocked area.
Each rule is matched against the start of the URL path, so Disallow: /search blocks /search, /search/ and /search-results alike. Google and Bing support two wildcards: an asterisk matches any run of characters, and a dollar sign marks the end of a URL, so Disallow: /*.pdf$ blocks PDF files. When an Allow and a Disallow rule both match, Google follows the most specific one, meaning the longest matching path, and if they tie, Allow wins. An empty Disallow line blocks nothing; Disallow: / blocks everything.
A crawler that respects robots.txt simply does not fetch disallowed URLs. That is all the rule does. It is a request about crawling, not an instruction about indexing, and badly behaved bots ignore it altogether.
Why it matters
Used well, Disallow keeps crawlers away from URLs that waste their time: filtered and sorted category pages on a shop, internal search results, basket and account pages, and endless calendar or parameter combinations. On a large site this protects crawl budget for the pages that matter.
Used badly, it does real damage. One stray line can block a whole site. It also cannot hide anything: if other pages link to a disallowed URL, Google can index the address without its content and show it with no description. The file is public too, so listing private folders in it advertises them.
Disallow is also the usual way to opt out of AI crawlers, by naming their user agents, such as GPTBot. Whether to do that is a business decision about AI visibility as much as a technical one.
Common mistakes
- Disallowing pages you want removed from Google. Google cannot read a noindex tag on a page it may not crawl. Use noindex and leave the page crawlable.
- Blocking CSS and JavaScript. Google needs these files to render pages, and blocking them can make pages look broken to it.
- Carrying a staging file to the live site. Disallow: / is sensible on a development server and disastrous after launch.
- Forgetting that paths are case-sensitive. /Blog/ and /blog/ are different paths.
- Treating robots.txt as security. Protect private areas with a login, not a Disallow line.
How to act on it
Open your own robots.txt in a browser and read every line. For each Disallow rule, note why it exists; a rule nobody can explain deserves testing. Google Search Console’s robots.txt report shows the version Google last fetched and any parsing problems, and the URL Inspection tool tells you whether a particular URL is blocked.
Before changing the file, test the new rules against a list of URLs you want crawled and a list you want blocked; most SEO crawlers can apply a custom robots.txt for this. After any launch or migration, check robots.txt on the first day. Reviewing crawl rules is a standard step in my technical SEO service, and for page-level control it is worth comparing Disallow with the meta robots tag.
