SEO

Disallow

Also called Disallow directive, Allow directive

The robots.txt rule that asks crawlers not to request certain URLs; Allow makes exceptions. It controls crawling, not indexing.

Quick facts: Disallow

Category
SEO
Also called
Disallow directive, Allow directive
Level
Intermediate
Affects
Crawling, crawl budget, rendering, AI crawler access
Where to see it
Your robots.txt file, Google Search Console robots.txt report, URL Inspection tool, SEO crawlers with custom robots.txt testing
In this article4
  1. How Disallow works
  2. Why it matters
  3. Common mistakes
  4. How to act on it

Disallow is the rule in a robots.txt file that tells search engine crawlers not to request certain URLs on your site. Its counterpart, Allow, makes an exception to a Disallow rule, so you can block a folder while still permitting one path inside it.

How Disallow works

Robots.txt sits at the root of a domain, for example example.co.uk/robots.txt. It is made of groups: each group names one or more crawlers in a User-agent line, followed by the rules for them. A simple group for an online shop might contain:

  • User-agent: * applies the group to every crawler without its own group.
  • Disallow: /basket/ blocks the basket and everything under it.
  • Disallow: /search blocks internal search results.
  • Allow: /search/help/ reopens one path inside the blocked area.

Each rule is matched against the start of the URL path, so Disallow: /search blocks /search, /search/ and /search-results alike. Google and Bing support two wildcards: an asterisk matches any run of characters, and a dollar sign marks the end of a URL, so Disallow: /*.pdf$ blocks PDF files. When an Allow and a Disallow rule both match, Google follows the most specific one, meaning the longest matching path, and if they tie, Allow wins. An empty Disallow line blocks nothing; Disallow: / blocks everything.

A crawler that respects robots.txt simply does not fetch disallowed URLs. That is all the rule does. It is a request about crawling, not an instruction about indexing, and badly behaved bots ignore it altogether.

Why it matters

Used well, Disallow keeps crawlers away from URLs that waste their time: filtered and sorted category pages on a shop, internal search results, basket and account pages, and endless calendar or parameter combinations. On a large site this protects crawl budget for the pages that matter.

Used badly, it does real damage. One stray line can block a whole site. It also cannot hide anything: if other pages link to a disallowed URL, Google can index the address without its content and show it with no description. The file is public too, so listing private folders in it advertises them.

Disallow is also the usual way to opt out of AI crawlers, by naming their user agents, such as GPTBot. Whether to do that is a business decision about AI visibility as much as a technical one.

Common mistakes

  • Disallowing pages you want removed from Google. Google cannot read a noindex tag on a page it may not crawl. Use noindex and leave the page crawlable.
  • Blocking CSS and JavaScript. Google needs these files to render pages, and blocking them can make pages look broken to it.
  • Carrying a staging file to the live site. Disallow: / is sensible on a development server and disastrous after launch.
  • Forgetting that paths are case-sensitive. /Blog/ and /blog/ are different paths.
  • Treating robots.txt as security. Protect private areas with a login, not a Disallow line.

How to act on it

Open your own robots.txt in a browser and read every line. For each Disallow rule, note why it exists; a rule nobody can explain deserves testing. Google Search Console’s robots.txt report shows the version Google last fetched and any parsing problems, and the URL Inspection tool tells you whether a particular URL is blocked.

Before changing the file, test the new rules against a list of URLs you want crawled and a list you want blocked; most SEO crawlers can apply a custom robots.txt for this. After any launch or migration, check robots.txt on the first day. Reviewing crawl rules is a standard step in my technical SEO service, and for page-level control it is worth comparing Disallow with the meta robots tag.

Do and do not

Do

  • Document why each Disallow rule exists
  • Test changes against URLs you want crawled and blocked
  • Check robots.txt on launch day

Do not

  • Use Disallow to remove pages from Google
  • Block CSS or JavaScript files
  • List private folders as a form of security

Questions people ask about this

Does Disallow stop a page appearing in Google?

Not reliably. Disallow stops Google crawling the page, but if other pages link to it, Google can still index the URL and show it in results, usually without a description. To keep a page out of results, allow crawling and use a noindex tag, or remove the page with a 404 or 410.

What is the difference between Disallow and noindex?

Disallow controls crawling: it asks bots not to fetch a URL. Noindex controls indexing: it tells search engines not to show a page they have fetched. They work against each other if combined, because a disallowed page cannot be crawled, so its noindex is never seen.

Should I use Disallow to block AI crawlers?

It depends on what you want from AI tools. At the time of writing (October 2026), several AI companies run separate crawlers for model training and for live search answers, each with its own user agent. Blocking training crawlers keeps your content out of future training data, while blocking the search crawlers can remove you from AI answers that might otherwise cite you.

Related terms

Found this useful?

Share it, or ask an AI to summarise it

Back to the glossary

Knowing the term is the easy part

Applying it to your own site and budget is the work. Book a call and I will tell you what actually applies to you.