Noindex is an instruction that tells search engines not to include a page in their search results. You add it either as a robots meta tag in the page’s HTML head or as an X-Robots-Tag in the HTTP response header, and once a search engine has crawled the page and seen it, the page drops out of the index.
How noindex works
The two forms do the same job:
- Meta tag. A line in the page’s head with name="robots" and content="noindex". You can address a single crawler by replacing “robots” with its name, such as “googlebot”. The meta robots tag can carry other directives alongside it.
- HTTP header. An X-Robots-Tag header with the value noindex, sent by the server. This is the only option for files with no HTML head, such as PDFs and images.
The crucial point is that a search engine has to crawl the page to read the instruction. If the same URL is blocked in robots.txt, Googlebot never fetches it, never sees the noindex, and may keep the address in results based on links pointing to it, usually shown without a description. The two do different jobs: robots.txt controls crawling, noindex controls indexing.
Removal is not instant. The page goes after the next crawl, which can take days or weeks for less important URLs. For something urgent, Search Console’s Removals tool hides a URL for about six months while the noindex takes effect. Google stopped honouring noindex rules written inside robots.txt in 2019, so that old shortcut no longer works.
Why it matters
Used deliberately, noindex keeps low-value pages out of search: thank-you pages after a form, internal search results, filtered views of a category, staging copies, login screens and thin tag archives. That keeps the index focused on pages that deserve to rank and reduces index bloat.
Used by accident, it is one of the most damaging single lines in SEO. WordPress has a setting under Settings, Reading called “Discourage search engines from indexing this site”, which adds noindex to every page. It is ticked while a site is being built and is often forgotten at launch, which is one of the commonest reasons a new UK business website gets no organic traffic for months. Some site builders and staging tools add a similar site-wide tag that survives the move to the live domain.
Common mistakes
- Going live with the development setting still on. View the source of the live homepage on launch day and search for “noindex”.
- Combining noindex with a robots.txt block. The block hides the instruction. Let the page be crawled until it has dropped out, then block it if you need to.
- Using noindex for near-duplicates. Where two pages are close copies, a canonical tag pointing at the preferred version consolidates signals. Noindex simply discards the duplicate, and putting both on one page sends mixed messages.
- Noindexing pages that lead to products. If category or pagination pages are the only route to deeper content, a long-term noindex can mean Google visits them less and finds those products less often.
- Listing noindexed URLs in the XML sitemap. A sitemap should contain only pages you want indexed.
How to act on it
Open Search Console’s Page indexing report and look for the reason “Excluded by ‘noindex’ tag”. Every URL listed there should be one you meant to exclude. If an important page appears, find where the tag comes from: the page settings in your SEO plugin, a theme template, or a header added by the server or a CDN. The URL Inspection tool shows exactly what Googlebot saw on its last visit.
Then crawl the site and list every noindexed URL alongside its traffic, links and place in the structure. Decide page by page whether to keep it excluded, canonicalise it, improve it, or remove it with a redirect. Checking the indexing set-up of a whole site, including launch-day checks like this, is a standard part of my technical SEO work.
