The meta robots tag is a line of HTML in a page’s head that gives search engines instructions about that page: whether to include it in results, whether to follow its links and how much of it to show in a snippet. It works page by page, which makes it more precise than robots.txt.
How the meta robots tag works
A typical tag looks like <meta name="robots" content="noindex, follow">. The name “robots” addresses all search engines; you can address one crawler, such as name="googlebot", when you need different rules for Google. If there is no tag, the default is to index the page and follow its links.
The directives you will use most are:
- noindex Keep this page out of search results. This is the one most sites need, covered in more depth under noindex.
- nofollow Do not follow any links on this page. Rarely needed site-wide; use link-level attributes instead.
- none Shorthand for noindex and nofollow together.
- nosnippet and max-snippet: stop or limit the text shown in results, including in AI features that quote the page.
- max-image-preview Control the size of image previews, often set to “large” for publishers who want bigger thumbnails in Discover.
- unavailable_after Drop a page from results after a given date, useful for time-limited offers or events.
The crawler has to fetch the page to read the tag. If robots.txt blocks the URL, Google never sees the noindex and may still list the address based on links pointing to it. For PDFs and other non-HTML files, the same directives go in an HTTP header called the X-Robots-Tag.
Why it matters
One misplaced directive can remove a whole site from Google. The most common case I see is the WordPress setting “Discourage search engines from indexing this site”, which adds a noindex to every page. It is meant for builds in progress and is easy to leave switched on at launch.
Used properly, the tag keeps low-value pages out of the index: basket and checkout pages, internal search results, thank-you pages, filtered views and thin tag archives. That helps Google concentrate on the pages you want people to find.
Common mistakes
- Launching with noindex left on. Staging settings carried over to the live site are a regular cause of “why are we not on Google?”
- Blocking in robots.txt and noindexing at once. The block stops Google reading the noindex, so the page may linger in results.
- Noindexing pages that are canonicalised elsewhere. Mixed signals confuse Google. Pick one approach per page.
- Conflicting tags. A theme and an SEO plugin each printing a robots tag can leave you with “index” and “noindex” on the same page; Google applies the most restrictive.
- Noindexing paginated category pages. Products linked only from page two onwards can become harder to discover.
How to act on it
Crawl the site and export the robots directives for every URL. Compare the list against the pages you want in search, and look for important pages marked noindex and junk pages without it. In Search Console, the Page indexing report shows URLs “Excluded by ‘noindex’ tag”, which is the fastest way to catch accidents.
Set rules at template level in your CMS or SEO plugin rather than page by page, so new pages inherit the right setting. After any launch or redesign, check the homepage and a few key pages with URL Inspection. Auditing indexability like this is a core part of my technical SEO work.
