Index bloat is when a search engine has indexed far more URLs from a website than the site has pages worth finding in search. The extra URLs are typically filtered, duplicated, empty or automatically generated pages that give searchers nothing new.
How index bloat works
Most websites generate far more URLs than their owners realise. Common sources on UK sites include:
- Faceted navigation on online shops, where each combination of size, colour, brand and price creates a new URL.
- Sorting and tracking parameters, such as ?sort=price or session IDs, creating copies of the same page.
- Internal site search results pages that are linked or crawlable.
- WordPress tag archives, date archives, author archives on single-author sites and, on older installations, image attachment pages.
- Location pages produced from one template for every town in a county, differing only in the place name.
- Staging or development copies of the site left open to crawling.
- Old pages kept “just in case”: past events, discontinued products with no replacement, expired offers.
Search engines find these URLs through internal links, sitemaps and links from elsewhere. If nothing tells them otherwise, they may crawl and index them alongside the pages you actually care about.
Why it matters
A bloated index causes three problems. First, crawling effort goes on worthless URLs, so new and updated pages can take longer to be picked up; on large shops this becomes a genuine crawl budget issue. Second, near-identical pages compete with each other for the same searches, and Google may show the wrong one. Third, a site where most indexed pages are thin can look lower in quality overall, and Google assesses sites as a whole as well as page by page.
For a small brochure site of 30 pages, bloat is usually minor. For a Shopify or WooCommerce store with a few hundred products and several filters, the combinations can easily run into tens of thousands of crawlable URLs.
Common mistakes
- Blocking bloated URLs in robots.txt while they are still indexed. Blocking stops crawling, so Google never sees a noindex tag on them and they can stay in the index.
- Adding noindex and a robots.txt block at the same time, for the same reason.
- Deleting hundreds of pages at once without checking whether any earn links or traffic.
- Relying on a site: search count, which is a rough estimate and not a reliable measure of what is indexed.
- Cleaning up the existing URLs but leaving the template that keeps creating new ones.
How to act on it
Compare two numbers: how many pages you want in Google, and how many Google says it has indexed in the Page indexing report in Search Console. A large gap is the clue. Crawl the site with an SEO crawler and group URLs by pattern to see which templates create the excess.
Then deal with each pattern on its merits. URLs that should not exist get a 404 or 410 status. Duplicates get a canonical tag pointing to the main version, or a redirect. Pages useful to visitors but not to searchers, such as filters and internal search, get noindex, and crawl controls only once they have dropped out. Thin pages worth keeping get improved; the rest get merged into stronger pages. Finally, stop the source, for example by switching off tag archives or limiting which filter combinations create indexable URLs. Diagnosing and clearing bloat is a core part of technical SEO work, especially on ecommerce sites.
