Indexing is the stage where a search engine analyses a page it has crawled and stores it in its index, the vast database it draws on to build search results. A page that is not indexed cannot appear in search at all, however good it is.
How indexing works
For Google, getting a page into search happens in stages:
- Discovery. Google learns that a URL exists, through links from other pages, an XML sitemap or a manual request.
- Crawling. Googlebot fetches the page. Crawling is a separate step, and a crawled page is not yet an indexed one.
- Rendering. Google runs the page’s JavaScript, much as a browser does, to see the content a visitor would see. Rendering problems often hide content from Google.
- Indexing. Google reads the text, headings, links, images and structured data, works out what the page is about, and decides whether to store it.
- Serving. When someone searches, Google chooses from its indexed pages and ranks them.
Indexing is a decision, not an automatic step. Google groups duplicate and near-duplicate pages, chooses one canonical version to keep, and may leave the rest out. It also leaves out pages it judges to be of little value, and anything it is told not to index through a noindex directive. Search Console labels these outcomes, for example “Crawled – currently not indexed” (fetched but not stored) and “Discovered – currently not indexed” (known about but not yet fetched).
Why it matters
Every other piece of SEO depends on it. If your new service page for Croydon is not indexed, no amount of content work or link building will make it appear. Indexing problems are also easy to cause by accident. On WordPress, the “Discourage search engines from indexing this site” setting adds a noindex instruction to every page; it is meant for sites under construction and is regularly left switched on after launch. A redesign that changes URLs, a plugin that adds noindex to the wrong templates, or a robots.txt rule copied over from a staging site can each remove pages from Google.
The opposite problem exists too. Indexing pages that should never be found, such as filter combinations, tag archives and thin duplicates, leads to index bloat and pulls attention away from the pages that matter.
Common mistakes
- Assuming that publishing a page means it is indexed. Check it.
- Requesting indexing in Search Console again and again for a page Google has decided not to keep, instead of fixing the reason.
- Leaving a staging site’s noindex or robots.txt block in place when the site goes live.
- Pointing canonical tags at the wrong URL, so Google indexes a different page from the one intended.
- Expecting every page on a large site to be indexed. Google does not promise that, and some exclusions are normal.
- Confusing indexed with ranking. An indexed page can still sit on page eight.
How to act on it
Use the URL Inspection tool in Google Search Console to check any important page. It shows whether the page is indexed, which canonical Google chose and when it was last crawled. For the whole site, the Page indexing report lists indexed and excluded pages with the reason for each exclusion. Work through the reasons in order of business impact, starting with the pages that should be bringing in enquiries.
Make important pages easy to reach through internal links and an accurate XML sitemap, keep low-value URLs out of the index, and recheck robots.txt and meta robots after every launch or redesign. Diagnosing and fixing indexing problems is central to my technical SEO work.
