An XML sitemap is a file that lists the URLs on your website you want search engines to find and index, written in a structured format machines can read. It can also show when each page last changed, helping crawlers decide what to revisit.
How an XML sitemap works
The file is a list of entries, each holding a page address in a loc tag and, optionally, a lastmod date. Search engines download it, add the URLs to their crawling queue and use it alongside the links they find on your pages. Three facts shape how it is used in practice:
- Size limits. One sitemap can hold up to 50,000 URLs or 50MB uncompressed. Larger sites split their URLs across several files and list them in a sitemap index.
- Dates. Google uses lastmod when it proves consistently accurate, and ignores it on sites where every date changes every day regardless of edits.
- Priority and change frequency. These old optional tags are ignored by Google, so there is no benefit in tuning them.
Extensions add details for images, video and news articles. Search engines learn where the sitemap is from a Sitemap line in robots.txt and from direct submission in Google Search Console and Bing Webmaster Tools. Google retired its old sitemap ping address in 2023, so submission and robots.txt are the routes that count.
WordPress has produced a basic sitemap at /wp-sitemap.xml since version 5.5. Most SEO plugins switch that off and generate their own, often at /sitemap_index.xml.
Why it matters
A sitemap is a hint, not an instruction. Google may crawl a listed URL and still decide not to index it. But for some sites it makes a real difference to how quickly pages are found:
- new sites with few links pointing to them, where crawlers have little else to follow;
- large sites, such as online shops with thousands of product pages several clicks from the homepage;
- sites that add or change content often, where accurate lastmod dates steer recrawling.
It is also a diagnostic tool. In Search Console, the Page indexing report can be filtered to show only URLs from your submitted sitemaps, so you can see which pages you consider important that Google has not indexed, and why.
Common mistakes
- Listing URLs that should not be indexed Redirects, error pages, noindexed pages, or URLs whose canonical tag points elsewhere. Each one tells Google the sitemap cannot be trusted.
- Fake lastmod dates Set to the moment the file is generated rather than when the content changed.
- Two competing sitemaps For example the WordPress core one and a plugin’s, both live and both submitted.
- Old sitemaps left in place after a migration Still listing the previous URL structure.
- Staging or development URLs leaking into the live file.
- Expecting a sitemap to force indexing of thin or duplicate pages.
How to act on it
Find your sitemap: try /sitemap.xml, /sitemap_index.xml and /wp-sitemap.xml, and check robots.txt for a Sitemap line. Open it and confirm it lists the pages you expect, with no staging addresses.
Then test the contents. Most crawlers can crawl a list of URLs taken from a sitemap; every one should return a 200 status, be indexable and carry a self-referencing canonical. Remove anything that fails. Submit the sitemap in Search Console and Bing Webmaster Tools, check that the Sitemaps report shows it was read successfully, and filter the Page indexing report by sitemap every month or so.
On a new site, submit the sitemap as soon as the site is live and open to search engines, not while it is still blocked. Getting sitemaps clean and keeping them that way is part of the technical SEO work I do, and one of the first checks in any site launch.
