---
title: "XML Sitemap Checker & Finder: Validate Any Sitemap"
description: "Find a website's XML sitemap and validate it: XML errors, URL limits, lastmod dates, duplicate and off-site URLs, sitemap index files and robots.txt references."
image: "https://www.blackseoanalyzer.com/static/images/black-seo-analyzer-og-image.png"
canonical: "https://www.blackseoanalyzer.com/en/tools/xml-sitemap-checker"
language: "en"
---

# XML Sitemap Checker

Enter a domain to find its XML sitemap, or paste a sitemap URL to validate it. The checker reads robots.txt, tries the usual sitemap locations, then checks the XML, the URLs, lastmod dates and sitemap index files against the sitemaps.org protocol and Google's rules.

## How to find a website's sitemap

Most sites put their sitemap in one of a handful of places. Check them in this order:

1. **robots.txt.** Open `https://example.com/robots.txt` and look for lines that start with `Sitemap:`. This is the most reliable place, because it's where site owners tell every crawler where the sitemap is. A file can list several.
2. **The common file names.** Try `/sitemap.xml`, `/sitemap_index.xml`, `/wp-sitemap.xml` and `/sitemap.xml.gz` at the root of the domain.
3. **Your CMS's default.** WordPress 5.5 and later generates `/wp-sitemap.xml`. Yoast SEO and Rank Math replace it with `/sitemap_index.xml`. Shopify, Wix and Squarespace serve `/sitemap.xml`.
4. **Search Console, for your own site.** The Sitemaps report in Google Search Console lists every sitemap that has been submitted for the property, with the date Google last read it and how many URLs it found.
5. **A search, as a last resort.** `site:example.com filetype:xml` sometimes turns up sitemap files that Google has indexed. It's hit or miss, because sitemaps usually aren't indexed.

The tool above does steps 1 and 2 for you: enter a domain and it reads robots.txt, tries the common locations, shows which ones returned 404, and validates the first sitemap it finds. If a sitemap exists but robots.txt doesn't mention it, it tells you.

A sitemap only shows what the site owner chose to list. To see every page a site actually links to, you need a crawl; see [how to see all pages on a website](https://www.blackseoanalyzer.com/en/blog/how-to-see-all-pages-on-a-website).

## What an XML sitemap is and when you need one

An XML sitemap is a file that lists the URLs you want search engines to crawl and index, with optional dates for when each one last changed. It's a hint, not a command: listing a URL doesn't guarantee Google will crawl or index it.

A small site whose pages are all linked from the navigation may not need one. Sitemaps matter most for large sites, new sites with few external links, sites with pages that are hard to reach through internal links, and sites that publish or change a lot of content. Most CMSs generate one automatically, so in practice the question is usually whether the generated sitemap is correct.

## Sitemap format

A regular sitemap is a `<urlset>` with one `<url>` per page. `<loc>` is required; `<lastmod>` is optional:

```
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://example.com/</loc>
    <lastmod>2026-09-01</lastmod>
  </url>
  <url>
    <loc>https://example.com/products?color=red&size=m</loc>
  </url>
</urlset>
```

A sitemap index lists other sitemaps instead of pages:

```
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <sitemap>
    <loc>https://example.com/sitemap-pages.xml</loc>
    <lastmod>2026-09-01</lastmod>
  </sitemap>
  <sitemap>
    <loc>https://example.com/sitemap-posts.xml.gz</loc>
  </sitemap>
</sitemapindex>
```

The rules the checker applies:

- The file must be UTF-8, and values must be entity-escaped. An unescaped `&` in a query string is the most common reason a sitemap isn't valid XML; write it as `&`.
- Every `<loc>` must be a full, absolute URL with the protocol and host, under 2,048 characters.
- URLs should be on the same host and protocol as the sitemap. A sitemap can list URLs for another host only if you prove you control both, for example with a `Sitemap:` line in the other host's robots.txt or by verifying both in Search Console.
- Google also accepts a plain text file with one URL per line and nothing else, and RSS or Atom feeds.

## Sitemap size limits

| Limit | Value |
| --- | --- |
| URLs per sitemap | 50,000 |
| File size per sitemap | 50 MB uncompressed |
| Sitemaps per sitemap index | 50,000 |
| Sitemap index files per site in Search Console | 500 |

Past either limit, split the URLs into several sitemaps and list them in a sitemap index. Gzip compression saves bandwidth but doesn't raise the 50 MB limit, which applies to the uncompressed file. Child sitemaps must be in the same directory as the index or below it, and Google has said it doesn't support an index that lists other indexes. See Google's guide to [managing large sitemaps](https://developers.google.com/search/docs/crawling-indexing/sitemaps/large-sitemaps).

## lastmod, priority and changefreq

`<lastmod>` must use the W3C Datetime format: `2026-09-16`, or a full timestamp with a time zone such as `2026-09-16T14:30:00+00:00`. Google uses lastmod only when it's consistently and verifiably accurate, so it should be the date the page's content last changed in a meaningful way, not the date the sitemap was generated. Two patterns tell Google the dates aren't reliable: dates in the future, and every URL sharing the same lastmod. The checker flags both.

Google ignores `<priority>` and `<changefreq>`. They don't hurt, but there's no reason to spend time tuning them.

## What not to put in a sitemap

List only the URLs you want to appear in search results. That rules out:

- **Redirects.** List the final URL, not one that 301s to it. Turn on the status check above to test the first 10 URLs.
- **Error pages.** 404, 410 and 5xx URLs waste crawls and can't be indexed.
- **Noindexed pages.** A sitemap that says "index this" and a page that says `noindex` send conflicting signals.
- **Non-canonical URLs.** Parameter variants, session IDs and duplicates should be left out in favor of the canonical version. The [canonical tag checker](https://www.blackseoanalyzer.com/en/tools/canonical-tag-checker) shows which URL a page declares.
- **URLs blocked by robots.txt.** Google can't crawl them, so listing them does nothing. Test them with the [robots.txt tester](https://www.blackseoanalyzer.com/en/tools/robots-txt-tester).
- **The http version of an https site**, or both `www` and non-`www` hosts.

## How to submit a sitemap

There are two ways that work for Google:

- **robots.txt.** Add a line anywhere in the file: `Sitemap: https://example.com/sitemap.xml`. Every crawler that reads robots.txt can then find it, not only Google.
- **Search Console.** Submit the sitemap URL in the Sitemaps report, or through the Search Console API. The report then shows whether Google could read it and how many URLs it discovered.

Google deprecated its sitemap "ping" endpoint in [June 2023](https://developers.google.com/search/blog/2023/06/sitemaps-lastmod-ping), and it stopped responding about six months later. Plugins that still ping it do nothing useful; use robots.txt and Search Console instead.

## Compare your sitemap with a crawl to find orphan pages

A sitemap and a crawl of your internal links should mostly agree. Where they don't, there's usually a problem:

- **In the sitemap but not reached by the crawl:** orphan pages with no internal links. Search engines can discover them through the sitemap, but visitors and crawlers following links can't reach them. Link to them, or remove them if they shouldn't exist.
- **Reached by the crawl but not in the sitemap:** pages the sitemap generator misses, or pages that shouldn't exist at all, like faceted or parameter URLs.

To get both lists, crawl the site from its home page, then crawl again starting from the sitemap, and compare the URLs. The [guide to seeing all pages on a website](https://www.blackseoanalyzer.com/en/blog/how-to-see-all-pages-on-a-website) walks through the method.

bseoa can start a crawl from a sitemap. Pass `--is-sitemap` (a URL ending in `.xml` is detected automatically), and it reads every URL in the sitemap, follows child sitemap URLs ending in `.xml` listed in an index, and then keeps crawling the links it finds on those pages:

```
black-seo-analyzer --url-to-begin-crawl https://example.com/sitemap.xml --is-sitemap --output-type csv --output-file sitemap-crawl.csv
```

It can also export a visual map of a crawl with `--output-type sitemap`. That's an HTML diagram of how pages link together, not an XML sitemap file for search engines.

## Check every page, not just one

These tools look at one URL at a time. bseoa crawls your whole site on your own machine and checks every page for 300+ technical SEO issues across 16 analysis modules.

[Start the 14-day free trial](https://www.blackseoanalyzer.com/en/free-trial)

## More free SEO tools

- [Title & Meta Description Length Checker](https://www.blackseoanalyzer.com/en/tools/meta-description-length-checker)
Type or paste a title and meta description and see, as you type, whether Google or Bing will cut them off, measured in pixels rather than characters.

- [Robots.txt Tester](https://www.blackseoanalyzer.com/en/tools/robots-txt-tester)
Test whether a URL is allowed or blocked for Googlebot, Bingbot, GPTBot or any user agent, and see the exact rule that decided it.

- [Schema Markup Validator](https://www.blackseoanalyzer.com/en/tools/schema-markup-validator)
Paste JSON-LD or enter a URL. Get JSON syntax errors with line numbers and a check of the properties Google requires for rich results.

- [Hreflang Checker](https://www.blackseoanalyzer.com/en/tools/hreflang-checker)
Validate a page's hreflang annotations, then fetch each alternate to confirm it links back, returns 200 and isn't canonicalized elsewhere.

- [Canonical Tag Checker](https://www.blackseoanalyzer.com/en/tools/canonical-tag-checker)
See which URL a page declares as canonical, in HTML or HTTP headers, and whether that target actually resolves to an indexable page.

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://www.blackseoanalyzer.com/#organization","name":"Fiscus Technology, LLC","url":"https://www.blackseoanalyzer.com/","logo":"https://www.blackseoanalyzer.com/static/images/black-seo-analyzer.png"},{"@type":"WebSite","@id":"https://www.blackseoanalyzer.com/#website","name":"bseoa","url":"https://www.blackseoanalyzer.com/","publisher":{"@id":"https://www.blackseoanalyzer.com/#organization"}},{"@type":"WebApplication","@id":"https://www.blackseoanalyzer.com/en/tools/xml-sitemap-checker#app","name":"XML Sitemap Checker","url":"https://www.blackseoanalyzer.com/en/tools/xml-sitemap-checker","description":"Find a site's sitemap from its domain, then validate the XML, the URLs in it, lastmod dates and sitemap index files.","applicationCategory":"DeveloperApplication","operatingSystem":"Any (web browser)","isAccessibleForFree":true,"offers":{"@type":"Offer","price":"0","priceCurrency":"USD"},"publisher":{"@id":"https://www.blackseoanalyzer.com/#organization"}},{"@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://www.blackseoanalyzer.com/"},{"@type":"ListItem","position":2,"name":"Free SEO Tools","item":"https://www.blackseoanalyzer.com/en/tools"},{"@type":"ListItem","position":3,"name":"XML Sitemap Checker","item":"https://www.blackseoanalyzer.com/en/tools/xml-sitemap-checker"}]}]}
```
