A sitemap is a list of URLs you hand to a crawler, and it is the only file where you state what your site contains. The sitemap limits are simple to state and easy to cross, and when one is crossed the file is fetched, read, and then partly or wholly ignored. Nothing anywhere reports that back to you.
You can validate one against the protocol without an account. These are the rules a crawler is applying.
A sitemap is a hint, not an instruction
Google says this in plain terms. Submitting a sitemap is merely a hint, and it does not guarantee that Google will download it or use it for crawling URLs on the site.
So a sitemap is not a way to force indexing, and its absence is not what stops a site being crawled. What it does is tell a crawler about URLs it might not otherwise have reached. New pages, pages buried deep in a structure, and pages that few things link to are the ones it helps with.
The sitemap limits that break a file
Two hard limits are in the sitemap protocol, and both are easy to cross on a large site.
50,000 URLs per file. Past that the file does not qualify, and the fix is a second file rather than a longer one.
50MB, which is 52,428,800 bytes. The number that matters is the size after uncompression. Compressing the file with gzip saves bandwidth when it is fetched, and the protocol allows it, but the uncompressed size still has to fit inside the limit.
Three format rules catch people too. The file must be UTF-8 encoded, every data value must be entity-escaped, and it must open with a urlset element carrying the protocol namespace. An unescaped ampersand inside a URL is the classic break, because a crawlable page address often contains one.
To go beyond either limit you split the site into several sitemaps and list them in a sitemap index file. An index can only reference sitemaps on the same site as the index itself.
What a crawler reads and what it ignores
This is where most sitemaps carry dead weight, and it is worth knowing before you spend an afternoon tuning them.
Google ignores the priority value outright. It is not a ranking signal and it never was. Every page being set to 1.0, or 0.5, or anything else, changes nothing at all.
Google ignores changefreq in the same way. Telling a crawler that a page changes weekly does not make it come back weekly, and the setting does not influence how often it comes.
The modification date is read, but only under a condition. Google uses it when it is consistently and verifiably accurate, which in practice means it agrees with the page's real last modification. A script that stamps today's date on every URL every night does not make a site look fresh. It teaches the crawler that the field cannot be trusted, and the field is then ignored for everything, including the pages where it was honest.
So the useful part of a sitemap entry is the address and, if you can keep it truthful, the modification date. The rest is optional and mostly cosmetic.
The URLs that disappear without a message
Every URL in a sitemap has to use the same protocol and reside on the same host as the sitemap itself. A sitemap at a www host cannot list URLs on a subdomain, and it cannot mix http and https addresses.
The protocol is explicit about what happens next. URLs that are not considered valid are dropped from further consideration. There is no error, no report and no indication in the response that anything was discarded, so a sitemap that looks complete can be contributing a fraction of what its author believes.
That single rule explains a lot of sites whose sitemap is valid, listed and crawled, and which still have pages that never appear. The addresses were fine. They were not eligible.
Validating one properly
Enter a hostname and the tool does what a crawler does. It looks for the file, checking the Sitemap line in robots.txt first and then the usual addresses, and it tells you which of those found it, because a sitemap nobody can find is a file nobody reads.
From there it identifies whether the document is a list of URLs or an index, counts the URLs and the bytes against both limits, lists the child sitemaps when it is an index, and reports each problem it finds with what to change. Parse errors, limit breaches, entries pointing off the host, and modification dates that cannot be honest all come back as specific findings rather than a pass or fail.
It is worth running on a site you have inherited. A sitemap tends to be generated once, linked once, and then left alone for years while the site around it changes shape.
The sitemap as a list of what is live
There is one more use for a sitemap that has nothing to do with crawlers. It is the closest thing most sites have to a written statement of which URLs exist, which makes it the natural reference to check the rest of the site against.
Anything in the sitemap that no longer resolves is a stale entry. Anything live and linked that is missing from the sitemap is a page a crawler has to find some other way. Both lists are short, both are concrete, and neither needs a crawler to produce. The broken link crawler covers the other half of that audit, which is the links between pages rather than the list of them.
EnterraHost