Skip to content
EnterraHost
Products
Tools
Hosting Email Support Client Area Find a domain
Easy utilities

Sitemap Validator

Checks a sitemap against the protocol and tells you what to change. Parse errors, limit breaches, bad dates and broken index entries, each with the fix.

Give it a domain and it checks any Sitemap line in robots.txt, then the usual addresses starting with the root. Or paste the full address of a sitemap anywhere on the site.

How the sitemap is found

The root first, then robots.txt, then the usual suspects.

Nearly every sitemap is at the site root, either /sitemap.xml or whatever your platform generates by default. That is where a crawler looks first, and it is where this tool ends up on the average site.

A Sitemap line in robots.txt is checked before that guess, because a declaration is a statement about the site and a guess is not. It is what gets a sitemap found when it lives somewhere else, such as a platform that generates it at a subpath.

With nothing declared, the usual addresses are tried in turn and the first one that returns a real document is the one reported. That starts at the root /sitemap.xml, then the WordPress variants such as /wp-sitemap.xml and /sitemap_index.xml, then a compressed or plain text sitemap. The result says which address answered and whether anything declared it.

A declared sitemap on a different host is reported rather than fetched. A sitemap there is only valid when it is verified in the same search property as this one, and following it would mean this tool requesting any address the site owner cared to name.

The failure nobody notices

A sitemap that does not parse is silent.

A sitemap is read as a single document. If one character in it is wrong, the parser stops and reports nothing at all, so every URL the file listed is lost rather than merely mis-described. The site looks like it has no sitemap, which is the opposite of the truth and points at the wrong problem.

What breaks it

An unescaped ampersand in a query string, an unclosed tag, or a character that is not valid in the declared encoding. Generating the file by hand or by string concatenation is what usually causes it, because both skip the escaping an XML writer would do.

What a fix looks like

Write & where you meant a bare ampersand, and build the file with an XML writer rather than a string. This tool reports the line number the parser objected to, which is the quickest way to find the one character that matters.

Two limits, one file each

50,000 URLs and 50 megabytes.

A single sitemap may hold up to 50,000 URLs and may be up to 50 MB uncompressed. Past either limit the file is invalid and the whole thing is discarded, not truncated. Larger sites split the URLs across several sitemaps and list those in a sitemap index.

An index has its own failure mode. Each loc in it has to resolve to a real sitemap, and one dead child takes every URL inside it out of the picture without anything reporting an error. This tool fetches the children on the same host and reads each one.

What this does not do

This validates the document, not the pages.

A sitemap can be perfectly written and still list URLs that 404, redirect or carry a noindex tag. That is a different check against real requests, and it lives in the next tool along.