Skip to content
EnterraHost
Products
Tools
Hosting Email Support Client Area Find a domain

Broken links. What the crawl found, and what to fix first

A crawl finds the broken links on a site, following what your pages point at rather than what you believe they point at. It is a scan rather than a check, which means it follows what your pages link to instead of what you believe they link to, and that difference is where the value is.

You can crawl a site and list every failure without an account. This is what the report means and which numbers deserve your afternoon.

A crawl report for example.com showing one page crawled and eleven other counters at zero, with one page not in the sitemap as the only non-zero figure.
A real crawl of example.com, captured 8 October 2026. Almost every counter is zero because the site is a single page with no links and no sitemap, and the one figure that is not zero is explained at the end of this guide.

What the crawl panel reports

Twelve counters come back, and they fall into four groups.

Reach. Pages crawled, pages skipped, and pages not reached by the crawl. These describe how much of the site the scan saw, which is the first thing to check before trusting any other number.

Links. Broken internal links, broken external links, and redirects found. This is the group most people came for.

The sitemap. Sitemap URLs, sitemap issues, and pages not in the sitemap. These compare what you publish in the sitemap against what the crawl found.

Presentation. Broken images and missing alt text, which are the failures a reader notices before a crawler does.

One counter sits apart from the rest. Possibly blocked is deliberately separate from the broken counters, because a page that answered 403 or 429 is not a dead link. It is a page that refused to be read, and treating that as broken sends you looking for a bug that is not there.

Broken links are a reader problem more than a crawl problem

The instinct is that broken links damage a site's standing in search. The measured position is narrower and more useful than that.

Google's own description of how status codes affect crawling says that 4xx codes, with the exception of 429, have no effect on crawl rate. A 404 does not cost you anything in crawling terms. It costs you the reader who clicked it, which is the argument for fixing it and also the reason it is not an emergency.

What does cost you is a 5xx. Server errors and 429 responses make crawlers slow down deliberately, so a page that fails intermittently is more expensive than a page that fails honestly. If the crawl reports both, fix the errors before the 404s.

The failure that returns 200

The worst result a crawl can report is also the one it reports as fine. A page that returns a success status while showing an error message, or nothing at all, is a soft 404, and Google names it in the same document. The content suggests a failure even though the status code claims otherwise, and search consoles surface it as a soft 404 error.

This is why the visit is worth making. A crawler can read the status line. Only a person can tell that the page it read says "product no longer available" in a clean 200 response.

Permanent and temporary redirects

A redirect is a link that still works, which is why it is counted separately rather than as a failure.

The distinction that matters is in how redirects are interpreted. A permanent redirect shows the new address in search results. A temporary one shows the original address instead, and is not taken as a signal that the target should be the canonical page. So a redirect used to paper over a permanent move keeps the wrong URL in front of people for as long as it stays temporary.

Chains are the other thing to look at. Each hop is a delay and another chance for a client to give up, and a chain that ends at a 404 is a broken link that a status check alone would have missed.

The sitemap half of the crawl

Three counters read your sitemap against reality, and they find a different class of problem. A sitemap URL that no longer resolves is a stale entry. A page the crawl reached that is missing from the sitemap is one you are not telling anyone about. Listing both is how you catch a site that has quietly changed shape since the sitemap was generated, which is most sites.

The sitemap validator checks the file itself, and this crawl checks the pages behind it. Run them in that order, because a sitemap with parse errors cannot be compared against anything.

Images and alt text

A broken image is a visible hole in a page, and an image with no alt text is invisible to everyone who cannot see it. The crawl reports both together because they are the same class of failure. Neither costs you a crawl budget and both cost you a reader.

Why the example above is almost all zeros

The report at the top of this page crawls one page and finds nothing wrong, because example.com is a single page with no links on it. That is a truthful report rather than a broken tool, and reading it is a useful exercise.

The one number that is not zero is the one worth understanding. Not in sitemap reads 1, because the crawl reached one page and the site publishes no sitemap at all, so nothing is listed. On a real site that number would be a prompt to check whether the missing pages are missing on purpose.

Run the crawl on a site you own and the shape inverts. Most counters stay at zero, a few come back with something, and the ones that do are the list you work from.

Every guide in this category is on the Mini Tools page.