Skip to content
EnterraHost
Domain Names Shared Web Hosting Business Email Our TLDs Mail Protection EnterraMon EnterraSEO Ping Test Traceroute Server Check IPv6 Test MAC Address Lookup Robots.txt Tester Broken Link Checker Support Status Client Area

Broken link checker.

Crawl a site, follow what it links to, and see what fails. Broken links, broken images, missing alt text, redirects that cost a round trip, and sitemap problems that waste crawl budget.

Up to 25 pages, plus a check of the sitemap. This fetches real pages, so give it up to a minute. Public addresses only.

The distinction that matters

Dead is not the same as refused.

A crawler that treats every non-200 status as broken will hand you a list that is mostly wrong, and you will stop reading it.

Genuinely dead

A 404 means nothing is there. The page was moved, deleted or never existed, and every visitor following that link hits a dead end. This is the list worth acting on, and it is usually short.

Possibly blocked

A 401, 403, 429 or 500 is something answering rather than nothing being there. Plenty of perfectly healthy pages refuse automated requests. Check these by hand before changing any of them, because most will be fine.

Crawl budget

The sitemap check finds more than missing pages.

Cross-checking the crawl against the sitemap does two obvious things. It finds a page the crawl reached that the sitemap does not list, and a sitemap URL the crawl could not reach. Both are worth knowing, and the second is the more interesting of the two, because a URL you have advertised and cannot reach is a contradiction you published yourself.

It also surfaces pages carrying a noindex tag or a canonical pointing at a different address. Neither of those is a broken link, and both waste crawl budget and send search engines a mixed signal about which page should rank. They are the kind of thing that sits unnoticed for years because nothing errors.

Fix the causes, not the list. If twelve pages report the same dead link, that link lives in a template, and repairing twelve pages by hand leaves the thirteenth to come back tomorrow.

Limits

What this crawl does not cover.

25 pages, not all of them

It crawls a sample from the address you give. A problem in a section the crawl never reached will not appear, so run it from a different starting point if you have a specific area in mind.

A sample in time

It reports what it found when you ran it. A link that breaks next week is not in this result, and nothing here is scheduled. Run it after a deploy rather than trusting an old report.

Pages behind a login

Anything requiring a sign-in is invisible to the crawl, which sees the login page and stops. That is a limit of unauthenticated crawling rather than a fault in the tool.

Next

Check what a crawler is allowed to see.

A crawl reports what it can reach. The rules file decides what any crawler may reach in the first place, and a single stray line there can hide a whole section.

Test a robots.txt file