Skip to content
EnterraHost
Domain Names Shared Web Hosting Business Email Our TLDs Mail Protection EnterraMon EnterraSEO Ping Test Traceroute Server Check IPv6 Test MAC Address Lookup Robots.txt Tester Broken Link Checker Support Status Client Area

Robots.txt tester.

Fetch a site robots file, list what it actually says, and test real paths against it. The rules are matched the way crawlers match them, not the way people assume.

Reads the robots.txt at the site root. Public addresses only, and redirects are not followed.

The part people get wrong

Robots.txt controls crawling, not indexing.

This is the single most expensive misunderstanding in the file. A URL you disallow is not removed from search, and can still rank. A crawler that is blocked never reads the page, so it never sees a noindex instruction, and a page linked from elsewhere can be indexed on the strength of that link alone.

To block crawling

Use robots.txt. This is what it is for, and it is the right tool for keeping a crawler out of paths it wastes time on, such as internal search results or faceted URLs.

To keep a page out of the index

Use a noindex robots meta tag or the equivalent HTTP header, and let the page be crawled so the instruction is actually read. Blocking the crawl and asking for noindex at the same time cancels the second request.

Matching

Three rules that decide the answer.

Prefix, not exact

Disallow /private also blocks /private-notes and /private/report. A path with no wildcard and no end anchor matches anything that starts with it, which is why a rule that looks tidy can be wider than intended.

Longest match wins

Allow and Disallow are not a first-match list. Every rule is tested and the most specific one, measured by pattern length, decides. A short Disallow cannot override a longer Allow.

Allow wins a tie

When an Allow and a Disallow match with the same pattern length, the allow rule wins. That is the behaviour in RFC 9309 and it is the opposite of what most people expect when they write the two lines in the wrong order.

Next

Once crawling is right, check what it finds.

A correct robots file does not tell you whether the pages behind it resolve. The next check follows the links and reports what breaks.

Check for broken links