Skip to content
EnterraHost
Products
Tools
Hosting Email Support Client Area Find a domain

Robots.txt allow and disallow. Which rule wins, and why

A robots.txt allow and disallow rule decides one thing only. May this crawler fetch this URL? It does not decide whether the URL can appear in search results, and it is not a way to remove a page. Those are different jobs with different tools, and confusing them is the most common mistake on the web.

You can fetch a robots file and test real paths against it without an account. This is how the decision is made.

What a robots.txt file is for

The file lives at the top level of a domain, at /robots.txt, and it tells crawlers which parts of the site they may fetch. Its purpose is to keep crawlers out of places they do not need to go, which protects your server from wasted requests and keeps your crawl budget pointed at pages you want indexed.

Google says this plainly in its own documentation. The file is used mainly to avoid overloading a site with requests, and it is not a mechanism for keeping a web page out of Google. Anything that claims otherwise is selling you a removal tool that does not remove anything.

How a robots.txt allow and disallow rule is chosen

A file is divided into groups. Each group starts with one or more user-agent lines naming the crawlers it applies to, and the allow and disallow lines that follow belong to that group until another user-agent line appears.

Three rules decide the outcome, and only the first two are well known.

The group is chosen by the most specific matching user agent. A crawler looks for the group that names it most precisely and ignores every other group, including the one for everyone. Google's own description of the specification puts it as finding the group with the most specific user agent that matches, with the rest ignored. So a site that writes a User-agent: * block and expects it to also cover Googlebot is wrong whenever a Googlebot group exists elsewhere in the file.

Inside the group, the longest match wins. RFC 9309, the standard that governs all of this, requires the most specific match to be used and defines most specific as the one with the most characters. A rule for /shop/private/ beats a rule for /shop/ every time, no matter which one appears first in the file. Order in the file is not precedence.

If an allow and a disallow match equally, allow wins. That tie-break is in the same section of the standard, and it is the reason an empty disallow line is safe. Disallow: with no path matches nothing and allows everything.

Two smaller details catch people out. A rule that appears before the first user-agent line belongs to no group and is ignored, so a file that opens with a bare disallow line blocks nothing at all. And wildcards are supported, with * standing for any run of characters and $ anchoring a match to the end of the path, though both are an extension to the original convention rather than part of the oldest implementations.

A blocked page can still appear in search

This is the part worth understanding properly, because it is where the money is spent on the wrong fix.

Blocking a URL tells a crawler it may not fetch it. It does not tell a search engine that the URL is unimportant. Google is explicit that a page blocked by robots.txt can still have its URL indexed if other pages link to it, and the crawler does that without ever visiting the page.

The trap that follows is worse. The usual advice for a page you want gone is to add a noindex tag. Adding one to a page that robots.txt already blocks achieves nothing, because the crawler is not allowed to fetch the page and therefore never reads the tag. You have blocked the only mechanism that could have removed it. Blocking and noindexing are opposites here, and using both gives you the worst outcome rather than the strongest one.

If the goal is removal, the order matters. Let the crawler fetch the page, serve a noindex tag, and block the URL in robots.txt only after the page has dropped out of the results. That is also why a robots.txt file is the wrong tool for anything sensitive. It is a public document, anybody can read it, and a disallow line is an announcement that the path exists.

What robots.txt will not do

It will not slow a crawler down on request. The crawl-delay field is not in the standard, and Google states that fields such as crawl-delay are not supported, so crawl rate is configured in a search console rather than in this file.

It will not secure anything either. A disallowed path is still a reachable path, crawlers are not the only things that read URLs, and this has never been an access control.

It also does nothing for rankings in the way people hope. There is no crawl budget stored somewhere that improves a page by keeping crawlers away from other pages. The file directs effort, and past that its effect on rankings is indirect.

The last limit is easy to hit and hard to notice. Google enforces a 500 KiB limit on the file, which is 512,000 bytes, and ignores everything past that point, so a giant generated file stops working quietly at the place it was most needed.

Testing a file properly

Reading a robots file is not the same as testing it, because the decision depends on the crawler and the path together. A path can be allowed for one agent and blocked for another, and the group logic above is exactly where a file stops doing what its author intended.

Enter a hostname and the tool fetches the file, lists the groups it found, and tests real paths against it. Each result says whether the path is allowed and which rule decided it, so a surprising answer points at the line responsible rather than leaving you to work it out.

Run it against the paths you care about and at least two user agents, one of them Googlebot. The case worth catching is a path that is allowed for everyone and blocked for Googlebot, or the reverse, because that is the shape of a file that was edited by somebody who assumed one group covered everything.

When you need Googlebot to crawl something you blocked

Blocking too much is a real failure, and it usually happens to stylesheets and scripts. Google renders pages to index them, and a renderer that cannot fetch your CSS and JavaScript sees a page that looks nothing like the one your visitors get.

If your pages are built by JavaScript, check that the paths serving those files are allowed. That single check explains a large share of pages that look fine in a browser and empty in search results, and the sitemap guide covers the other half of the crawl problem, which is whether the URLs you are advertising can be read at all.

Every guide in this category is on the Mini Tools page.