How to Fix “Blocked by Robots.txt” in Google Search Console

On this page

“Blocked by robots.txt” means you stopped Google from crawling a URL. Crawling and indexing are separate stages, and robots.txt governs only the first. Google’s introduction to robots.txt says so directly: robots.txt is used mainly to avoid overloading your site with requests, and “it is not a mechanism for keeping a web page out of Google.” A blocked URL can still appear in search results, because Google can learn about it from links without ever reading the page. Match the directive to what you want: use robots.txt to stop crawling, noindex to stop a page appearing in search, and a password to keep something private.

Crawling and indexing are different stages

Robots.txt acts before the fetch. Googlebot checks it first, and if a rule disallows a URL, Google doesn’t download the page. Indexing is a later decision about whether a URL gets a place in search results. The consequence: robots.txt stops Google reading a page, not Google knowing it exists.

Google’s robots.txt introduction spells out what that means for a blocked page. The URL and, potentially, other publicly available information such as anchor text in links to the page can still appear in Google Search results. To keep a page out of Google, it says, block indexing with noindex or password-protect the page.

So the status tells you crawling is being prevented. Whether that is a problem depends on whether the page should be crawled at all. For a private path you also protect with a password, the block is expected. For a page you meant to remove from search, the block is the wrong tool.

When a blocked page is indexed anyway

The Page indexing report has a separate warning for this: “Indexed, though blocked by robots.txt.” Google says it means the page was indexed despite the block. Google always respects robots.txt, but that doesn’t necessarily prevent indexing if someone else links to the page. Google won’t request and crawl the page, but it can still index it using information from the page that links to it. Because of the robots.txt rule, any snippet shown for the page will probably be very limited.

A blocked page can therefore end up in results with a thin, uninformative listing. To check for it, open the warning’s examples in the report or inspect a sample URL in URL Inspection.

Why noindex fails on a blocked page

One source of “my noindex isn’t working” is using both directives on the same page: a noindex tag in the HTML and a disallow in robots.txt. Google’s guide to blocking indexing states the rule: for the noindex rule to be effective, the page must not be blocked by a robots.txt file. If Google can’t crawl the page, it never sees the noindex, whether it is in a meta tag or in an X-Robots-Tag HTTP header. The header arrives with the response to a fetch, and a blocked URL is never fetched.

This explains a loop that can run for weeks. Someone adds noindex, sees the page still listed, assumes the tag is broken, and adds a robots.txt disallow to “force” it out. The disallow is what stops Google from ever reading the tag. Each step feels like a tighter restriction and moves further from the goal. The two directives don’t stack. Robots.txt sits in front of the fetch, and noindex sits inside the response.

When a page that should be noindexed stays indexed, check robots.txt first.

Over-broad rules catch real pages

The second class of problem is a disallow wider than intended. A rule that blocks every URL with a query string, written to stop filter and tracking parameters, also blocks parameterized pages you want indexed. Platform defaults can include broad rules that cover real content, so read what each rule matches before you ship it.

The tool for this is the robots.txt report in Search Console. Google says it shows which robots.txt files Google found for the top 20 hosts on your site, when they were last crawled, and any warnings or errors. It also lets you request a recrawl of a robots.txt file after you fix an error or make a critical change.

Fix it by intent

Decide what you want for each blocked URL, then use the directive that does that job.

  • To keep a page out of search results: allow crawling and add noindex, as a meta tag or an X-Robots-Tag header. Google has to fetch the page to see the instruction. To remove a page from results, you have to let Google in far enough to be told to leave.
  • To stop Google fetching a URL: keep the disallow, and accept that if other pages link to it, the URL may still be listed with little or no snippet. Robots.txt manages crawling. It does not guarantee a page never appears.
  • To keep content private, such as a staging site, use a password. Google’s guide to controlling what you share recommends password protection for private content, which also keeps it out of search. If staging URLs are already indexed, a new robots.txt block is not a way to remove them. Password protection will, in time: Google says it eventually removes content that already appears in search.

Check before and after the change

Because the directives interact, check what Google is doing rather than reasoning about it. Open an affected URL in URL Inspection and see whether crawling is allowed and whether noindex is detected. If you remove a disallow so Google can finally see a noindex, expect a sequence, not an instant change:

  1. Google recrawls the now-allowed URL.
  2. It reads the noindex.
  3. It drops the page from results.

Requesting indexing after the change asks for that recrawl, and Google says a request doesn’t guarantee the outcome. When the warning changes from “Indexed, though blocked by robots.txt” to the noindex status (“URL marked ‘noindex'” in Google’s help page; check the exact label in your report), the fix has taken hold.

Robots.txt is still the right tool for crawling

None of this makes robots.txt obsolete. It remains the right tool for managing crawling. Google’s crawl budget guide recommends it for URLs you don’t want crawled at all, such as differently sorted versions of the same page and infinite scrolling pages that duplicate information on linked pages. The same guide warns against using noindex to save crawl budget, because Google still requests the page before it sees the tag.

Frequently asked questions

Why does my robots.txt-blocked page still show in Google?

Because robots.txt blocks crawling, not indexing. If other pages link to the URL, Google can index it without fetching it, and Google says any snippet will probably be very limited. To remove it, allow crawling and add noindex.

Can I noindex a page that is blocked in robots.txt?

No. Google says that for noindex to work, the page must not be blocked by robots.txt. Remove the disallow first, so Google can crawl the page and see the instruction.

What should I use to hide a staging site?

A password. Google’s guidance on private content points to password protection, and robots.txt is not a way to keep pages out of Google.

How do I check which robots.txt rules Google is using?

Use the robots.txt report in Search Console. It shows the robots.txt files Google found for your top hosts, when they were last crawled, and any errors, and it lets you request a recrawl after a critical change.

Leave a comment

Your email address will not be published. Required fields are marked *