How to Diagnose and Fix Crawl Traps

On this page

A crawl trap is a pattern on your site that generates an endless supply of low-value URLs: a calendar that pages forward forever, a session ID added to every link, a relative link that keeps extending the path, internal search that turns every query into an address. The cost is attention. Google’s guide to faceted navigation puts it plainly: if crawling is spent on useless URLs, the crawlers have less time for new, useful ones. The cure is to separate what users can do from what Googlebot can reach: keep the feature, remove the crawlable path, and use the right tool at each layer.

Unbounded generators are the real traps

Certain URL sources are large; traps are unbounded.

  • Calendars. A booking or events calendar with a crawlable “next month” link always has another month. Googlebot can follow it into years nobody will ever book.
  • Session IDs in URLs. An identifier stamped into the address means every visit, including every crawl, can produce a new URL for content Google already has.
  • Relative link loops. A malformed relative link that appends a path segment on each page produces addresses such as /shop/shop/shop/, each one a valid response.
  • Internal search. A search form that submits with GET creates a URL for every query, including queries typed by bots.

Filters are the multiplying version of the same problem: independent parameters can combine into far more URLs than there are products. That case has its own playbook; the same principles apply to it.

Diagnose from the crawl data

Two sources show which generator is responsible.

Crawl Stats in Search Console. The Crawl Stats report shows request volume, host status and breakdowns by response and purpose, with example URLs in each group. A trap can show up as requests far out of proportion to your real page count, or a rising trend with no matching content growth. The example URLs can name the pattern outright.

Your server logs. Filter to verified Googlebot requests and group them by URL pattern with the variable values removed, so that /events?month=2031-04 and /events?month=2032-09 fall into one bucket, /events?month=. For example, imagine one pattern taking a majority of Googlebot’s requests while mapping to a handful of real pages: that pattern is your generator. The status codes inside each bucket tell you whether Google is being served real pages, errors or redirects.

Stop creating the URL

The fix that removes a trap at the source is to stop generating the address. Filtering, sorting and searching are user actions; they don’t have to create new crawlable URLs.

  • Filters and sorting: keep the filter state after a #. Per the faceted navigation guide, a filtering mechanism based on URL fragments has no impact on crawling, because fragments aren’t something Google Search generally supports in crawling and indexing. The same guide recommends a 404 for filter combinations with no results.
  • Calendars: stop rendering crawlable links beyond the range where real events or availability exist.
  • Session IDs: keep session state out of the URL, for example in a cookie.
  • Relative link loops: fix the link so it resolves to the intended address, then let the looped URLs return an error.

Every URL you generate and then try to manage still costs something, because Google has to discover it and may need to fetch it to read your instruction. Stopping the URL from existing beats managing it after it exists.

The right tool for each layer

Tool What it does Use it for
No crawlable URL (fragment or client-side state) Nothing new to crawl Filters, sorting and search that change the view only
robots.txt disallow Stops crawling of the pattern; doesn't keep URLs out of results Patterns you never want fetched, once nothing from them is indexed
<!–INLINECODE1–> Allows crawling, keeps the page out of results URLs that must stay crawlable but mustn't appear in search
<!–INLINECODE2–> or <!–INLINECODE3–> Hints about preference Secondary measures; the faceted guide calls them generally less effective in the long term

The conflict to remember: Google has to crawl a page to see its noindex. A URL that is blocked in robots.txt and also carries noindex never has the noindex read.

Order matters when trap URLs are already indexed

The mistake that can make a trap stick: you find thousands of trap URLs in the index and add a robots.txt disallow. They may not go away. Google’s robots.txt introduction says a URL blocked by robots.txt can still appear in search results, without a description, and that robots.txt is not a mechanism for keeping a page out of Google.

So choose the order by what is already in the index:

  • Pattern already indexed: allow crawling, apply noindex to the pattern, and watch the indexed URLs for that pattern fall as Google recrawls them. Then add the robots.txt block to stop future crawling.
  • Pattern not indexed: block it in robots.txt now, or better, stop generating it.

For an urgent case, Google’s Removals tool hides a URL from results for about six months. It isn’t a fix; the durable solution is still noindex first, then block.

Frequently asked questions

Will robots.txt remove trap pages that are already indexed?

No. robots.txt controls crawling, not indexing, and Google’s introduction to it says a blocked URL can still appear in results without a description. Blocking also stops Google from seeing a noindex. Apply noindex first, let it process, then block.

How do I tell a crawl trap from a large but legitimate site?

Group Googlebot’s requests by URL pattern. A large legitimate site spreads crawling across its real pages. A trap concentrates requests on one parameterized or unbounded pattern that maps to a small number of unique pages.

What should an empty filter combination return?

A 404. Google’s faceted navigation guide recommends a 404 status when a filter combination doesn’t return results, rather than an empty page.

Leave a comment

Your email address will not be published. Required fields are marked *