Parameter URLs Are Eating Your Crawl Budget
On this page
Filter, sort and tracking parameters multiply URLs, and the result is two separate problems: crawling spent on copies, and copies sitting in the index. They need different tools, and the tools depend on each other. robots.txt stops crawling but doesn’t remove anything from the index. A canonical signals which version should consolidate the duplicates, but only once Google has crawled the page and read it. So work in sequence: classify each parameter, normalize parameter order in code, consolidate what is already indexed, block what never needs crawling, and keep the sitemap to clean URLs.
How the explosion scales
The growth is multiplicative. A category with five filter dimensions produces every combination of selected values across them. A sort parameter multiplies that set again, and pagination inside each combination multiplies it once more.
Parameter order adds a layer of its own. To a crawler, ?color=blue&size=large and ?size=large&color=blue are two URLs with the same content. A server that doesn’t normalize order creates a duplicate for every ordering of every combination. The pages look fine to a person, so the duplication shows up in logs, a crawl or the Page indexing report rather than on the page.
Crawl waste and index bloat are different problems
- Crawl waste is Googlebot requesting low-value parameter URLs instead of your important pages. Google’s crawl budget guide warns that time spent on URLs Google shouldn’t crawl can leave the rest of a site unexplored.
- Index bloat is near-duplicate parameter URLs already in the index, splitting signals and competing with the clean version.
robots.txt only addresses the first. Google’s robots.txt introduction warns that blocking a URL doesn’t keep it out of the index when other pages link to it. And a blocked page is a page Google can’t read, so any canonical on it goes unseen.
Classify every parameter
Put each parameter in one of three groups:
- Filters with their own search demand. A category narrowed to one attribute people search for can deserve an indexable URL. The demand has to be shown in query data, not assumed.
- Deep filter combinations. The intersection of three or four filters can be a long-tail page with no demand of its own; where that is the case, it shouldn’t compete in the index.
- Sort, session and tracking parameters. Re-sorting or tagging the same products creates no new content. These variants should never be indexed.
When unsure, treat a parameter as group two or three until data promotes it.
The order of operations
The right sequence depends on whether the URLs are already indexed.
- Parameter URLs already in the index. Point their canonical at the clean URL and let Google crawl them. Google’s guide to consolidating duplicate URLs says not to use robots.txt for canonicalization, so a disallow at this stage can leave the duplicates indexed and unconsolidated. Once they have consolidated, you can add a disallow for patterns that no longer need crawling.
- Parameter patterns not yet crawled at scale. Prevent crawling from the start. Google’s faceted navigation guide recommends robots.txt disallow rules, or putting filters in URL fragments, for filter URLs that don’t need to be indexed.
- Parameter order, in both cases. Normalize it server-side so every combination has one canonical form. This is a code change, and it removes a whole layer of duplicates before any signal has to work.
Skip noindex here: the consolidation guide advises against it for choosing a canonical within one site. For parameter duplicates, the canonical is the practical consolidation signal.
Your own tracking parameters
Campaign parameters are a self-inflicted source of duplicates. A link tagged with UTM parameters may get shared and crawled, leaving Google a tagged copy of a page that should exist only in its clean form. Two fixes:
- keep campaign parameters off internal links entirely, since they are meant for traffic arriving from outside;
- emit a canonical to the clean URL on every tagged variant, ideally in middleware so each variant is covered without per-page work.
Diagnose with logs, curate the sitemap
Complete server logs show every request, so grouping Googlebot requests by URL pattern shows which parameter group consumes the most crawling and how that shifts after your changes. Search Console’s Crawl Stats gives the site-level view; the logs give the per-pattern view.
The sitemap should list only the clean URLs you want indexed. The consolidation guide says every page listed in a sitemap is suggested as a canonical, so a generated sitemap that includes parameter permutations argues against your own canonicals.
In the Page indexing report, three reasons show the scale of the problem: “Discovered – currently not indexed”, “Crawled – currently not indexed” and “Duplicate without user-selected canonical”. Export the affected URLs by reason and group them by pattern to find your largest parameter group, then watch the same groups after deployment. A rise in “Crawled – currently not indexed” alongside a fall in “Discovered – currently not indexed” can mean Google is now reaching the parameter URLs and declining to index them, which is the intended outcome.
The Removals tool isn’t a fix here. It hides URLs for about six months, and the consolidation guide rules it out for canonicalization, since a removal applies to every version of the URL.
Expect a slow recovery
Crawl distribution can start to shift once the canonicals are processed, while the long tail of parameter URLs takes longer to be recrawled and consolidated. If a platform migration is coming, start the cleanup before it, so the work is underway rather than starting on launch day.
Frequently asked questions
Should I use robots.txt or canonical tags on parameter URLs?
It depends on their state. If the parameter URLs are already indexed, use canonicals first and let Google crawl them, then block patterns that no longer need crawling. If they aren’t crawled or indexed yet, blocking them from the start is the more direct route.
Why are my “Discovered – currently not indexed” counts so high?
Google’s help page ties that status, typically, to server load rather than to a verdict on the pages. On a site with parameter sprawl, those URLs compete for the same limited crawling. Find the dominant parameter pattern, consolidate or block it, and take it out of the sitemap.
Do I still configure parameters in Search Console?
No. Parameter handling lives in your canonicals, robots.txt rules, server-side normalization and sitemaps.