Google Is Indexing Your Search Results Pages
On this page
Internal search results pages get indexed for a simple reason: they are crawlable URLs like any other. A search box that submits with a GET request turns every query into a unique, linkable address such as ?q=blue+widgets, and once any of those addresses is linked, Google can find and crawl it. The result can be a supply of machine-generated pages that repeat your category pages with less substance. Keep these pages out of the index first, then stop them being crawled, preserve the ones that earned external links, and close the form behavior that creates them.
Confirm the problem first
Measure the exposure before changing anything. A site: search combined with your search parameter, such as site:example.com inurl:q=, gives a quick look at which search URLs Google has indexed. Treat it as a spot check: Google’s documentation on the site: operator says it may not list all indexed URLs. The Page indexing report’s list of indexed pages is the firmer view.
What you need from this step is the parameter pattern and a sense of scale. The pattern is what your rules will match, and the scale tells you whether any of these URLs should be preserved.
Why these pages can be a liability
They are thin and duplicative. A results page is a list of links in a template, sometimes with no results at all. A results page for “blue widgets” overlaps with your “blue widgets” category page, so a weaker page may end up competing for the same queries as the page you want to rank.
They spend crawling on the wrong pages. Every search URL Googlebot fetches is a fetch that didn’t go to a product, category or article page. With combinations of query strings, the number of possible URLs has no natural limit.
They are an opening for spam. Google’s help for the Manual Actions report lists internal search pages among the places where third-party spam can appear, and gives as an example internal search results where the user’s query appears geared at promoting a third-party website or service. If anyone can type a query and your site turns it into an indexable page, someone else can publish on your domain.
The two-layer fix, in order
Layer one: keep the pages out of the index. Add noindex to the search results template, as a meta robots tag or as an X-Robots-Tag HTTP header.
Layer two: stop the crawling. Once the existing search URLs have been recrawled and have dropped out of the index, add a robots.txt rule that matches your parameter pattern, such as Disallow: /*?q=, plus Disallow: /*&q= if the parameter can follow others in the URL, so Googlebot stops fetching them.
The order is the whole game. Google’s guide to blocking indexing with noindex says the page must not be blocked by robots.txt for the rule to be effective: if the crawler can’t fetch the page, it never sees the noindex, and the page can still appear in search results, for example if other pages link to it. Add the Disallow first and the URLs already in the index can stay there. Let the noindex be seen and honored, confirm the URLs have dropped, then add the Disallow. From then on, Google can’t see the noindex on search URLs that get linked later, and the same guide warns that a blocked page can still appear if other pages link to it. That is one more reason to fix the form itself.
Preserve the ones that earned links
A search URL without external links needs nothing beyond the two layers. Occasionally a search URL has earned real links because someone found a filtered view useful. For those specific URLs only, add a permanent redirect to the closest relevant category or landing page. Google’s redirects documentation says a permanent redirect is a signal that the target should be canonical.
Keep this surgical. A sweeping rule that redirects thousands of unlinked search URLs adds complexity for little benefit, and sending a large set of URLs to one loosely related page isn’t what the redirect is for.
Close the root cause, then monitor
The durable fix is in the form. A search form that submits with GET creates a URL for every query; a form that submits with POST, or one that fetches results with JavaScript without changing the address, doesn’t create crawlable URLs in the first place. Where you can change the implementation, doing so removes the source instead of leaving you to keep deindexing what a leaky form produces.
Then keep watching, because the problem can come back. Re-run the site: check and watch the Page indexing report for the parameter pattern, especially after template changes or when another team ships a new search feature without the noindex template.
Don’t use the Removals tool as the fix. Google’s help for the Removals tool says a temporary removal lasts only about six months, and Google can keep crawling the URL. The two layers are the mechanism.
Frequently asked questions
Should I just add a Disallow in robots.txt and be done?
No. A Disallow stops crawling, but it doesn’t remove pages already indexed, and it stops Google from seeing a noindex you add later. Add noindex first, confirm the pages have dropped, then add the Disallow.
How do I check how bad the problem is?
Run a site: search with your search parameter for a quick look, then check the Page indexing report. Google says site: may not list every indexed URL, so treat it as a spot check, not a count.
Will switching the search form to POST fix the pages already indexed?
No. It stops new ones from being created. The URLs already in the index still need noindex to drop out, and the ones with external links still need their redirects.
Can spammers use my search pages?
Google’s Manual Actions help names internal search pages among the places third-party spam can appear. Keeping search results out of the index removes the payoff.