Your RSS Feed Is Creating Thousands of Indexed URLs
On this page
If a site: search on your domain turns up feed URLs, your feeds have become indexable documents in their own right. Google’s list of file types indexable by Google includes XML, and a CMS can generate a feed not only for the whole site but for every category, tag and author, and sometimes for comments. Each one is a URL that can be crawled and, once linked or referenced, indexed alongside the articles it repeats. Aim for a targeted repair: keep the feeds out of the index with an HTTP header, turn off the feed types you don’t use, and keep the main feed working for the subscribers and systems that rely on it.
Where the feed URLs come from
A default WordPress install exposes a main feed at /feed/, plus one for every category, tag and author archive: /category/recipes/feed/, /tag/quick-dinners/feed/, /author/jane/feed/, and in certain setups a comments feed per post. Multiply categories, tags and authors, and a modest blog can carry hundreds of feed endpoints nobody deliberately created. Paginated feeds such as /feed/?paged=2 add more.
These URLs aren’t in your navigation, but they are easy to find. The <link rel="alternate" type="application/rss+xml"> tag in your page head points at them, feed readers and aggregators request them, and other sites sometimes link to a feed directly.
Duplication, not a penalty
An indexed full-text feed repeats the article it syndicates in a second format. Google’s canonicalization documentation says “Some duplicate content on a site is normal and it’s not a violation of Google’s spam policies.” The cost is quieter: two URLs carry the same content, and when the feed is the one shown for a query, a bare XML file stands in for your designed page with its navigation, context and calls to action. The goal is one indexable URL per article, the article itself.
That framing also rules out the overcorrection. Disabling RSS entirely would remove the duplication, but it breaks every subscriber and aggregator for no gain over a targeted fix.
Measure the scope first
Search site:yourdomain.com inurl:feed to see which feed URLs Google shows. Treat it as a spot check: Google’s documentation on the site: operator says it may not list all indexed URLs. Then look at the Page indexing report’s list of indexed pages for URLs containing /feed/. That list is a sample too, capped at 1,000 rows, so use both views to learn which feed families are involved, not to get an exact count.
Before changing anything, check what your site already sends. Request a feed URL and read its response headers. If an X-Robots-Tag: noindex header is already there, for example from an SEO plugin, the feeds you see in search may simply not have been recrawled yet.
The fix: noindex in the HTTP header
You can’t put <meta name="robots" content="noindex"> in an RSS feed. A feed is an XML document with no HTML head to hold a meta tag. Google’s robots meta tag documentation says that to block indexing of non-HTML resources, you use the X-Robots-Tag response header instead.
Two ways to send it:
- Your CMS or SEO plugin’s feed settings, if they offer a
noindexor disable option for feed types. - A conditional header at the server or CDN. Match requests whose path contains /feed/, which covers variants such as /feed/atom/ and the paginated feeds, and add
X-Robots-Tag: noindexto the response. This covers every feed family with a /feed/ path at once and doesn’t depend on a plugin.
Subscribers aren’t affected. noindex governs whether a URL appears in search results; the feed is still served to anyone who requests it.
Keep the main feed working
Don’t treat every feed as clutter. Google’s guide to building a sitemap says that if your CMS generates an RSS or Atom feed, you can submit the feed’s URL as a sitemap, and that Google accepts RSS 2.0 and Atom 1.0. The same guide mentions WebSub for broadcasting changes to search engines when you use Atom or RSS. So keep the main feed switched on: it is a way to tell Google about new articles, separate from whether the feed URL itself appears in results.
Full-text or excerpt feeds
Decide deliberately between full-text and excerpt feeds, because the choice shapes both the duplication and where people read. A full-text feed publishes the whole article, which is the heaviest repetition and lets readers finish the piece without visiting the site. An excerpt feed publishes a summary and a link, which repeats less and points the reader to your page. If on-site visits matter to you, excerpts are the trade to consider; if readable-in-reader syndication is the goal, full text is a legitimate choice. Pick on purpose rather than accepting the CMS default.
Prune the feeds you don’t offer
An index stays tidier when fewer endpoints exist to begin with. If you don’t promote category, tag or author feeds, turn those feed types off rather than only adding noindex. URLs that are no longer generated can’t be indexed again.
Then watch the feed URLs leave. Removal depends on recrawling, so it won’t be instant. Re-check the site: query and the Page indexing report at intervals, including the paginated feeds.
Frequently asked questions
Will noindexing my feeds stop people from subscribing?
No. noindex only affects whether the feed URL appears in search results. The feed is still served on request, so readers and subscribers receive it as before.
Should I disable RSS to be safe?
No. Disabling RSS breaks legitimate subscribers and removes a feed you can submit as a sitemap. Add noindex to the feeds you keep, and turn off only the feed types you don’t offer.
Why can’t I use a meta robots tag on a feed?
A feed is XML, not HTML, so it has no head to carry a meta element. Google’s documentation says to use the X-Robots-Tag HTTP header for non-HTML resources.