How to Handle Duplicate Content Issues

On this page

Duplicate content is a canonicalization problem, not a penalty. Google’s canonicalization guide says some duplicate content on a site is normal and isn’t a violation of its spam policies. The same guide names two costs: the same content under many URLs can be a bad experience for people, who wonder which page is right, and it can make it harder to track how your content performs. Two practical costs follow: Google picks one version to show and may not pick the one you intended, and crawling spends time on copies. Tell Google which URL you want, using the mechanism that matches the cause; the work is URL architecture more than rewriting.

Fix by cause

Diagnose where the duplicates come from, then apply the matching signal. Google’s guide to consolidating duplicate URLs lists the signals by strength: redirects first, then rel="canonical", then sitemap inclusion.

  • Protocol and host variants (http and https, www and non-www, trailing slash or not). Redirect every variant to one version with a permanent redirect. This is the cleanest case and belongs at the server. The consolidation guide adds that Google prefers HTTPS pages over equivalent HTTP pages as canonical, unless there are problems or conflicting signals.
  • Parameter, sort and session duplicates. Decide first whether the parameter URLs need to be crawled at all. If they carry signals you want consolidated, such as links or shares, point a canonical at the clean URL and let Google crawl them. If they are pure crawl waste, block the pattern in robots.txt. Don’t do both on the same URLs: robots.txt stops the fetch, and with it any chance of Google reading the canonical. The consolidation guide says not to use robots.txt for canonicalization.
  • Near-duplicate variants such as size or color pages, or location pages that differ only in a name. Consolidate them into one page with a selector, or canonicalize them to a primary version, depending on whether each variant has its own search demand.
  • Syndicated copies on other domains. Agree with the partner before publication how their copy will be handled. Google’s guide to fixing canonicalization issues says the canonical link element isn’t recommended for avoiding duplication by syndication partners, because the pages are often very different, and that the most effective solution is for partners to block indexing of your content. Get that agreement in writing, because the partner controls that page.

Faceted navigation: decide filter by filter

Filters can generate huge numbers of near-duplicate URLs, so the real decision is which filtered pages deserve to be indexed at all. A useful test has three parts, all of which have to pass:

  1. Demand. People search for that filtered combination.
  2. Distinct value. The page offers something beyond a reordering of its parent.
  3. Inventory. Enough products sit behind it to make a substantial page.

Filters that fail a test shouldn’t be indexable pages; filters that pass all three deserve clean URLs. That is a per-pattern decision, not one rule for every parameter.

For the filters you do want indexed, Google’s faceted navigation guide sets URL rules of its own, covering parameter separators, filter order and empty results.

For the filters you don’t want indexed, the same guide says to prevent crawling of those URLs.

A canonical is a signal, not a command

Put a self-referential canonical on the pages you want indexed; the consolidation guide recommends including the canonical link on the canonical page itself. It can lower the chance Google improvises a different choice when it meets a variant.

But the declaration can lose. Google chooses the canonical from all the signals it has, so a canonical pointing at one URL while internal links, sitemaps and redirects favor another is a contradiction Google has to resolve. When your declared canonical is overridden, read it as a diagnostic: another signal may disagree with it. Align the signals, and the declaration has a better chance of holding.

One thing not to use: the consolidation guide says it doesn’t recommend noindex to prevent selection of a canonical page within a single site, because it blocks the page from Search completely.

Verify in the Page indexing report

Google’s Page indexing report help names three duplicate-related reasons, and each says something different:

  • “Duplicate without user-selected canonical”: Google found duplicates, no canonical was declared, and it chose one itself.
  • “Duplicate, Google chose different canonical than user”: you declared a canonical and Google picked another. This is the one to chase, because it shows where your signals and Google’s judgment split.
  • “Alternate page with proper canonical tag”: the page is a recognized duplicate pointing at its canonical. That is the intended state.

On similarity, Google’s canonicalization documentation gives no percentage that turns a page into a duplicate. Treat any “X percent similar” figure as a triage heuristic of your own, not a Google rule.

Frequently asked questions

Does duplicate content cause a Google penalty?

Not by itself. Google says some duplicate content is normal and isn’t a violation of its spam policies. The costs are practical: a confusing experience for people, harder performance tracking, Google showing a version you didn’t intend, and crawling spent on copies.

Should I rewrite duplicate pages to make them unique?

Not as the first move. Duplication from protocol and host variants, parameters and filters is structural, and signals are the tool for it; syndicated copies are handled by the partner blocking indexing of their copy. Rewriting matters only when two pages should be different and currently aren’t.

Can I use a canonical and a robots.txt block on the same URLs?

No. If robots.txt blocks a URL, Google never fetches it and never sees its canonical. Choose one: canonical for URLs whose signals you want consolidated, a crawl block for URLs that are pure waste.

Leave a comment

Your email address will not be published. Required fields are marked *