How to Find and Fix Orphan Pages

On this page

An orphan page is a page no internal link points to. Visitors browsing your site have no path to it, and Google can only learn about it from a sitemap, a crawl request or a link from another site. Finding orphans isn’t a single report; it is a comparison between the URLs a link-following crawl reaches and the full list of URLs you know exist. Fixing them means building real paths to the pages that deserve to be found, such as related-content modules, breadcrumbs, hub pages and subcategories, and leaving alone the pages that are orphaned on purpose.

How Google finds pages

Google’s page indexing documentation is plain about discovery: for Google to learn about a page, you must submit a sitemap or a crawl request, or Google must find a link to the page somewhere. An orphan has no internal link, so on your own site it depends on the sitemap.

A sitemap helps, but it isn’t a guarantee. Google’s sitemap overview says a sitemap helps search engines discover URLs but doesn’t guarantee that everything in it will be crawled and indexed, and notes that on large sites it is harder to make sure every page is linked by at least one other page. A page in the sitemap and nowhere else is known, not connected.

Don’t read orphan status into an indexing label. The page indexing report defines “Discovered – currently not indexed” as a page Google found but hasn’t crawled yet, typically because crawling it was expected to overload the site, so Google rescheduled the crawl. That status can apply to any page, linked or not. Treat the orphan check and the indexing report as two separate diagnostics.

Detection is a comparison

Finding orphans takes two datasets and a comparison, whichever tool runs it.

  1. Crawl from the homepage with a site crawler. The crawl follows links and gives you the set of URLs reachable through internal linking.
  2. Build the list of URLs that should exist from your XML sitemap, a CMS export and, if you have them, server logs.
  3. Compare the two. URLs in the second list that the crawl never reached are orphans. Site crawlers can do this directly when you add the sitemap as a second source and filter for URLs with zero internal links.

Server logs add evidence the other sources don’t: which URLs Googlebot has requested over time. A URL Googlebot has never requested and that has no internal link path is an orphan in every sense.

Orphans and near-orphans on large catalogs

On large catalogs, the problem takes two forms.

  • Buried by pagination. Products so deep behind page after page of category listings that they sit dozens of clicks from the homepage. They are reachable in theory and hard to reach in practice, for visitors and crawlers alike.
  • Dropped entirely. A catalog or query bug removes products from the listings altogether, so no path reaches them.

Platform migrations deserve a specific warning because they can orphan pages silently. When templates, URL structures or query logic change during a re-platform, links that used to render can simply stop rendering, and pages that were linked yesterday become orphans with no error to alert you. Run the crawl-versus-list comparison after every migration.

The JavaScript filter trap

A “fix” can fail to de-orphan anything if the links it adds aren’t real links. Google’s guidance on crawlable links says Google can generally only crawl a link if it is an <a> element with an href attribute. Its pagination guide adds that Google’s crawlers don’t click buttons and generally don’t trigger JavaScript functions that require user actions, and its faceted navigation guide says Google Search generally doesn’t support URL fragments in crawling and indexing. Filters or related-product widgets that change state only in the browser, or use # fragments, look like linking to a person and may give a crawler nothing to follow. Only real URLs in <a href> links de-orphan a page.

Prioritize before you fix

Not every orphan is a problem, so triage first.

  • Fix the valuable ones first: live products that should be selling, service pages, content meant to rank.
  • Leave intentional orphans alone, such as paid-search landing pages and thank-you pages kept out of the link graph on purpose.
  • Retire the dead ones: for pages with no demand and no value, the answer is noindex or removal, not a new link.

Fixes that don’t require a re-platform

Orphan fixes are ordinary internal-linking work:

  • Related-content modules so each page gets contextual links from its neighbors.
  • Breadcrumbs so parent pages link down to their children.
  • Crawlable hub or category pages that link to the orphaned set.
  • Subcategories that split oversized categories, so products sit at a reasonable depth instead of on page 28 of one giant listing.

On a large catalog, a broad fix is auto-populating related-product links for the zero-link set, which removes the “no path” status across thousands of pages at once. Know the limit: de-orphaning gets a page found and crawlable; whether it then gets indexed depends on the page itself. The page indexing report describes “Crawled – currently not indexed” as a page Google crawled but didn’t index, which may or may not be indexed in the future. De-orphaning helps discovery; it isn’t sufficient for indexing.

Measure the fix on each subsequent crawl: the count of URLs with zero internal links should fall to the intentional set, and the count with only one link should drop. Then track the indexing status of the formerly orphaned pages in Search Console.

Frequently asked questions

If a page is in my sitemap, isn’t it discoverable?

Discoverable, yes. Google’s sitemap overview says a sitemap helps discovery but doesn’t guarantee crawling or indexing. A page that exists only in the sitemap is known to Google but has no path from the rest of your site, for visitors or crawlers. An internal link is what connects it.

Will adding the orphaned page to my navigation fix it?

It removes the orphan status. For a large set of orphans, though, adding everything to the global navigation bloats the menu for visitors. Contextual links from relevant pages, such as related-content modules, in-body links and breadcrumbs, connect each page where it belongs.

How is an orphan page different from a page Google chose not to index?

An orphan has no internal link path; the problem is discovery. A page reported as “Crawled – currently not indexed” was found and crawled but not indexed. De-orphaning solves the first. It doesn’t guarantee the second.

Leave a comment

Your email address will not be published. Required fields are marked *