How to Do SEO for a Site with 1M+ Pages

On this page

Google’s crawl budget guide is written with million-page sites in mind. Its first audience line is large sites with 1 million or more unique pages whose content changes moderately often, about once a week, and Google adds that its numbers are rough estimates, not exact thresholds. At this size the job changes. It stops being “get everything indexed” and becomes “get the right pages crawled, indexed and refreshed, and stop spending crawl on the rest.” That work starts with measurement, not with more content.

Measure what Google crawls

Two sources, used together:

  • Crawl Stats. Search Console’s Crawl Stats report shows the site-level trend, with breakdowns by response, file type, crawl purpose and Googlebot type. It tells you whether the host is healthy and where the requests are going in aggregate.
  • Server logs. Logs give you the per-URL view that Crawl Stats doesn’t: every request, which you can group by template (product, category, filtered listing, article, profile, utility pages).

Before you count anything in the logs, make sure the requests are Google’s. A user-agent string can be copied by any bot. Google’s guide to verifying Google’s crawlers describes command-line checks for one-off lookups, and for large-scale lookups it recommends an automatic solution that matches a crawler’s IP address against Google’s published list of IP addresses. At a million pages, the automatic route is the one that fits.

Then, for each template, answer two questions: how often Google requests it, and whether those requests land on URLs you want in the index.

Cut crawl waste at the source

The crawl budget guide’s first recommendation is to manage your URL inventory. It says that if Google spends too much time crawling URLs it shouldn’t, its crawlers might not explore the rest of your site, or might not increase your crawl budget. Its tools, in the order the guide lists them:

  1. Consolidate duplicate content, so crawling focuses on unique content rather than unique URLs.
  2. Block with robots.txt the URLs you don’t want crawled and can’t consolidate. The guide’s examples are differently sorted versions of the same page and infinite scrolling pages that duplicate information on linked pages.
  3. Return 404 or 410 for pages you have removed for good, and fix soft 404s.
  4. Avoid long redirect chains, which the guide says have a negative effect on crawling.
  5. Make pages efficient to load. The guide says that if Google can load and render your pages faster, it might be able to read more content from your site.

Faceted navigation can be a large source of waste on catalog sites, when filters multiply thousands of products into a huge number of URL combinations. Google’s faceted navigation guide gives a clear split. If you don’t need the filtered URLs indexed, prevent them from being crawled, either with robots.txt disallow rules or by putting filters in URL fragments, which Google Search generally doesn’t support in crawling and indexing. Using rel="canonical" or nofollow on filter links, it says, is generally less effective in the long term. Pick one mechanism per URL pattern. A disallow and a noindex on the same URLs work against each other: a blocked page is never fetched, so its noindex is never read. And a disallow doesn’t guarantee absence: Google’s consolidation guide says it may still index URLs disallowed in robots.txt, without their content.

If you remember Search Console’s URL Parameters tool, don’t plan around it. Google announced on March 28, 2022 that it was deprecating the tool in one month, and that its crawlers would learn how to deal with URL parameters automatically.

Manage the index, not just the crawl

“Crawled – currently not indexed” means Google fetched the page and did not index it. Google’s Page indexing report help says such a page may or may not be indexed in the future, and that there is no need to resubmit it. The crawl budget guide explains the step in between: after crawling, each page is evaluated, consolidated and assessed for its suitability for the index. A large and growing count in that status means a large part of your page population isn’t clearing that evaluation.

Set an indexing threshold: a bar a page has to clear to deserve a URL in the index, such as unique content, enough substance and real search demand. Then decide what happens below the bar:

  • Near-duplicates: consolidate into a stronger page and redirect.
  • Pages nobody needs: remove them with a 404 or 410.
  • Pages users need but searchers don’t: keep them with a noindex. Know what that buys. It keeps the page out of results, but the crawl budget guide says not to use noindex to manage crawling, because Google still requests the page before it sees the rule.

Google’s ranking systems guide also notes that site-wide signals are used alongside page-level ones, which is one more reason not to leave a large thin tail in the index.

Direct attention with architecture and sitemaps

Internal links shape where crawlers find new and important URLs. Link priority pages from the sections Google already revisits, such as the home page and major hubs, and give them more internal links.

Split XML sitemaps by section and page type. A sitemap doesn’t command Google to crawl anything; at this scale, it also becomes a diagnostic tool. The Page indexing report has a sitemap filter: you can view all known pages, only submitted pages, only unsubmitted pages, or the URLs of one specific sitemap. With sitemaps split by section, that last option turns one site-wide number into a per-section view of what is indexed and what isn’t.

Keep lastmod honest. Google’s sitemap guide says it uses the value when it is consistently and verifiably accurate, and a million-URL sitemap set that restamps every date on each deploy gives it no reason to.

Tier the content investment

You can’t hand-write a million pages. Set the investment by template and by the value of the page:

  • Top tier: fully written content for the highest-value, most competitive pages.
  • Middle tier: strong templates enriched with page-specific data, context and any real user content.
  • Long tail: the minimum that still clears the indexing threshold. Anything that doesn’t clear it is excluded.

Template pages clear the bar through data that is distinct per page. A million pages built from one template with a single swapped word are candidates for “Crawled – currently not indexed”, and if they were built mainly to manipulate rankings, Google’s spam policies treat that as scaled content abuse.

Get new pages discovered fast

A site that publishes constantly has a second crawl problem: new URLs waiting to be found. Link new pages from sections Google already revisits, list them in sitemaps with accurate lastmod values, and watch the Discovery share in Crawl Stats, which counts requests for URLs Google has never crawled before. After that first crawl, later visits count as Refresh.

Frequently asked questions

Should I try to get all million-plus pages indexed?

No. Aim to get the pages that clear your threshold crawled, indexed and kept fresh, and exclude the rest deliberately. Consolidate duplicates, remove what nobody needs, and keep the pages that serve users but not searchers out of results.

Do I need log files, or is Search Console enough?

You need both. Crawl Stats shows site-level trends and breakdowns; logs show which templates and URLs get the requests. Verify that the requests come from Google before trusting the counts.

Does noindex save crawl budget?

No. It keeps a page out of search results, but Google still has to request the page to read the rule. To stop the crawling itself, consolidate the URLs or block them with robots.txt.

Leave a comment

Your email address will not be published. Required fields are marked *