Google Crawls Daily But Your New Pages Take Months to Index

On this page

Daily Googlebot activity in your logs and slow indexing of new pages can both be true, because they answer different questions. A recrawl asks whether a page Google already knows has changed. A new URL first has to be found, then crawled, then chosen for the index. You can be crawl-rich and index-poor at the same time, and more sitemap resubmissions or indexing requests won’t fix a site-wide problem at the finding or the choosing step. For scale, Google’s guide to asking Google to recrawl says “Crawling can take anywhere from a few days to a few weeks.” New pages that take months point to a problem at one of three steps: finding, crawling or choosing.

What the daily crawl is spent on

Google’s crawl budget guide describes crawl demand in terms of popularity and staleness: URLs that are more popular tend to be crawled more often to keep them fresh, and Google’s systems want to recrawl documents often enough to pick up changes. An established page with links and a history of updates can earn that attention. A URL published yesterday has neither, so the attention your old pages receive doesn’t extend to it.

You don’t have to guess how your own crawl divides. Search Console’s Crawl Stats report groups requests by purpose: “Discovery” for a URL Google has never crawled before, “Refresh” for a recrawl of a known page. If Refresh dominates your daily requests and Discovery doesn’t rise after you publish, the daily crawl isn’t reaching your new pages, and more crawl activity elsewhere won’t change that.

Give new pages a path in

Google’s link best practices say Google uses links to find new pages to crawl, and that every page you care about should have a link from at least one other page on your site. A new product or article whose only link sits deep in a paginated archive has a long path in.

The Crawl Stats report also shows you where to put the links. Click into the Refresh group and it lists example URLs, which are pages Google is revisiting. Link new pages from those, and keep the path short:

  • A “latest” section on pages Google revisits, such as the home page or a main category page, that links straight to each new URL.
  • Contextual links from established articles or products on the same topic, so discovery doesn’t depend on archive pagination.
  • Fewer clicks from the home page to anything new, so the path stays short as the site grows.

The selection step

Being crawled doesn’t mean being kept. The Page indexing report shows two places where a new page can stall: “Discovered – currently not indexed” (found, not yet crawled) and “Crawled – currently not indexed” (crawled, not kept). The first points back to the path in and to crawl capacity. The second sends you to the page itself.

For the second, look at what the page adds. A manufacturer description reused across retailers, location pages that differ only by city name, or syndicated copy give Google little that it doesn’t already have. More crawling won’t change that; differentiating the page can. This is also why “publish more pages” and “get pages indexed faster” can pull against each other: a large batch of similar, thin pages competes for places none of them clearly earn.

What sitemaps and indexing requests can and can’t do

  • Sitemaps help Google learn that URLs exist. Google’s guide to building a sitemap says Google ignores the <priority> and <changefreq> values, and uses <lastmod> only when it is consistently and verifiably accurate. Resubmitting an unchanged sitemap gives Google nothing new to act on.
  • Request indexing in URL Inspection asks Google to crawl a URL. Google’s URL Inspection documentation says a request doesn’t guarantee the page will appear in the index, and the recrawl guide adds that there’s a quota and that repeated requests for the same URL won’t speed things up. Use it for a page you have improved, not as a bulk indexing method.
  • The Indexing API is not a general shortcut. Its quickstart says it can only be used for pages with JobPosting, or BroadcastEvent embedded in a VideoObject, and that usage beyond the default testing quota requires approval.

Measure which lever moved

After you change the paths in or the pages themselves, watch two things:

  1. Crawl Stats: do Discovery requests rise after you publish? If they do, Google is finding new URLs on the site.
  2. Page indexing: do new URLs move out of “Discovered – currently not indexed” (discovery improved) or out of “Crawled – currently not indexed” (the pages improved)?

The two statuses point to which change worked, and so to whether you fixed the real bottleneck or added more crawling to a site that was never short of it.

Frequently asked questions

If Google crawls my site every day, why doesn’t it index my new pages?

Because the daily crawl may be recrawling pages Google already knows. Check Crawl Stats by purpose: if Discovery requests don’t rise after you publish, the new pages aren’t being reached, and the first fix to try is linking them from pages Google revisits.

Does requesting indexing force a page into the index?

No. Google says a request doesn’t guarantee indexing, there’s a quota, and repeating the request for the same URL won’t make it faster.

Do sitemap priority and change frequency values help?

No. Google’s sitemap guide says it ignores <priority> and <changefreq>. An accurate <lastmod> is the value it can use.

Leave a comment

Your email address will not be published. Required fields are marked *