Google Discovers Your Site But Won’t Index It
On this page
When Search Console lists pages as “Discovered – currently not indexed”, their content has not been evaluated and rejected. Google hasn’t crawled them yet. Google’s Page indexing report help defines the status plainly: the page was found by Google, but not crawled yet. “Typically, Google wanted to crawl the URL but this was expected to overload the site; therefore Google rescheduled the crawl.” That is why the report shows no last crawl date for these URLs. Google has not read the page, so the page’s content isn’t what it assessed. The question for this status is why the crawl hasn’t happened, and Google’s own explanation points first to your server.
Give new pages time first
Part of a Discovered list is a queue. Google’s help says that after a URL is known it can take some time, up to a few weeks, before Google crawls some or all of a site. It also says indexing is never instant, even with a crawl request, and that Google doesn’t guarantee every page will be indexed. A batch of new URLs that has sat in Discovered for a week is still inside the window Google describes. A list that keeps growing for months is past it.
Check whether your server can take the crawl
Because the typical cause Google gives is load, the first place to look is the Crawl Stats report in Search Console. Google describes it as a report for advanced users that shows how many requests Google made, what your server returned and any availability issues it encountered.
Host status summarizes availability over the last 90 days. A green status means Google found no significant crawl availability issues. Google assesses availability in three categories, and a significant error in any of them can lower the status:
- robots.txt fetching. Google requests robots.txt regularly. If the request doesn’t return a valid file or a “does not exist” response, Google slows or stops crawling until it gets one. A 200 with any file, even an empty or invalid one, counts as a success. So do 403, 404 and 410, which mean there is no file. A 429 or a 5xx counts as a failure. Google’s timeline for a failure is specific:
- for the first 12 hours it stops crawling the site;
- from 12 hours to 30 days it uses the last robots.txt it fetched successfully;
- after 30 days it crawls as if there were no file if the home page is available, and stops crawling if it isn’t.
- DNS resolution. Failures show when your DNS server didn’t recognize the hostname or didn’t respond. Google’s example of an issue is DNS resolution failing for more than 5% of requests on a given day.
- Server connectivity. This shows when the server was unresponsive or didn’t return a full response.
The report also shows average response time and a breakdown by response code. A server that slows down or returns errors under crawling fits the situation Google’s definition describes. When Google announced the retirement of its crawl rate limiter tool, it said that if a server persistently returns HTTP 500 errors for a range of URLs, Googlebot will slow down crawling automatically and almost immediately.
Two details in the report change how you read it:
- Redirects count as separate requests. If page 1 redirects to page 2, which redirects to page 3, each is its own request. Chains cost crawl requests.
- Crawl purpose separates discovery from refresh. A “Discovery” request is for a URL Google has never crawled before; “Refresh” is a recrawl of a known page. Google says that after adding a lot of new content or submitting a sitemap, you should ideally see a bump in discovery crawls. If the Discovered list is growing and discovery crawls are not, the crawl that would clear it isn’t happening.
At scale, it can be a crawl budget problem
Google’s crawl budget guide is written for very large or fast-changing sites. It also names sites with “a large portion of their total URLs classified by Search Console as Discovered – currently not indexed.” Its advice centers on the URLs you ask Google to crawl. It warns that if Google spends too much time on URLs it shouldn’t crawl, its crawlers might not explore the rest of the site.
Practices the guide recommends, and traps it names, include:
- Consolidate duplicate content, so crawling goes to unique content rather than unique URLs.
- Block unimportant URLs with robots.txt, such as differently sorted versions of the same page, if you can’t consolidate them.
- Don’t use
noindexto save crawl budget. Google still requests the page, then drops it when it sees thenoindex, which wastes crawling time. - Don’t use robots.txt to move budget around temporarily. Google says it won’t shift freed-up crawl budget to other pages unless it is already hitting your site’s crawl capacity limit.
- Return 404 or 410 for pages removed for good. A 404 is a strong signal not to crawl that URL again, while blocked URLs stay in the crawl queue much longer.
- Eliminate soft 404s. They continue to be crawled and waste budget.
- Keep sitemaps up to date, with
lastmodfor updated content. - Avoid long redirect chains, and make pages efficient to load.
The trap to avoid is the idea that blocking pages with robots.txt automatically moves the freed crawl to the pages you care about. Google says that only happens when it is already limited by your site’s capacity. Consolidating duplicates and returning 404 or 410 for dead URLs cuts wasted crawling; neither guarantees a reallocation.
Look at how Google reaches the page
On a healthy server with a modest number of URLs, look at discovery itself. A URL that Google knows only from the sitemap has no link path to prompt a crawl, and one buried several clicks from any page Google revisits has only a long one. Links from pages that are crawled regularly, such as the home page, section hubs and recent articles, give it a path that doesn’t depend on the sitemap alone.
Read the string literally before you act. “Discovered” means not yet crawled. “Crawled – currently not indexed” means Google fetched the page and did not index it, and it calls for a look at the page itself. Treating one as the other sends the work in the wrong direction.
What doesn’t clear the list
- Mass resubmission. Request indexing in the URL Inspection tool asks Google to crawl one URL. Google says a request doesn’t guarantee indexing, and there is a daily limit. It is for a handful of important pages, not for a backlog of thousands.
- Adding
noindexto “save” budget. Google still requests those pages. - Waiting on a server that keeps failing. If host status shows availability problems, fix them first; Google slows crawling while they persist.
There is no fixed date for a Discovered page to move. Fix availability, cut wasted URLs, give the important pages a path, and watch the discovery crawls and the list together in Search Console.
Frequently asked questions
Does “Discovered – currently not indexed” mean Google thinks my page is low quality?
Not according to Google’s definition. The page has not been crawled, so its content has not been evaluated. Google says the typical reason is that crawling was expected to overload the site, so it rescheduled the crawl.
How long should I wait before worrying?
Google says it can take up to a few weeks after a URL is known before it is crawled, and that indexing is never instant. A list that keeps growing for months calls for checks, starting with host status, discovery crawls and how the pages are linked.
Will blocking low-value pages help my other pages get crawled?
It stops wasted crawling, but Google says it won’t shift the freed crawl budget to other pages unless it is already hitting your site’s crawl capacity limit. The crawl budget guide puts consolidating duplicates first, uses robots.txt only for URLs that can’t be consolidated, and recommends 404 or 410 for pages removed for good.
Can a robots.txt problem cause this?
Yes. If robots.txt returns a 429 or a server error, Google stops crawling for the first 12 hours, then uses the last good copy for up to 30 days. Check robots.txt fetching in the Crawl Stats report’s host status.