Google Keeps Crawling URLs You Deleted Years Ago

On this page

Deleting a page doesn’t make Google forget its URL. Google’s crawl budget guide says it plainly: Google won’t forget a URL that it knows about, but a 404 status code is a strong signal not to crawl that URL again. So a “Not found (404)” list full of addresses you removed years ago is expected, and on its own it isn’t a problem. Your job is narrower than clearing the report: stop your own site from feeding Google dead URLs, rescue the ones that still carry value, and avoid the fixes that fall short or backfire.

Why old URLs keep coming back

Google finds URLs through links, and it keeps the ones it has found. Three things keep old addresses in play:

  • Links on other sites. A page that linked to your old article may still link to it. When Google crawls that page, it meets your old URL again. You don’t control those pages.
  • Your own site. Internal links, navigation, old sitemaps and feeds that still list a deleted URL keep pointing Google at it. These you control.
  • Variants. Every tracking-parameter version of an old URL that was ever linked, such as ?utm_source= or ?ref=, is a separate URL, so one deleted page can show up as several URLs.

What Google does with them

Google’s documentation on HTTP status codes describes the handling of 4xx responses:

  • Google doesn’t index URLs that return a 4xx status code, and indexed URLs that start returning one are removed from the index.
  • Newly encountered 404 pages aren’t processed, and the crawling frequency gradually decreases.
  • The 4xx status codes, except 429, have no effect on crawl rate.
  • All 4xx errors, except 429, are treated the same.

That last point settles the old 404-versus-410 debate for Google Search: a 410 doesn’t make Google drop a URL faster than a 404. Use whichever is accurate for your users.

Google’s Page indexing report help lists a 404 for a page that you’ve removed and have no replacement for among the right reasons for a URL not to be indexed. The report is telling you what Google found, not asking you to act on each line.

Three fixes that fall short or backfire

Blocking old URLs in robots.txt. It does stop Google fetching them, but then Google can’t see the 404 either, and removing an indexed URL for its 4xx response depends on fetching it. The crawl budget guide adds that blocked URLs stay part of your crawl queue much longer and are recrawled when the block is removed. Let the URLs return 404.

Using the Removals tool. Google’s Removals help says a successful request lasts only about six months and that blocking a URL doesn’t prevent Google from crawling it, only from showing it in results. It is for urgent hiding, not for making Google stop asking.

Redirecting everything to the home page. Google’s site move guide says not to redirect many old URLs to one irrelevant destination, such as the home page, because it can confuse users and might be treated as a soft 404.

What to do instead

  1. Stop feeding Google dead URLs. Remove deleted URLs from your XML sitemaps and feeds, and crawl your site to find internal links that still point at them. Repoint each link to a live page or remove it.
  2. Rescue the ones with value. A deleted URL that still earns links from other sites, or still gets visits, deserves a permanent redirect in one hop to its real replacement. The rest can stay 404.
  3. Handle a legacy pattern with one rule. If an old permalink structure left a whole family of dead URLs, a single pattern redirect to the matching current URLs clears the family, as long as each old URL has a true counterpart.
  4. Ignore probes. Random paths and exploit attempts that show up in the list have no links to save and no relationship to your content. Leave them returning 404.

Reading the report

Open the “Not found (404)” examples in the Page indexing report and group them by pattern before judging them: an old date-based path, a discontinued product directory, a parameter signature, or random probe strings. The group points to the source. A pattern that matches your own old structure points at internal links, sitemaps or a missing pattern redirect. Scattered one-off URLs with outside links are rescue candidates. Probe strings need nothing.

Then watch the trend, not the count. Because Google keeps URLs it knows, the list may never reach zero. What should change after the cleanup is that Google stops finding dead URLs through your own site.

Frequently asked questions

Do all these 404s hurt my rankings?

Google’s Page indexing help lists a 404 for a removed page with no replacement as a right reason not to be indexed: for that page, it is the expected state, not a fault to fix. Act on the dead URLs your own site still links to and the ones with outside links you want to keep.

Should I return 410 instead of 404 so Google gives up faster?

For Google Search, all 4xx errors except 429 are treated the same, so a 410 doesn’t speed anything up. Choose the code that is accurate for visitors.

Why do URLs I deleted years ago still appear?

Google won’t forget a URL it knows about, and links on other sites can keep leading Google back to it. The crawling frequency for a 404 URL gradually decreases, but the URL stays known.

Leave a comment

Your email address will not be published. Required fields are marked *