Cloudflare Is Caching Your Noindex Tags

On this page

You add noindex at your origin, and Google keeps treating the page as indexable. If Cloudflare is caching your HTML, the directive may be fine and the delivery at fault: the edge keeps serving the old cached page, old robots tag included, until the cache expires or you purge it. A crawl inside that window records the old directive. First check the condition, though. Cloudflare’s documentation on default cache behavior says the Cloudflare CDN does not cache HTML or JSON by default. This problem exists when a cache rule or setting has made your HTML cacheable.

The same trap runs in reverse: you remove a noindex, and the cache keeps serving it. The fix in both directions is to purge the cache as part of the deploy that changes a directive.

Confirm what Googlebot receives

Your browser may show the origin’s fresh response while the edge serves the cached one, so the browser proves nothing. Request the URL with a normal GET and print the response headers:

curl -s -D - -o /dev/null https://example.com/page-you-noindexed/

Read the cf-cache-status value. Cloudflare’s page on cache responses defines them:

  • HIT: the resource was found in Cloudflare’s cache.
  • DYNAMIC: Cloudflare determined the asset isn’t eligible for cache, so the request went to your origin without a cache lookup.

If you see DYNAMIC, Cloudflare isn’t caching this HTML, and the cause lies elsewhere. If you see HIT, fetch the body and check whether the served HTML carries the directive you expect:

curl -s https://example.com/page-you-noindexed/ | grep -i "robots"

A HIT with the old directive in the served HTML proves the cache is serving a stale version. Check the response headers too, since an X-Robots-Tag header is cached along with the page.

The timing trap

The damage comes from sequence. Suppose HTML is cached at the edge for days. You add noindex at the origin, and a crawl shortly afterward is served the cached, indexable page. Google records what it received; your correct directive sits at the origin, hidden behind the cache.

Two delays stack. First the cache has to expire or be purged before the origin’s directive is served at all. Then Google has to come back and crawl the page to read it, and for a page Google revisits at long intervals, that second wait can be long. The first delay is the part you control.

The reverse case follows the same logic. You remove a noindex to bring a page back, the origin serves an indexable response, and the cache keeps serving the old noindexed HTML. The page stays out, and the directive change gets the blame when the cache is the holdout.

The fix workflow

Order the deploy so the cache can’t serve a stale directive into a crawl:

  1. Change the directive at the origin.
  2. Verify it at the origin with the cache bypassed, and confirm the served HTML and headers carry the intended directive.
  3. Purge the affected URLs immediately, as part of the same deploy. Cloudflare’s guide to purging the cache covers the options.
  4. Check cf-cache-status and the served directive again after the purge.

Purging days later is the failure mode. A crawl during the gap may already have recorded the wrong directive, and a late purge only changes what Google sees on its next visit.

For priority pages, request indexing in URL Inspection once the purge is done, so Google’s next fetch meets the corrected response.

When a page must disappear fast

If a page is surfacing in results and has to go quickly, the Search Console Removals tool can hide it, but know what it does. Google’s help for the Removals tool says a temporary removal lasts about six months, that Google can continue to crawl the URL, and that you should make the removal permanent or the page could appear again. Use it as a stopgap while the directive and the purge take effect, not as the fix.

Cache strategy and automation

If you do cache HTML, keep its cache lifetime short, and make the purge part of any deploy that changes a directive, triggered through the Cloudflare API rather than left to memory. A manual purge can get skipped under deadline pressure, and the skip is what can produce the stale crawl.

The same discipline covers canonical tags and redirects, because they live in the cached response too. Change a canonical at the origin while the edge serves the old one, and Google keeps reading the old canonical. Add a redirect while a cached 200 response is still being served, and the crawler doesn’t see the redirect until the cache clears. Treat any change to a robots directive, a canonical or a redirect as a cache event, and purge as part of the deploy.

Frequently asked questions

Does purging the cache deindex the page?

No. Purging makes sure Google receives your current response on its next crawl. The page leaves the index when Google recrawls it, reads the noindex and processes it.

Why does my browser show the noindex while Google ignores it?

Your browser may be getting the origin response or a different cache state. Check cf-cache-status with a curl GET: a HIT with the old directive in the served HTML confirms the edge is serving a stale copy.

Cloudflare doesn’t cache HTML by default. Can this still happen to me?

Only if something makes your HTML cacheable, such as a cache rule. Check cf-cache-status on the page: DYNAMIC means Cloudflare sent the request to your origin without a cache lookup, so the cache isn’t the cause.

Leave a comment

Your email address will not be published. Required fields are marked *