What Is Crawl Budget and Does It Matter for Small Sites?
On this page
Google gives small sites a test of its own. Its crawl budget guide says that if your site doesn’t have a large number of pages that change rapidly, or if your pages seem to be crawled the same day they are published, you don’t need to read the guide. Keeping your sitemap up to date and checking the Page indexing report regularly is enough. A small site that passes that test and still has pages out of the index has a different problem to look at first, such as page value or duplication. There are exceptions to check, and they have nothing to do with page count: a large share of URLs sitting in “Discovered – currently not indexed”, or URL sprawl from filters and parameters.
What crawl budget is
Google defines a site’s crawl budget as the set of URLs it can and wants to crawl. Two things set it.
- Crawl capacity limit. Also called hostload, this limits how much time your server spends holding connections open for Google. Google starts from a conservative default and adjusts it over time if there is demand to crawl more and the site stays healthy. If your site responds consistently and its response times stay stable or improve, the limit goes up. If responses slow down or return server errors, it goes down.
- Crawl demand. How much Google wants to crawl your URLs. The guide names three factors: perceived inventory (without guidance, Google tries to crawl all or most of the URLs it knows about), popularity, and staleness.
Two details change how you read your own site. Even if the capacity limit isn’t reached, Google crawls less when demand is low. And the capacity limit is shared across all of Google’s crawlers, so heavy demand from one, such as the crawler for ads or Shopping, can reduce what is left for the others.
Crawling is also not the same as indexing. The guide says that not every page that is crawled will necessarily be indexed: after crawling, each page is evaluated, consolidated and assessed for its suitability for the index.
Google’s own size estimates
The crawl budget guide names the sites it is written for:
- large sites, 1 million or more unique pages, with content that changes moderately often, about once a week;
- medium or larger sites, 10,000 or more unique pages, with very rapidly changing content, daily;
- sites with a large portion of their URLs classified by Search Console as “Discovered – currently not indexed”.
Google calls the numbers a rough estimate to help you classify your site, and says they are not exact thresholds. Google’s own list includes the 10,000 figure, and it comes with a condition: content that changes daily. A 5,000-page site that updates ten pages a month isn’t the case the guide describes.
When a small site should look at crawling anyway
The third line of that list doesn’t depend on size. Google’s Page indexing report help explains “Discovered – currently not indexed” as a page Google found but hasn’t crawled yet, typically because crawling it was expected to overload the site. On a small site with a large share of pages in that state, start by checking the server.
URL sprawl is the other case that ignores page count. The perceived-inventory factor means Google tries to crawl the URLs it knows about, so a small site whose filters, sort orders, session parameters or calendar pages generate thousands of URLs has a much larger crawlable inventory than its real page count suggests.
When crawling isn’t the issue
“Crawled – currently not indexed” is the opposite case. Google fetched the page. The Page indexing help says such a page may or may not be indexed in the future, and that there is no need to resubmit it for crawling. Requesting more crawling isn’t the lever. What can change it is the page: whether it adds something the rest of your site and the web don’t already have, and whether it duplicates another of your pages closely enough to be consolidated with it.
For a small site, that is content and structure work: make each page earn its place in the index, merge near-duplicates into one stronger page, link to important pages from the pages that already get traffic, and keep the sitemap limited to the URLs you want indexed.
A ten-minute check
Four reads can show whether crawling is a factor at all.
- Crawl Stats. In Search Console, the Crawl Stats report is available only for root-level properties: a Domain property, or a URL-prefix property at the root of the site. Check host status first. It should be green; if it isn’t, the details break availability down into robots.txt availability, DNS resolution and host connectivity. Then look at the average response time over the period.
- Crawl purpose. The same report separates Discovery, a URL Google has never crawled before, from Refresh, a recrawl of a known page. New pages should show up as Discovery requests soon after you publish them.
- Page indexing report. Compare how many pages sit in “Discovered – currently not indexed” with how many sit in “Crawled – currently not indexed”. The first points toward crawling; the second points toward the pages.
- URL Inspection on a new page. The URL Inspection tool shows the last time Google crawled the page. If new pages are crawled the same day you publish them, you have passed Google’s own test.
If host status is green, response times are stable, new pages are crawled promptly and most excluded pages are “Crawled”, crawl budget is not your constraint. The work is on the pages.
Keep the URL inventory clean at any size
At any size, four habits help keep Google’s attention on the pages you care about:
- Stop filter, sort and tracking parameters from generating crawlable duplicates.
- Fix soft 404s. The crawl budget guide says soft 404 pages will continue to be crawled.
- Return a 404 or 410 for permanently removed pages, which the guide calls a strong signal not to crawl that URL again.
- Point redirects straight at the final URL. In Crawl Stats, every hop in a redirect chain shows up as its own request.
This is hygiene, not an emergency. On a clean small site, it is where the crawl budget question ends.
Frequently asked questions
My small site’s pages aren’t indexed. Is crawl budget the cause?
Check which status they are in. If they are “Crawled – currently not indexed”, Google fetched them and crawling isn’t the constraint; improve or consolidate the pages. If a large share are “Discovered – currently not indexed”, check server health in Crawl Stats first.
Is there a page count where crawl budget starts to matter?
Google gives rough estimates, not thresholds: 1 million or more pages changing weekly, or 10,000 or more changing daily. It also includes any site with a large share of URLs in “Discovered – currently not indexed”.
Does a faster server give me more crawl budget?
It can raise the capacity limit. Google says the limit goes up when a site responds consistently and its response times stay stable or improve. But capacity is only half of it: if demand for your URLs is low, Google still crawls less.