How to Leverage Log File Analysis for SEO Decisions
On this page
Server logs are the fullest record of what crawlers requested from your site, when they include your edge traffic: every URL, response code, timestamp and user agent. Search Console tells you what Google did with your pages and summarizes the crawl side; complete logs show every request. When new content on a large site is crawled slowly, the logs show where the crawling goes instead, and that turns “Google is ignoring my pages” into a distribution you can measure and change.
Logs and Search Console answer different questions
Search Console already separates two of the failure points. Its Page indexing report distinguishes “Discovered – currently not indexed” (Google knows the URL but hasn’t crawled it) from “Crawled – currently not indexed” (Google fetched it and didn’t index it), and URL Inspection shows the last crawl date for any URL you check. What Search Console doesn’t give you is the full request history across thousands of URLs at once. Logs do:
- how often each URL, directory or template was requested;
- what each request returned;
- how long a new URL waited before its first request;
- which Google crawler made the request.
Used together, Search Console tells you the status, and the logs show the crawl pattern around it across the whole site.
Get data you can trust
Make sure the logs contain crawler traffic. Certain CDN and edge setups serve cached responses without logging an origin hit, or keep bot traffic out of the exports you receive. Check that verified Googlebot requests appear at all before you analyze anything. If they don’t, you need origin logs or an edge log export that includes crawlers.
Verify the crawler. A user-agent string can claim anything. Google’s guide to verifying its crawlers walks through a reverse DNS lookup confirmed by a forward lookup for a single request; at log scale, it points to matching addresses against the lists Google publishes.
Separate the three kinds of Google traffic. The same guide sorts Google’s crawlers and fetchers into three categories, each with its own reverse DNS masks and published IP ranges:
| Category | Example | robots.txt |
|---|---|---|
| Common crawlers | Googlebot | always respected for automatic crawls |
| Special-case crawlers | AdsBot | may or may not be respected |
| User-triggered fetchers | Google Site Verifier | ignored, because a user asked for the fetch |
For crawl budget analysis, start with the first category. Counting a user-triggered fetch as crawling inflates the numbers and can hide what Googlebot is doing.
Read the crawl distribution
With verified common-crawler requests isolated, group them and look at where the crawling goes:
- requests per directory or template, to see which sections absorb crawling and which are starved;
- the share of requests on parameter URLs (sort, filter, session, tracking) against clean content URLs;
- the share returning 404 or 410, crawling spent on pages that no longer exist;
- the share returning 301 or 302, crawling spent passing through redirects.
The output is a waste reading: how much crawling lands on URLs you don’t need indexed. The figure is specific to your site, so measure it rather than borrowing someone else’s. On a store with faceted navigation, the parameter share can exceed the content share by a wide margin, and that one comparison can explain slow indexing of new pages.
Check which Googlebot is crawling
The logs separate Googlebot Smartphone from Googlebot Desktop by user agent. Google’s mobile-first indexing guide says Google uses the mobile version of a site’s content, crawled with the smartphone agent, for indexing and ranking. If desktop crawling dominates your logs, find out why; separate mobile URLs and a server that responds differently by device are two places to look.
From logs to decisions
One decision metric is time to first crawl. For a sample of recently published URLs, find the first verified Googlebot request and measure how long after publication it came. Slow or missing first crawls on important new pages, while parameter and error URLs are requested repeatedly, suggest self-inflicted waste.
The fixes follow from the distribution:
- Cut the waste. Block crawl-trap parameter patterns that never need crawling, fix soft 404s, and collapse redirect chains into single hops.
- Improve discovery for starved sections. Link new and important pages from the pages your logs show Googlebot requesting most.
- Keep sitemaps clean. List only indexable, canonical URLs that return 200.
Then rerun the distribution. The content share should rise and the parameter and error share should fall. That before-and-after, measured on actual crawler behavior, is the evidence the analysis exists to produce.
Format, volume and retention
Access logs vary by server: the Common and Combined formats from Apache and Nginx, or a custom field order. A busy site can produce gigabytes a day, so plan the storage and query approach before you start: a log-analysis tool, a hosted pipeline, or parsed lines loaded into a database. Keep the fields the analysis uses (timestamp, URL, status code, user agent and client IP for verification) and drop the rest.
Retention is a constraint to settle early. Time-to-first-crawl analysis needs logs covering the period before a URL was published through its first crawl. A server that rotates logs away after days can’t answer the question, so set retention long enough for your realistic publish-to-crawl window before you need it.
Frequently asked questions
Do small sites need log file analysis?
Not as a first step. Google’s crawl budget guide has its own test for this: if your site doesn’t have a large number of pages that change rapidly, or pages seem to be crawled the same day they are published, the guide isn’t aimed at you. Search Console’s reports are enough there. Logs earn their cost on large or fast-changing sites and on stores with faceted navigation.
Can I use Search Console’s Crawl Stats report instead?
It is a good starting point: totals, response codes, file types, crawl purpose and Googlebot type, with example URLs. It doesn’t give you the full request history for every URL, or custom ratios for your own URL patterns. For those, you need verified logs.
Why verify Googlebot if the user agent already says Googlebot?
Because the user agent is self-reported and easy to copy. Without verification, crawl numbers can include impostors, and any conclusion about Google’s behavior carries the same risk.