Log file analysis for SEO: see what Googlebot really crawls

Server logs show every request Googlebot actually makes. How to verify real Googlebot, find crawl errors and wasted crawls, and spot pages Google never visits.

A crawler shows you what could be crawled. Search Console shows you a sample of what Google reports. Your server’s access logs show what actually happened: every request, from every bot, with the exact URL, time and status code.

That makes logs the most direct evidence of how Googlebot treats your site. This guide covers what to look for and how to avoid the most common mistake in log analysis: trusting fake Googlebots.

What can server logs tell you about SEO?

Each line in an access log records one request: the client IP, the time, the requested URL, the status code, the response size and the user agent. Filter those lines to search engine crawlers and you can answer questions no other tool answers directly:

Step 1: get the right logs

Ask your host or developer for the raw access logs covering at least a few weeks; a month is a good start. Common formats are Apache and Nginx logs in the combined or common format, and IIS logs in W3C format. They’re often rotated daily and gzipped, which is fine.

Logs contain IP addresses, so treat them as personal data: keep them only as long as you need them. Crawlens analyses imported logs on your own computer and doesn’t upload them.

Step 2: verify that Googlebot is really Googlebot

Anyone can send a request with a Googlebot user agent, and scrapers and vulnerability scanners do it all the time. Analyse unverified “Googlebot” traffic and you’ll draw conclusions from requests Google never made.

Google’s guide to verifying Googlebot gives the method:

  1. Run a reverse DNS lookup on the IP from your logs.
  2. Check the host name ends in googlebot.com, google.com or googleusercontent.com.
  3. Run a forward DNS lookup on that host name.
  4. Check it resolves back to the original IP.

Google also publishes its crawler IP ranges as JSON files for verification at scale. Bingbot can be verified the same way, with host names ending in search.msn.com.

Crawlens does the reverse-and-forward DNS check on every crawler IP when you import logs, and reports fake Googlebot requests separately so they don’t skew the analysis.

Step 3: look for these four problems

Errors Googlebot keeps hitting

URLs where Googlebot gets 4xx or 5xx responses. Server errors matter most: Google’s crawl budget guide says that when a site responds with server errors or slows down, Google crawls less. For 404s, redirect pages that moved and remove links to pages that are gone.

Important pages Googlebot never crawls

Indexable pages from your crawl that don’t appear in the logs at all during the period. Changes to them aren’t picked up, and new pages may not get indexed. Link to them from strong, frequently crawled pages and include them in your XML sitemap.

Requests for URLs your crawl never found: retired pages, parameter combinations, old campaign URLs. They’re often orphan pages or crawl traps. Redirect retired URLs, link to orphans that should rank, and block endless parameter spaces in robots.txt.

Where crawl activity goes

Group requests by section and status code. If most Googlebot hits land on filters, internal search results or old pagination while key product or article pages get few, your internal linking and URL hygiene are steering crawlers to the wrong places.

Does crawl budget matter for your site?

Probably less than you’ve read. Google’s crawl budget guide is written for large sites with 1 million+ unique pages that change moderately often, or 10,000+ pages that change daily. For most sites, Google crawls everything it wants to.

Logs are still worth reading on smaller sites. Not for “crawl budget”, but because they show errors Googlebot hits, pages it never visits, and old URLs that still attract crawls.

What about AI crawlers?

AI companies’ crawlers announce themselves in the user agent: GPTBot and OAI-SearchBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot, and others. Your logs show which pages they read and how often, which helps you decide what to allow in robots.txt. Crawlens groups them as AI crawlers next to Googlebot and Bingbot.

How Crawlens checks logs

Import Apache, Nginx or IIS logs (plain or .gz, several files at once) into a project with a recent crawl. Crawlens verifies Googlebot and Bingbot by DNS, then runs three checks against the crawl:

Check What it flags
Googlebot gets errors URLs where verified Googlebot received 4xx or 5xx, sorted by hits
Indexable pages Googlebot did not crawl Pages from your crawl with no Googlebot request in the log period
URLs Googlebot crawls that the crawl did not find Successful Googlebot requests to pages your site doesn’t link to

Checklist

Frequently asked questions

What is log file analysis in SEO?

It's the practice of reading your web server's access logs to see which URLs search engine crawlers request, how often, and which status codes they get. Unlike a crawl or Search Console, logs record what bots actually did on your server.

How do I know if a request is really from Googlebot?

Don't trust the user agent alone; it's easy to fake. Google recommends a reverse DNS lookup on the IP, checking that the name ends in googlebot.com, google.com or googleusercontent.com, then a forward DNS lookup to confirm it resolves back to the same IP.

Do small sites need log file analysis?

Crawl budget mainly matters for very large or fast-changing sites; Google's guide targets sites with 1 million+ pages, or 10,000+ pages that change daily. But logs are still useful on smaller sites to find errors Googlebot hits and pages it never visits.

Which log formats can I analyse?

Most servers write Apache or Nginx access logs in the combined or common format, or IIS logs in W3C format. Crawlens reads all of these, plain or gzipped, and several files at once.

Can logs show AI crawlers like GPTBot or ClaudeBot?

Yes. AI crawlers identify themselves in the user agent, so logs show which pages GPTBot, ClaudeBot, PerplexityBot and others request, alongside Googlebot and Bingbot.

· Founder, Crawlens

Dien builds Crawlens, a desktop crawler for technical SEO audits, and writes about the checks it runs: crawling, indexing, JavaScript rendering and Search Console data.

Audit your own site with Crawlens

Crawl, run 70 checks, and see them next to Search Console and Core Web Vitals data.

Download free for Windows