Quick answer
AI crawler log file analysis means reading your raw server or CDN logs to see which AI bots (GPTBot, ClaudeBot, PerplexityBot and others) request which URLs, how often, and what status code they get back. It's the only way to see this traffic, because Search Console doesn't report it and analytics tags don't fire for bots.
Here's a situation I keep running into. Someone tells me their brand "isn't showing up in ChatGPT" and asks what content to write. Nobody has checked whether ChatGPT's crawler can even get through the front door. Half the time it's hitting a wall of 403s, or a redirect chain, or a page that's blank until JavaScript runs.
Your logs answer that in about ten minutes. This guide walks through how, in the order I'd do it.
1. Why logs, and not your dashboards
Search Console covers Google. GA4 covers people running JavaScript in a browser. AI crawlers are neither. Most of them fetch raw HTML and leave, so no tag fires and no report exists. If you want to know what they did on your site, the server is the only witness.
It also matters more now than it did a couple of years ago. Whether an AI system can fetch, read and later cite your page depends on what happens at request time. That is exactly what a log records. If you're tracking the outcome side, our piece on the AI citation rate metric pairs well with this one: citations are the result, logs show the access that made them possible.
2. The bots you'll actually meet
Not all AI bots do the same job, and that changes how you treat them. Some collect training data, some build a search index, and some fetch a page live because a user just asked a question. Blocking one doesn't block the others.
| User-agent token | Operator | Typical purpose |
|---|---|---|
GPTBot | OpenAI | Training data collection |
OAI-SearchBot | OpenAI | Search index for ChatGPT |
ChatGPT-User | OpenAI | Live fetch on a user's request |
ClaudeBot | Anthropic | Training data collection |
Claude-User, Claude-SearchBot | Anthropic | User-triggered fetch and search |
PerplexityBot | Perplexity | Search index |
CCBot | Common Crawl | Open dataset many models draw on |
Google-Extended and Applebot-Extended are robots.txt controls, not separate crawlers. You won't see them in logs. The fetching is done by Googlebot and Applebot. Bot names change, so check each vendor's docs before you build rules around this list.3. Getting the logs
Where they live depends on your setup.
- Your own VPS or server: Nginx usually writes to
/var/log/nginx/access.log, Apache to/var/log/apache2/access.log. Rotated files are gzipped, so usezcatto include them. - Behind a CDN (Cloudflare, Fastly, CloudFront): the CDN often answers before your origin does, so origin logs miss requests. Use the CDN's log export or analytics instead.
- Managed hosting: many hosts have a log download in the dashboard. If yours doesn't, ask support. It's a routine request.
Pull at least 30 days. One week can be badly skewed by a single crawl burst.
4. Verify before you believe
A user-agent is just text the client chooses to send. Scrapers pretend to be GPTBot all the time because they know sites treat it kindly. So before you draw conclusions, check that the traffic is real.
OpenAI and Perplexity publish IP ranges for their bots, so compare the requesting IP against those lists. Google's method is a reverse DNS lookup followed by a forward lookup, documented in Google's verification guide. Where a vendor publishes no ranges, treat the hits as "claims to be" and weigh them accordingly.
5. Four commands to start with
You don't need a paid tool for a first pass. These assume the common Nginx/Apache combined log format. Swap in your file path.
grep -oEi "GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-User|PerplexityBot|CCBot" access.log | sort | uniq -c | sort -rn
grep -i "GPTBot" access.log | awk '{print $9}' | sort | uniq -c | sort -rngrep -i "ClaudeBot" access.log | awk '{print $7}' | sort | uniq -c | sort -rn | head -25grep -i "PerplexityBot" access.log | grep "robots.txt" | head
For anything bigger than a few hundred MB, load the logs into a proper tool. Our roundup of AI SEO tools and the Claude AI SEO automation guide cover ways to speed up the boring parts.
6. What to look for
A wall of 403s or 429s
Your firewall or bot protection may be blocking bots you want. This is the most common cause of "why don't AI tools mention us". Fix the rule, not the content.
Redirects and 404s
If a bot spends its visits on old URLs that redirect twice or die, it may never reach your best pages. Clean internal links and update sitemaps. The same logic sits behind crawl budget optimisation.
Important pages nobody visits
Compare the top URLs in the logs against the pages you'd want cited. If your pricing, comparison or guide pages are missing, check internal linking and whether they're in the sitemap.
Content that only exists after JavaScript
Many AI crawlers fetch the raw HTML and don't render it. If your key text is injected client-side, they may see an empty shell. Our JavaScript SEO guide shows how to test this.
When I open logs for a new site, I don't start with the totals. I start with the status-code split for each bot. If more than a small slice is anything other than 200, that's where the work is, before any content changes.
7. Fixes, including Core Web Vitals
Once you know what's happening, the fixes are usually plain.
- Set deliberate rules in robots.txt. Decide bot by bot: training, search and user-triggered fetching can each get a different answer. The robots.txt and AI crawlers guide has the syntax.
- Consider an llms.txt file. It's a proposed convention, not a guarantee, and support varies by vendor. Our llms.txt guide is honest about that.
- Serve fast, stable HTML. Slow server responses hurt crawlers just as they hurt visitors, and they show up in your logs as timeouts. Improving TTFB helps LCP, and Core Web Vitals is a visitor-facing benefit on top. Start with the site speed and Core Web Vitals guide, and see web.dev on Web Vitals for the metrics themselves.
- Fix internal links to point at final URLs. No chains, no dead ends.
8. A monthly routine
Fifteen minutes, once a month
- Export the last 30 days of logs
- Count hits per AI bot and compare with last month
- Check the status-code split for each bot
- Verify IPs for any sudden spike
- List top URLs, then compare against pages you want cited
- After any robots.txt or firewall change, re-check within a week
Pair this with Search Console for Google's side and GA4 for human visits, and you'll have all three views of your site's visibility. If you want the wider picture, the technical SEO guide is the place to go next.
Frequently asked questions
Does Google Search Console show AI crawler traffic?
No. Search Console reports on Google's own crawling. GPTBot, ClaudeBot, PerplexityBot and similar bots only show up in your server or CDN logs.
Can I trust the user-agent string to identify an AI bot?
Not on its own. Anyone can send a fake user-agent. Where the vendor publishes IP ranges, check the requesting IP against them before you treat the traffic as genuine.
Does Google Analytics track AI crawlers?
No. Most AI crawlers don't run JavaScript, so the GA4 tag never fires. Logs are the only reliable source.
How often should I review AI crawler logs?
Monthly is enough for most sites. Check sooner after a migration, a robots.txt change, or a big content launch.
Sources & references
- OpenAI: Overview of OpenAI crawlers
- Perplexity: Crawlers
- Google Search Central: Verifying Googlebot
- web.dev: Web Vitals
Decide which bots get in, bot by bot, and write the rules correctly.
Read the guide →Stop crawlers wasting visits on redirects, duplicates and dead URLs.
Read the guide →Faster HTML for people and for crawlers.
Read the guide →What to do on the content side once crawlers can reach you.
Read the guide →