If you want to know whether ChatGPT, Claude or Perplexity are reading your site, your analytics tool will not tell you. GA4, and every other script-based tool, only records a visit when a browser runs JavaScript, and AI crawlers fetch the HTML and leave. The only place they show up is your server or CDN logs. On my own site I built a small logger that records every request from a known AI crawler into a database table and shows it in an "AI discovery" report in my admin dashboard. This piece explains the method, what each crawler is for, how far you can trust a user agent, and what I actually do with the data. For the wider strategy on getting cited by AI, see my AI marketing guide; this article is only about measurement.
Why your analytics tool cannot see AI crawlers
Client-side analytics works by loading a script in the visitor's browser. A crawler such as GPTBot or ClaudeBot requests the page, reads the HTML and moves on without executing that script, so no pageview is sent. Even the bots that do render JavaScript are usually filtered out by analytics vendors as known bot traffic, which is the right default for reporting on people but useless when the bots are what you want to see. That leaves three places to look. Raw server access logs, if your host gives you them. CDN or edge logs, which Cloudflare, Fastly and Vercel all expose in some form. Or your own middleware, which is what I use. My site runs on Next.js, and a proxy function sees every page request before it is served. When the user agent matches a known AI crawler, it writes a row with the bot name, vendor, purpose, path and country to a table, after the response has been sent so the visitor never waits. Requests to robots.txt, llms.txt and the sitemap are logged too, which tells me which bots read the discovery files and which go straight to content.
The bots worth logging, grouped by purpose
A raw list of crawler names is not very useful. What matters is why each one is fetching your page, because that changes what a spike means. I group them into five purposes. Training crawlers collect content that may be used to train models: GPTBot for OpenAI, ClaudeBot for Anthropic, plus CCBot from Common Crawl, Bytespider and others. Search index crawlers build the index an assistant searches when it answers: OAI-SearchBot for ChatGPT search, Claude-SearchBot, and PerplexityBot. Live answer fetchers visit a page because a person asked a question right now: ChatGPT-User, Claude-User and Perplexity-User. Search engines matter as well, because Googlebot feeds AI Overviews and AI Mode, and Bingbot feeds Copilot. Finally there are agents, such as ChatGPT Agent, that browse on someone's behalf. OpenAI, Anthropic and Perplexity each document these roles on their own bot pages, linked in the sources below. My list lives in one file with a pattern and a purpose per bot, so the logger, the analytics and the report all agree on the same definitions.
User agent is a claim, the IP address is the proof
Matching on user agent is cheap and catches the honest bots, but anyone can put "GPTBot" in a request header. Google says this plainly in its crawler documentation: user agent strings can be spoofed, so check the IP or hostname as well. The main AI vendors publish their crawler IP ranges as JSON files. OpenAI lists separate files for GPTBot, OAI-SearchBot and ChatGPT-User. Anthropic publishes one for its bots. Perplexity publishes ranges for PerplexityBot and Perplexity-User. For Googlebot, Google describes a reverse DNS lookup followed by a forward lookup, or a check against its published IP lists. The opposite problem also exists. In August 2025 Cloudflare published a report alleging that Perplexity used undeclared crawlers with browser-like user agents to reach sites that had blocked it; Perplexity disputed the report. Either way, the lesson is that user-agent logs undercount some traffic and overcount spoofed traffic. My own logger identifies by user agent only, and the report says so. I treat it as a directional signal and run an IP check before any number goes into a client decision.
What to store for each request
Keep the record small and consistent. For each crawler request I store a timestamp, the bot name, the vendor, the purpose, the path requested and the country header from the edge. That is enough to answer the questions that matter: which bots visit, how often, which pages they read, and whether that changes after you publish or restructure something. If you are doing this from raw logs instead of middleware, add the IP and the response status code, because a bot hitting a stream of 404s or redirects is telling you something about your site, not about AI. Do not log anything about human visitors in the same table, and do not store full request headers you do not need. I also exclude API routes and the admin area from logging, and prune rows older than a little over a year so the table does not grow forever. Write the log after the response, not before. A logging failure should never slow down or break a page, so the insert is wrapped so that it can fail silently and only print an error to the server console.
Reading the data: which signals actually matter
Not all crawler hits are equal. Training crawls tell you your content may end up in a future model, which is slow and indirect. Search index crawls tell you an assistant is keeping a copy of your page available to cite. Live answer fetches are the most interesting, because they happen while a person is asking a question and the assistant decided your page was worth opening. In my report I show these three as separate lines over time, plus a list of the pages AI opened while answering. That list is the closest thing to a citation log you can get from your own data, though it is still not proof that the page was quoted in the final answer. I also keep a list of pages no AI crawler has read in the period. If an important service page never appears there, I check that it is in the sitemap, linked internally and not blocked. Finally, compare crawl activity with real visits from assistants in your first-party analytics, covered in my piece on measuring AI search traffic. Crawls without clicks is a normal pattern for training bots; it is a warning sign for live answer bots.
Turning crawler logs into decisions
Crawler data is only useful if it changes something. Here is what I use it for. First, access: if a search index bot you want, such as OAI-SearchBot, never appears, check robots.txt, your CDN's bot protection and any firewall rules before anything else. Some CDN and hosting dashboards include AI bot blocking settings, and a single toggle can make you invisible to an assistant. Second, coverage: compare the pages bots read with the pages you want cited. If they keep reading old blog posts and ignore the service pages, internal linking and the sitemap need work. Third, freshness: when I update an article, I watch for index and live answer fetches on that path in the following days. Fourth, policy: the purpose breakdown makes the robots.txt conversation with a client concrete, because you can show what blocking training crawlers would and would not change. I cover that policy in detail in my piece on GPTBot versus OAI-SearchBot. Results for my own site are tracked in my dashboard, and I do not publish them as benchmarks, because one consultant's site is not a sample anyone should plan around.
Sources
OpenAI, Overview of OpenAI crawlers: https://developers.openai.com/api/docs/bots. Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler. Perplexity, Perplexity crawlers: https://docs.perplexity.ai/guides/bots. Google, Google's common crawlers: https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers. Google, Verifying Googlebot and other Google crawlers: https://developers.google.com/search/docs/crawling-indexing/verifying-googlebot. Cloudflare, Perplexity is using stealth, undeclared crawlers to evade website no-crawl directives (4 August 2025): https://blog.cloudflare.com/perplexity-is-using-stealth-undeclared-crawlers-to-evade-website-no-crawl-directives/
FAQ
No. GA4 records visits through a JavaScript tag, and AI crawlers fetch the HTML without running it. You need server logs, CDN logs or middleware that inspects the user agent on each request to see them.
Check the requesting IP against the JSON list of ranges OpenAI publishes for that bot. Anthropic and Perplexity publish similar lists, and Google documents a reverse and forward DNS check for Googlebot. A user agent alone can be faked.
Live answer fetchers such as ChatGPT-User, Claude-User and Perplexity-User, because they visit while a person is asking a question. Search index crawlers come next. Training crawls are the least direct signal of visibility.