Docs
Browse the docs

AI crawler tracking from server logs

See which AI crawlers like GPTBot, ClaudeBot and PerplexityBot, and which search bots, read your pages, from server or CDN logs with no IP addresses stored.

View as Markdown

Crawlers don't run JavaScript, so the tracker never sees them. The AI crawlers report is built from your Cloudflare analytics or your server or CDN logs instead.

What you see#

Daily hits per crawler and per path, grouped into:

Category Examples
AI (training, AI search and assistants) GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, Claude-User, PerplexityBot, Google-Extended, Applebot-Extended, Bytespider, CCBot, Amazonbot, Meta-ExternalAgent
Search Googlebot, Bingbot, DuckDuckBot, YandexBot, Baiduspider, Applebot
SEO tools AhrefsBot and similar
Other Any other bot

Use it to decide what to allow in robots.txt, to see which pages AI assistants fetch when answering users, and to spot crawlers ignoring your rules.

AI assistants sending people to your site is different: those visits appear in the AI assistants channel.

Connect Cloudflare#

If your site is behind Cloudflare, this is the quickest way. There is nothing to deploy.

  1. Open the site's AI crawlers page and choose Connect Cloudflare.
  2. Approve read-only access on Cloudflare's own consent page. We ask for two permissions: listing your zones and reading zone analytics.
  3. Pick the zone that serves the site.

The first import brings in up to 30 days, as far back as your Cloudflare plan keeps analytics. After that the report updates every three hours, and Sync now imports straight away. Only requests to the site's own hostname (with and without www) are counted, so other subdomains in the same zone stay out.

Good to know:

  • Cloudflare samples analytics on busy zones, so hit counts are close estimates, not exact log counts.
  • Each day keeps the 10,000 busiest crawler and path pairs.
  • Use Cloudflare sync or log shipping for a site, not both, or the same hits are counted twice.
  • Disconnect removes the hits imported from Cloudflare and revokes our access. Hits from your own logs stay.

Over the API and MCP, crawlers_cloudflare_zones, crawlers_cloudflare_connect, crawlers_cloudflare_sync and crawlers_cloudflare_disconnect do the same once the Cloudflare account has been approved in the browser.

Connect your logs#

Not on Cloudflare, or want exact counts? Send logs yourself.

Logs are sent with the site's server ingest key (Site settings → Tracking → Server ingest key) to the endpoint shown on the site's AI crawlers page. Connectors:

  • Cloudflare: a Worker (or Logpush job) that forwards request logs.
  • Vercel and Netlify: a log drain.
  • Nginx or Caddy: a small log shipper that posts lines in the "combined" log format.

Privacy#

Log lines contain visitors' IP addresses, so we parse them immediately and keep only the day, the path, the status code and the crawler's name. Lines from normal browsers are skipped. The IP, the referrer and any user field are discarded as soon as a line is split and are never stored.

The Cloudflare sync never receives IP addresses at all. Cloudflare sends us request counts grouped by User-Agent and path. We use the User-Agent to name the crawler and store only the name, the path and the daily count.