# AI crawler tracking from server logs

> See which AI crawlers like GPTBot, ClaudeBot and PerplexityBot, and which search bots, read your pages, from server or CDN logs with no IP addresses stored.

Crawlers don't run JavaScript, so the tracker never sees them. The **AI
crawlers** report is built from your **Cloudflare analytics** or your
**server or CDN logs** instead.

## What you see

Daily hits per crawler and per path, grouped into:

| Category | Examples |
|---|---|
| **AI** (training, AI search and assistants) | GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, Claude-User, PerplexityBot, Google-Extended, Applebot-Extended, Bytespider, CCBot, Amazonbot, Meta-ExternalAgent |
| **Search** | Googlebot, Bingbot, DuckDuckBot, YandexBot, Baiduspider, Applebot |
| **SEO tools** | AhrefsBot and similar |
| **Other** | Any other bot |

Use it to decide what to allow in `robots.txt`, to see which pages AI
assistants fetch when answering users, and to spot crawlers ignoring your
rules.

AI assistants sending *people* to your site is different: those visits
appear in the **AI assistants** [channel](/docs/metrics/channels).

## Connect Cloudflare

If your site is behind Cloudflare, this is the quickest way. There is
nothing to deploy.

1. Open the site's **AI crawlers** page and choose **Connect Cloudflare**.
2. Approve read-only access on Cloudflare's own consent page. We ask for
   two permissions: listing your zones and reading zone analytics.
3. Pick the zone that serves the site.

The first import brings in up to 30 days, as far back as your Cloudflare
plan keeps analytics. After that the report updates every three hours, and
**Sync now** imports straight away. Only requests to the site's own
hostname (with and without `www`) are counted, so other subdomains in the
same zone stay out.

Good to know:

- Cloudflare samples analytics on busy zones, so hit counts are close
  estimates, not exact log counts.
- Each day keeps the 10,000 busiest crawler and path pairs.
- Use Cloudflare sync or log shipping for a site, not both, or the same
  hits are counted twice.
- **Disconnect** removes the hits imported from Cloudflare and revokes our
  access. Hits from your own logs stay.

Over the API and MCP, `crawlers_cloudflare_zones`,
`crawlers_cloudflare_connect`, `crawlers_cloudflare_sync` and
`crawlers_cloudflare_disconnect` do the same once the Cloudflare account
has been approved in the browser.

## Connect your logs

Not on Cloudflare, or want exact counts? Send logs yourself.


Logs are sent with the site's **server ingest key** (Site settings →
Tracking → Server ingest key) to the endpoint shown on the site's **AI
crawlers** page. Connectors:

- **Cloudflare:** a Worker (or Logpush job) that forwards request logs.
- **Vercel** and **Netlify:** a log drain.
- **Nginx or Caddy:** a small log shipper that posts lines in the
  "combined" log format.

## Privacy

Log lines contain visitors' IP addresses, so we parse them immediately and
keep only the day, the path, the status code and the crawler's name.
Lines from normal browsers are skipped. The IP, the referrer and any user
field are discarded as soon as a line is split and are never stored.

The Cloudflare sync never receives IP addresses at all. Cloudflare sends
us request counts grouped by User-Agent and path. We use the User-Agent
to name the crawler and store only the name, the path and the daily count.
