# Bot filtering rules

> Bot filtering rules for User-Agents, crawlers and data centers, and how to send server-side events with the visitor's User-Agent and IP so they count.

Every hit, browser or server, goes through the same bot checks before it's
counted. Bots are counted in aggregate (reason, country) with **no IP
stored**, and don't count toward your usage.

## Checks, in order

1. **Empty User-Agent** → `bot_empty_user_agent`.
2. **Known crawlers** (Googlebot, Bingbot, GPTBot and the rest of our UA
   list) → `bot_known_bot`.
3. **Automation and HTTP libraries** → `bot_automation`. User-Agents
   containing any of: `headless`, `phantomjs`, `puppeteer`, `playwright`,
   `selenium`, `webdriver`, `electron`, `slimerjs`, `lighthouse`,
   `pagespeed`, `gtmetrix`, `pingdom`, `uptimerobot`, `monitage`,
   `prerender`, `python`, `curl`, `wget`, `httpclient`, `okhttp`,
   `go-http`, `java/`, `axios`, `node-fetch`, `scrapy`.
4. **Data-center networks** (major cloud and hosting providers, by ASN) →
   `bot_datacenter`. VPN and relay networks such as iCloud Private Relay
   aren't on the list, so real people behind them count.
5. **Referrer spam** domains → `bot_referrer_spam`.

## What this means for server-side events

- Always send the **visitor's** `user_agent`. If you leave it out, or your
  HTTP library's default (`python-requests/2.x`, `node-fetch`, `curl/8`,
  `Go-http-client`…) ends up in `user_agent`, the event is dropped.
- Send the **visitor's** `ip`, or none at all. Your server's IP is often a
  cloud IP and would be filtered as a data center.
- The request's own headers don't matter: the checks use the `user_agent`
  and `ip` fields in the body.

## Visits with no input seen

Some bots run a real browser with an ordinary User-Agent, so the checks
above let them through. They tend to load one page and leave without
touching it. To help you spot them, the tracker reports whether it saw any
real input during a visit: a pointer press or move, a key press, a touch or
a wheel turn. It sends a yes or no only, never what the input was (see
[What the tracker sends](/docs/tracker/payload#interaction-flag)).

Every visit then has one of three values in the `interaction` dimension:

| Value | Meaning |
|---|---|
| `seen` | Input was seen on at least one page of the visit |
| `none` | The tracker could tell, and saw none |
| `unknown` | Nothing can tell: server-side events, the no-JavaScript pixel, a self-hosted copy of an older tracker, and all data from before 1 October 2026 |

These visits are **not dropped**, and they count toward your usage like any
other. A person who opens a page and closes it without touching it also
shows as `none`, so read it as a signal, not as proof.

- **Overview → Technology → Interaction** breaks visits down by the three
  values. Click a row to filter the dashboard by it.
- To leave them out of a report or an API call, filter with
  `["interaction", "is_not", "none"]`. Save it as a
  [segment](/docs/dashboard/filters-and-segments) to reuse it.
- **Site settings → Bots** shows how many visits in the last 30 days had no
  input seen, with links to both views.

## Turning filtering off

**Site settings → Bots** lets you switch bot filtering off for a site, for
diagnostics. Everything is then counted, including crawlers. Turn it back
on when you're done.

The bot summary on the same tab shows the last 30 days by reason and
country, and the Overview's footer shows how many hits were filtered. For AI
crawler visibility (which never runs JavaScript), use the
[AI crawler logs](/docs/features/ai-crawlers).
