# Advanced Bot Detection and Scoring

> Set up honeypot pages that only bots can reach, so Privatus Analytics learns what bots on your site look like. With example code, best practices and what not to do.

Bots pollute analytics. They inflate pageviews, distort conversion rates
and hide what real people do. Simple bots announce themselves and are
filtered already. Advanced bots run a real browser and look like people.

**Advanced Bot Detection** is part of the **Business** and **Enterprise**
plans. You find it in the site menu.

## Where this is going

We are building a machine learning model that gives every page view and
visit a **bot score**, so you can filter bot traffic out of your reports.
The bot score is in development and is not in your reports yet.

Today the feature collects the data the model learns from: hits on
**honeypots** that you set up on your site. Each bot that walks into one
shows us how bots behave on your site.

## What a honeypot is

A honeypot is a link to a page on your site that no person can ever click.
The link is in the page HTML, so bots that read the HTML find it. It is
hidden from view, from the keyboard and from screen readers.

1. A bot reads your HTML and finds the hidden link.
2. The bot follows the link and loads the honeypot page.
3. The tracking script on the honeypot page tells us about the request.
4. We record what the bot looked like.

The tracking script **must** be installed on the honeypot page. Without
the script we never hear about the visit.

## Set up a honeypot

The **Advanced Bot Detection** page in the app has the exact code for
your site, including the snippet that marks a page as a honeypot. This
guide covers the parts around it.

### 1. Create the honeypot page

Make a plain page at an address that looks ordinary. Add the honeypot
snippet from the app, and a robots meta tag so search engines never index
it.

```html
<!doctype html>
<html lang="en">
<head>
  <meta charset="utf-8">
  <meta name="robots" content="noindex, nofollow">
  <title>Partner archive</title>
  <!-- The honeypot snippet from the Advanced Bot Detection page -->
</head>
<body>
  <h1>Partner archive</h1>
  <p>This page is no longer updated.</p>
</body>
</html>
```

### 2. Add the hidden link

Put the link in the HTML of your real pages, for example in the footer or
between two sections.

```html
<a href="/partners/archive-2019" class="site-extra" rel="nofollow"
   aria-hidden="true" tabindex="-1">Partner archive</a>
```

`aria-hidden="true"` keeps it away from screen readers, `tabindex="-1"`
keeps it out of keyboard navigation, and `rel="nofollow"` tells search
engines not to follow it.

### 3. Hide the link in your stylesheet

Hide it with a class in your CSS file, not with an inline style. Simple
bots skip links that carry `display:none` or the `hidden` attribute right
in the HTML.

```css
.site-extra {
  position: absolute;
  left: -9999px;
  width: 1px;
  height: 1px;
  overflow: hidden;
}
```

### 4. Tell good crawlers to stay away

Disallow the page in `robots.txt`. Search engines and other well-behaved
crawlers obey it, so the bots you catch are the ones that ignore the rules.

```text
User-agent: *
Disallow: /partners/archive-2019
```

### 5. Check it

Open the honeypot page yourself in a browser. One hit appears in the
report within about a minute. That hit is you, so ignore it.

## Warning: only mark a page that people cannot reach

A hit on a honeypot page is stored with its IP address, User-Agent and
request headers for 12 months. Never put the honeypot snippet on a real
page, in your site-wide layout or in a shared template. Your visitors
would be recorded as bots and left out of your statistics.

## Best practices

### Do

- Use several honeypots in different places, such as the header, the
  footer and the middle of long pages.
- Give each one an ordinary address and link text, like an old archive or
  a partner page.
- Hide the link with a class in your stylesheet, and add
  `aria-hidden="true"`, `tabindex="-1"` and `rel="nofollow"`.
- Disallow every honeypot page in `robots.txt` and add a `noindex` robots
  meta tag to the page.
- Keep the honeypot page small and harmless, with a little plain text and
  nothing else.
- Change the addresses now and then. Bot authors learn to skip known
  traps.
- Mention bot detection in your own privacy policy.

### Don't

- Don't put the honeypot snippet on a page people can reach, or in a
  layout that every page shares.
- Don't list a honeypot page in your sitemap, menus, search results or
  internal search.
- Don't use words like honeypot, trap or bot in the address, the link
  text or the class name.
- Don't hide the link only with an inline `display:none` or the `hidden`
  attribute.
- Don't leave the link reachable with the Tab key or readable by a screen
  reader.
- Don't put forms, real content, personal data or links back to your site
  on the honeypot page.
- Don't block or ban on a single hit. Link checkers, security scanners
  and browser extensions sometimes follow hidden links too.
- Don't build a maze of pages that link to each other without end. It
  wastes your server and proves nothing.

## The report

| Part | What it shows |
|---|---|
| **Honeypot hits** | Hits in the chosen period (7, 30 or 90 days, or 12 months) |
| **Automated browsers** | Hits where the script saw an automation tool driving the browser |
| **No interaction** | Hits with no pointer, key or touch input |
| **Hits by honeypot page** | Each honeypot address, with its hits and last hit |
| **What the standard filter saw** | How the standard bot filter would have judged each hit. **Undetected** hits are the valuable ones: the standard filter would have counted them as people |
| **Recent hits** | The latest 50 hits with time, page, country, browser and automation |

Honeypot hits are never counted as pageviews or visits, and they do not
use your monthly events.

## What we store

For every other request, the IP address and User-Agent are used in memory
and never stored. Honeypot hits are the one exception, because no person
can reach a honeypot page.

| Stored with a honeypot hit | Why |
|---|---|
| IP address, User-Agent and request headers (encrypted) | What the bot score learns from |
| Page, country, browser, operating system, device | The report |
| Automation signal, interaction, screen width, language | What the script saw on the page |

Cookies and authorization headers are never kept. The IP address,
User-Agent and headers are never shown in the app or returned by the API.
Everything is deleted after 12 months, or when the site is deleted.

### A sentence for your own privacy policy

> We use hidden pages that people cannot reach to detect automated
> traffic. If automated software requests one of these pages, our
> analytics provider, Privatus Analytics, stores the IP address, user
> agent and request headers of that request for up to 12 months to tell
> bots from people.

Only hits sent by the tracking script count. The no-JavaScript pixel and
the server-side events API cannot report a honeypot hit.

## API and MCP

`GET /sites/{site_id}/bot-detection.json` and the MCP tool
`bot_detection_report` return the same report. Workspaces on other plans
get `402` with the code `plan_limit`.
