Documentazione
Sfoglia la documentazione

Advanced Bot Detection and Scoring

Set up honeypot pages that only bots can reach, so Privatus Analytics learns what bots on your site look like. With example code, best practices and what not to do.

Visualizza come Markdown

Bots pollute analytics. They inflate pageviews, distort conversion rates and hide what real people do. Simple bots announce themselves and are filtered already. Advanced bots run a real browser and look like people.

Advanced Bot Detection is part of the Business and Enterprise plans. You find it in the site menu.

Where this is going#

We are building a machine learning model that gives every page view and visit a bot score, so you can filter bot traffic out of your reports. The bot score is in development and is not in your reports yet.

Today the feature collects the data the model learns from: hits on honeypots that you set up on your site. Each bot that walks into one shows us how bots behave on your site.

What a honeypot is#

A honeypot is a link to a page on your site that no person can ever click. The link is in the page HTML, so bots that read the HTML find it. It is hidden from view, from the keyboard and from screen readers.

  1. A bot reads your HTML and finds the hidden link.
  2. The bot follows the link and loads the honeypot page.
  3. The tracking script on the honeypot page tells us about the request.
  4. We record what the bot looked like.

The tracking script must be installed on the honeypot page. Without the script we never hear about the visit.

Set up a honeypot#

The Advanced Bot Detection page in the app has the exact code for your site, including the snippet that marks a page as a honeypot. This guide covers the parts around it.

1. Create the honeypot page#

Make a plain page at an address that looks ordinary. Add the honeypot snippet from the app, and a robots meta tag so search engines never index it.

html
<!doctype html>
<html lang="en">
<head>
  <meta charset="utf-8">
  <meta name="robots" content="noindex, nofollow">
  <title>Partner archive</title>
  <!-- The honeypot snippet from the Advanced Bot Detection page -->
</head>
<body>
  <h1>Partner archive</h1>
  <p>This page is no longer updated.</p>
</body>
</html>

Put the link in the HTML of your real pages, for example in the footer or between two sections.

html
<a href="/partners/archive-2019" class="site-extra" rel="nofollow"
   aria-hidden="true" tabindex="-1">Partner archive</a>

aria-hidden="true" keeps it away from screen readers, tabindex="-1" keeps it out of keyboard navigation, and rel="nofollow" tells search engines not to follow it.

Hide it with a class in your CSS file, not with an inline style. Simple bots skip links that carry display:none or the hidden attribute right in the HTML.

css
.site-extra {
  position: absolute;
  left: -9999px;
  width: 1px;
  height: 1px;
  overflow: hidden;
}

4. Tell good crawlers to stay away#

Disallow the page in robots.txt. Search engines and other well-behaved crawlers obey it, so the bots you catch are the ones that ignore the rules.

text
User-agent: *
Disallow: /partners/archive-2019

5. Check it#

Open the honeypot page yourself in a browser. One hit appears in the report within about a minute. That hit is you, so ignore it.

Warning: only mark a page that people cannot reach#

A hit on a honeypot page is stored with its IP address, User-Agent and request headers for 12 months. Never put the honeypot snippet on a real page, in your site-wide layout or in a shared template. Your visitors would be recorded as bots and left out of your statistics.

Best practices#

Do#

  • Use several honeypots in different places, such as the header, the footer and the middle of long pages.
  • Give each one an ordinary address and link text, like an old archive or a partner page.
  • Hide the link with a class in your stylesheet, and add aria-hidden="true", tabindex="-1" and rel="nofollow".
  • Disallow every honeypot page in robots.txt and add a noindex robots meta tag to the page.
  • Keep the honeypot page small and harmless, with a little plain text and nothing else.
  • Change the addresses now and then. Bot authors learn to skip known traps.
  • Mention bot detection in your own privacy policy.

Don't#

  • Don't put the honeypot snippet on a page people can reach, or in a layout that every page shares.
  • Don't list a honeypot page in your sitemap, menus, search results or internal search.
  • Don't use words like honeypot, trap or bot in the address, the link text or the class name.
  • Don't hide the link only with an inline display:none or the hidden attribute.
  • Don't leave the link reachable with the Tab key or readable by a screen reader.
  • Don't put forms, real content, personal data or links back to your site on the honeypot page.
  • Don't block or ban on a single hit. Link checkers, security scanners and browser extensions sometimes follow hidden links too.
  • Don't build a maze of pages that link to each other without end. It wastes your server and proves nothing.

The report#

Part What it shows
Honeypot hits Hits in the chosen period (7, 30 or 90 days, or 12 months)
Automated browsers Hits where the script saw an automation tool driving the browser
No interaction Hits with no pointer, key or touch input
Hits by honeypot page Each honeypot address, with its hits and last hit
What the standard filter saw How the standard bot filter would have judged each hit. Undetected hits are the valuable ones: the standard filter would have counted them as people
Recent hits The latest 50 hits with time, page, country, browser and automation

Honeypot hits are never counted as pageviews or visits, and they do not use your monthly events.

What we store#

For every other request, the IP address and User-Agent are used in memory and never stored. Honeypot hits are the one exception, because no person can reach a honeypot page.

Stored with a honeypot hit Why
IP address, User-Agent and request headers (encrypted) What the bot score learns from
Page, country, browser, operating system, device The report
Automation signal, interaction, screen width, language What the script saw on the page

Cookies and authorization headers are never kept. The IP address, User-Agent and headers are never shown in the app or returned by the API. Everything is deleted after 12 months, or when the site is deleted.

A sentence for your own privacy policy#

We use hidden pages that people cannot reach to detect automated traffic. If automated software requests one of these pages, our analytics provider, Privatus Analytics, stores the IP address, user agent and request headers of that request for up to 12 months to tell bots from people.

Only hits sent by the tracking script count. The no-JavaScript pixel and the server-side events API cannot report a honeypot hit.

API and MCP#

GET /sites/{site_id}/bot-detection.json and the MCP tool bot_detection_report return the same report. Workspaces on other plans get 402 with the code plan_limit.