# Site audit

> A polite crawler audits your site for broken links, missing tags and slow pages, scores it, and ranks every issue by the traffic it touches.

## What the audit does

The site audit crawls your site the way a search engine does. It follows
links from the home page, reads your `robots.txt` and sitemaps, and runs
a list of checks on every page it reaches. When it finishes you get:

- a **score** from 0 to 100,
- the **issues** it found, as errors, warnings and notices,
- what is **new** and what was **fixed** since the audit before,
- and every issue ranked by the **pageviews** of the pages it touches,
  because Privatus Analytics already knows which pages people visit.

Every site is audited by default. The first audit starts when the site
is added. Open **Site audit** in the site's sidebar to see the results.

Audited pages never count as events, and audits keep running while
collection is paused at the monthly limit.

## Schedule and limits

Each plan sets how many pages one audit crawls and how often a site is
audited. The numbers are the **Site audit pages per audit** and
**Site audit interval (days)** rows of the table on the
[Plans](/docs/billing/plans#limits) page.

- The page allowance counts HTML pages on your site. Images, scripts,
  stylesheets and links to other sites are checked too and do not count.
- When a crawl stops at the allowance with pages left, the overview says
  how many pages were found and how many were audited.
- **Run now** starts an audit by hand. On a paid plan you can start one
  every 24 hours. On the Free plan a manual audit takes the place of the
  scheduled one, so there is one audit per interval.
- Only one audit runs per site at a time. **Cancel audit** stops it, and
  nothing of the unfinished crawl is kept.

## Crawler

The audit is run by our own crawler, `PrivatusAuditBot`. It sends this
`User-Agent`:

```
Mozilla/5.0 (compatible; PrivatusAuditBot/1.0; +https://privatusanalytics.com/docs/features/site-audit#crawler)
```

It is built to be gentle with your server:

- It fetches **one page at a time** per site, over one kept-alive
  connection, and waits between requests.
- It reads `robots.txt` first and obeys it, following the rules for
  `PrivatusAuditBot` and then the rules for `*`.
- It honors `Crawl-delay`, up to 10 seconds.
- When your server answers `429` or `503` it backs off, obeys
  `Retry-After`, and gives up politely if the site keeps refusing. That
  audit ends as **Partial** with the reason shown.
- It only requests pages. It never submits forms, never logs in and
  never runs your JavaScript.

To slow the crawler down, add a crawl delay to `robots.txt`:

```
User-agent: PrivatusAuditBot
Crawl-delay: 5
```

To keep it out of part of the site:

```
User-agent: PrivatusAuditBot
Disallow: /internal/
```

To block it completely, use `Disallow: /`. The next audit then fails
with a note that `robots.txt` does not allow the crawler. You can also
turn the audit off in its [settings](#settings).

If a firewall or bot protection blocks the crawler, allow the
`PrivatusAuditBot` user agent there.

## Scope and subdomains

The crawl starts at the site's home page. If the home page redirects
(for example from `example.com` to `www.example.com`), the crawler
adopts the final address. The apex domain and `www` always count as the
same site.

By default the crawler also follows links to **subdomains** of the
site's domain, such as `blog.example.com`. Turn that off in the
settings to stay on one host. Everything else is an external link. An
external link is checked once with a light request to see whether it
still works, and its page is never crawled.

Besides links, the crawler picks up URLs from your sitemaps and from
the pages that got pageviews in your analytics, so pages that nothing
links to are found too.

## What is stored

The audit stores **metadata about your pages**: the URL, status code,
response time and size, title, meta description, first H1, canonical,
robots directives, language, hreflang, structured data types, social
tags, a short list of response headers, word count, link counts and the
links between pages.

It **never stores the text or the HTML of a page**. Content is read in
memory to count words and to detect duplicates with a hash, and then
dropped. Nothing about your visitors is involved.

The pages, links and findings are kept for the two latest finished
audits of a site. Older audits keep their score and their totals per
check, so the trends still have their history.

## Score

The score is a public formula:

```
score = 100 x (1 - (E + 0.3 x W) / N) - 5 x S
```

- `N` is the number of pages crawled on your site.
- `E` is the number of pages with at least one error.
- `W` is the number of pages with warnings and no error.
- `S` is the number of site-level errors (such as an unreachable
  `robots.txt`), counted up to 4, so they cost at most 20 points.

Notices never change the score. A score of 90 or more is **Good**, 70
to 89 **Needs improvement**, and anything lower **Poor**. Muted checks
are left out.

## Overview

The overview shows the latest finished audit:

- the score with its band and the change since the audit before,
- pages crawled, and errors, warnings and notices, each with what is
  new and what was fixed,
- the crawled pages split into healthy, with issues, redirects, broken
  and blocked,
- the score over time,
- the **top issues**, ordered by severity and then by the pageviews of
  the pages they affect,
- a card per category with its own score,
- status codes, click depth and response times,
- and what the crawler read in `robots.txt` and your sitemaps.

While an audit runs, the page shows how many pages were crawled so far
and updates by itself. The results of the audit before stay visible
below. A failed audit shows why it failed.

## Issues

**Issues** lists one row per check that found something: the number of
URLs affected, how many are new, how many were fixed, and the pageviews
those pages got in the 30 days before the audit. Filter by severity,
by category, or to new issues only.

Open an issue to read **why it matters** and **how to fix it**, and to
see every affected URL. Sort the URLs by pageviews to fix the pages
people visit first. **CSV** downloads the list.

## Pages

**Pages** is the explorer of everything the crawler saw. Filter by
status code, crawl state, indexable or not, in the sitemap or not, with
or without issues, click depth, or text in the URL or title. Sort by
pageviews, links in, depth, response time or word count. The
**Resources** and **External links** tabs list the images, scripts,
stylesheets and outbound links that were checked.

Open a page to see every stored field, its issues, the first 100 links
in and out with their anchor text, and where it redirects to.
**Page analytics** opens the same page in your analytics, and the
analytics page detail links back to the audit.

## History and compare

**History** lists every audit of the site with its status, score and
counts. Open a finished audit to see its overview.

**Compare two audits** shows, per check, the URLs affected before and
after and what was added and fixed. For an audit and the one right
before it these are the exact new and fixed URLs. For audits further
apart they are the net change.

## Muting

Some findings are on purpose. **Mute this check** on an issue page
hides the check for the whole site. **Mute** next to a URL hides the
check for that URL only. Muted findings are still recorded, but they
stay out of the counts, and out of the score from the next audit on.
**Muted** lists every mute and lets you remove it.

Muting needs the `content.write` permission.

## Email

When an audit finishes, the workspace's owners and admins who can see
the site get a summary by email: the score, what is new, what was fixed
and the top issues, with a link to the audit.

To stop it for one site, use the unsubscribe link in the email or
**Turn off** under **Audit summary emails** on the audit overview. To
stop it for every site, untick **Site audit summary emails** on your
[account page](/docs/dashboard/account-settings).

## Settings

**Settings** on the audit page, and the **Site audit** tab of the site
settings, hold the same form. Changing it needs the `sites.manage`
permission.

| Setting | What it does |
|---|---|
| **Audit this site on a schedule** | Turns the audit on or off. Off also disables **Run now** |
| **Follow links to subdomains** | Crawls subdomains of the site's domain. On by default |
| **Check links to other sites** | Checks that outbound links still work. On by default |
| **Paths to skip** | Paths the crawler does not fetch, one per line, each starting with a slash |
| **Query parameters to ignore** | Parameters that never make a different page, such as `sort`. Tracking parameters like `utm_source` are always ignored |
| **Extra start URLs** | Full URLs to crawl as well as the home page, for pages nothing links to |

## Checks

Every check the audit runs today:

| Check | Category | Severity |
|---|---|---|
| Canonical tags point to different URLs | Crawlability | Error |
| Canonical in the header and in the tag differ | Crawlability | Error |
| Canonical loop | Crawlability | Error |
| Canonical tag is outside the head | Crawlability | Error |
| Canonical points to a missing page | Crawlability | Error |
| Canonical points to a page with a server error | Crawlability | Error |
| Content type is missing or wrong | Crawlability | Error |
| Head contains an element that belongs in the body | Crawlability | Error |
| Page that could not be reached | Crawlability | Error |
| Page on a host that does not resolve | Crawlability | Error |
| Robots directives contradict each other | Crawlability | Error |
| Robots meta tag is outside the head | Crawlability | Error |
| robots.txt could not be fetched | Crawlability | Error |
| Old AJAX crawling URL in use | Crawlability | Warning |
| Base tag has an invalid URL | Crawlability | Warning |
| More than one base tag | Crawlability | Warning |
| Base tag is outside the head | Crawlability | Warning |
| Canonical chain | Crawlability | Warning |
| Canonical tag is empty or invalid | Crawlability | Warning |
| Canonical points to another site | Crawlability | Warning |
| More than one canonical tag | Crawlability | Warning |
| Canonical points to the other protocol | Crawlability | Warning |
| Canonical points to a page blocked by robots.txt | Crawlability | Warning |
| Canonical points to the homepage | Crawlability | Warning |
| Canonical points to a noindex page | Crawlability | Warning |
| Canonical points to a redirect | Crawlability | Warning |
| Page is canonicalized and noindex at once | Crawlability | Warning |
| Character encoding is not declared | Crawlability | Warning |
| Doctype is missing | Crawlability | Warning |
| Page uses frames | Crawlability | Warning |
| Page is set to nofollow in the X-Robots-Tag header | Crawlability | Warning |
| Page is set to nofollow in a robots meta tag | Crawlability | Warning |
| Page is set to noindex in the X-Robots-Tag header | Crawlability | Warning |
| Page is set to noindex in a robots meta tag | Crawlability | Warning |
| Image, script or stylesheet blocked by robots.txt | Crawlability | Warning |
| More than one robots meta tag | Crawlability | Warning |
| robots.txt has lines that cannot be read | Crawlability | Warning |
| Canonical URL contains a fragment | Crawlability | Notice |
| Canonical is set in both the header and the tag | Crawlability | Notice |
| Canonical URL is missing | Crawlability | Notice |
| Canonical URL is relative | Crawlability | Notice |
| Canonical target has no internal links | Crawlability | Notice |
| Form submits with GET | Crawlability | Notice |
| Nofollow is set in both the meta tag and the header | Crawlability | Notice |
| Noindex is set in both the meta tag and the header | Crawlability | Notice |
| Page blocked by robots.txt | Crawlability | Notice |
| Robots meta tag has an unknown directive | Crawlability | Notice |
| No robots.txt file | Crawlability | Notice |
| Page could not be fetched | Redirects | Error |
| Page timed out | Redirects | Error |
| Redirect to a broken page | Redirects | Error |
| Redirect chain of four or more hops | Redirects | Error |
| Redirect from HTTPS to HTTP | Redirects | Error |
| Redirect loop | Redirects | Error |
| Soft 404 | Redirects | Error |
| Page returns a 4xx error | Redirects | Error |
| Page returns a 5xx error | Redirects | Error |
| Homepage does not redirect http to https | Redirects | Warning |
| Meta refresh tag | Redirects | Warning |
| Redirect chain | Redirects | Warning |
| Temporary redirect | Redirects | Warning |
| Refresh header in the response | Redirects | Warning |
| Site answers on both www and the bare domain | Redirects | Warning |
| Internal link redirected for letter case | Redirects | Notice |
| Internal link redirected for a trailing slash | Redirects | Notice |
| Redirect with no internal links | Redirects | Notice |
| Image, script or stylesheet that redirects | Redirects | Notice |
| Sitemap that cannot be read | Sitemaps | Error |
| Missing page in the sitemap | Sitemaps | Error |
| Page with a server error in the sitemap | Sitemaps | Error |
| Sitemap lastmod dates in the future | Sitemaps | Warning |
| Sitemap lastmod dates that cannot be read | Sitemaps | Warning |
| No XML sitemap found | Sitemaps | Warning |
| Page in the sitemap with no internal links | Sitemaps | Warning |
| Sitemap over the size limit | Sitemaps | Warning |
| Page in the sitemap blocked by robots.txt | Sitemaps | Warning |
| http URL in the sitemap | Sitemaps | Warning |
| Noindex page in the sitemap | Sitemaps | Warning |
| Non-canonical page in the sitemap | Sitemaps | Warning |
| Redirecting URL in the sitemap | Sitemaps | Warning |
| Page in the sitemap that timed out | Sitemaps | Warning |
| Sitemap URLs with no lastmod date | Sitemaps | Notice |
| Sitemap not named in robots.txt | Sitemaps | Notice |
| Indexable page missing from the sitemap | Sitemaps | Notice |
| URL listed in more than one sitemap | Sitemaps | Notice |
| Internal link to a broken page | Internal links | Error |
| Link points to a local address | Internal links | Error |
| Page with only nofollow internal links | Internal links | Warning |
| Link has no anchor text | Internal links | Warning |
| Link to the HTTP version of the site | Internal links | Warning |
| Internal link is nofollow | Internal links | Warning |
| Link URL is malformed | Internal links | Warning |
| Link has no real URL | Internal links | Warning |
| More than 1,000 links on the page | Internal links | Warning |
| Page has no internal links | Internal links | Warning |
| Orphan page | Internal links | Warning |
| Page more than three clicks from the homepage | Internal links | Warning |
| Page with both followed and nofollow internal links | Internal links | Notice |
| Page linked only from non-indexable pages | Internal links | Notice |
| Link has generic anchor text | Internal links | Notice |
| Internal link to a redirect | Internal links | Notice |
| Link URL is over 2,000 characters | Internal links | Notice |
| Page with only one internal link | Internal links | Notice |
| Paginated page with no internal links | Internal links | Notice |
| Broken image from another site | External links | Warning |
| External link to a broken page | External links | Warning |
| External link that redirects to a broken page | External links | Warning |
| Broken script from another site | External links | Warning |
| Broken stylesheet from another site | External links | Warning |
| Link opens a new tab without noopener | External links | Notice |
| External link that refused our crawler | External links | Notice |
| External link is nofollow | External links | Notice |
| Page has no visible text | Content | Error |
| Duplicate page with no canonical | Content | Error |
| Title tag is missing | Content | Error |
| Title tag is outside the head | Content | Error |
| Exact duplicate content | Content | Warning |
| Near duplicate content | Content | Warning |
| Thin content | Content | Warning |
| Duplicate meta description | Content | Warning |
| Meta description is missing | Content | Warning |
| More than one meta description | Content | Warning |
| Meta description is outside the head | Content | Warning |
| H1 heading is missing | Content | Warning |
| Placeholder text on the page | Content | Warning |
| Page uses a legacy plugin | Content | Warning |
| Duplicate title | Content | Warning |
| More than one title tag | Content | Warning |
| Title is too long | Content | Warning |
| Same page at more than one URL | Content | Warning |
| Meta description repeats the title | Content | Notice |
| Meta description is too long | Content | Notice |
| Meta description is too short | Content | Notice |
| Duplicate H1 heading | Content | Notice |
| More than one H1 heading | Content | Notice |
| H1 heading is too long | Content | Notice |
| No H2 headings | Content | Notice |
| Heading levels are skipped | Content | Notice |
| Low text to HTML ratio | Content | Notice |
| Title and H1 are identical | Content | Notice |
| Title is too short | Content | Notice |
| Broken image | Images | Error |
| Image has no alt attribute | Images | Warning |
| Image loaded over HTTP on a secure page | Images | Warning |
| Linked image has no alt text | Images | Warning |
| Image over 100 KB | Images | Warning |
| Image alt text is over 100 characters | Images | Notice |
| Image has no width and height | Images | Notice |
| Image uses a legacy format | Images | Notice |
| Page uses an image map | Images | Notice |
| Images further down the page are not lazy loaded | Images | Notice |
| JSON-LD cannot be parsed | Structured data | Error |
| Open Graph URL differs from the canonical | Structured data | Warning |
| Article structured data is incomplete | Structured data | Warning |
| Breadcrumb structured data is incomplete | Structured data | Warning |
| FAQ structured data is incomplete | Structured data | Warning |
| Organization structured data is incomplete | Structured data | Warning |
| Product structured data is incomplete | Structured data | Warning |
| Structured data has no type | Structured data | Warning |
| Structured data type is not valid | Structured data | Warning |
| No favicon link | Structured data | Notice |
| Open Graph tags are incomplete | Structured data | Notice |
| Open Graph tags are missing | Structured data | Notice |
| No structured data | Structured data | Notice |
| Twitter card tag is missing | Structured data | Notice |
| Hreflang has an invalid language or region code | International | Error |
| Hreflang link is outside the head | International | Error |
| hreflang to a broken page | International | Error |
| Page named under conflicting language codes | International | Warning |
| Hreflang lists a language more than once | International | Warning |
| hreflang with no return tag | International | Warning |
| Hreflang has no self reference | International | Warning |
| Hreflang URL is not absolute | International | Warning |
| Hreflang uses one URL for several languages | International | Warning |
| hreflang to a page blocked by robots.txt | International | Warning |
| hreflang to a noindex page | International | Warning |
| hreflang to a non-canonical page | International | Warning |
| hreflang to a redirect | International | Warning |
| Language and hreflang disagree | International | Warning |
| Language code is not valid | International | Warning |
| Language is not declared | International | Warning |
| Hreflang has no x-default | International | Notice |
| Page reachable only through hreflang | International | Notice |
| Hreflang is set but the language is not | International | Notice |
| Viewport meta tag is missing | Mobile | Error |
| Viewport has a fixed width | Mobile | Warning |
| Viewport does not use the device width | Mobile | Warning |
| Zooming is disabled | Mobile | Warning |
| Separate mobile URL | Mobile | Notice |
| Viewport initial scale is not 1 | Mobile | Notice |
| More than one viewport meta tag | Mobile | Notice |
| Very slow server response | Performance | Error |
| Broken JavaScript file | Performance | Error |
| Broken stylesheet | Performance | Error |
| HTML is over 2 MB | Performance | Warning |
| HTML is not compressed | Performance | Warning |
| Scripts or stylesheets served without compression | Performance | Warning |
| Slow server response | Performance | Warning |
| Layout shifts for real visitors (CLS) | Performance | Warning |
| Slow response to input for real visitors (INP) | Performance | Warning |
| Slow loading for real visitors (LCP) | Performance | Warning |
| More than 100 script and stylesheet files | Performance | Notice |
| Scripts or stylesheets with no browser caching | Performance | Notice |
| Certificate has expired | Security | Error |
| Certificate does not match the domain | Security | Error |
| Form submits to an HTTP address | Security | Error |
| Homepage not served over https | Security | Error |
| Mixed content | Security | Error |
| Page is served over HTTP | Security | Error |
| Password field on an insecure page | Security | Error |
| Certificate expires within 14 days | Security | Warning |
| Strict-Transport-Security header is missing | Security | Warning |
| Outdated TLS version | Security | Warning |
| Content-Security-Policy header is missing | Security | Notice |
| Protocol-relative resource URL | Security | Notice |
| Referrer-Policy header is missing | Security | Notice |
| X-Content-Type-Options header is missing | Security | Notice |
| Page can be framed by other sites | Security | Notice |
| Internal link has tracking parameters | URL structure | Warning |
| URL contains a double slash | URL structure | Warning |
| URL repeats a path segment | URL structure | Warning |
| Session ID in the URL | URL structure | Warning |
| URL contains spaces | URL structure | Warning |
| Internal search results can be indexed | URL structure | Notice |
| URL contains non-ASCII characters | URL structure | Notice |
| URL has an empty query string | URL structure | Notice |
| URL is over 200 characters | URL structure | Notice |
| URL has more than 4 parameters | URL structure | Notice |
| URL contains underscores | URL structure | Notice |
| URL contains uppercase letters | URL structure | Notice |
| AMP page has no canonical | Pagination | Error |
| Paginated page is canonicalized to page one | Pagination | Warning |
| Pagination URL is not linked on the page | Pagination | Warning |
| More than one next or previous page | Pagination | Warning |
| Paginated page is noindex | Pagination | Warning |
| AI crawlers blocked in robots.txt | AI search | Notice |
| Not modified in over 6 months | AI search | Notice |
| llms.txt is not in the expected format | AI search | Notice |
| No llms.txt file | AI search | Notice |
| Very long page | AI search | Notice |
| Little semantic HTML | AI search | Notice |
| Content security policy blocks the Privatus Analytics script | Traffic | Error |
| Broken page that still gets visitors | Traffic | Error |
| High-traffic page with errors | Traffic | Error |
| Page may depend on JavaScript | Traffic | Warning |
| Privatus Analytics script is missing | Traffic | Warning |
| Search or AI crawlers requesting a missing page | Traffic | Warning |
| Orphan page with traffic | Traffic | Warning |
| Non-indexable page with search clicks | Traffic | Warning |
| Page with traffic that the audit did not reach | Traffic | Notice |
| Redirecting page that still gets visitors | Traffic | Notice |

## API and MCP

Everything on these screens is in the JSON API and the MCP server. Add
`.json` to a page's address, or call the tool.

| Tool | Endpoint | What it does |
|---|---|---|
| `site_audit_get` | `GET /sites/{site}/audit` | The overview. `audit_id` picks an older audit |
| `site_audit_history` | `GET /sites/{site}/audit/history` | Every audit of the site |
| `site_audit_run` | `POST /sites/{site}/audit` | Starts an audit now |
| `site_audit_cancel` | `POST /sites/{site}/audit/cancel` | Cancels the running audit |
| `site_audit_compare` | `GET /sites/{site}/audit/compare` | Two audits, per check |
| `site_audit_issues_list` | `GET /sites/{site}/audit/issues` | One row per check |
| `site_audit_issues_get` | `GET /sites/{site}/audit/issues/{check_key}` | One check and its URLs |
| `site_audit_pages_list` | `GET /sites/{site}/audit/pages` | The page explorer |
| `site_audit_pages_get` | `GET /sites/{site}/audit/pages/{id}` | One page in full |
| `site_audit_mutes_list` | `GET /sites/{site}/audit/mutes` | Muted checks |
| `site_audit_mutes_create` | `POST /sites/{site}/audit/mutes` | Mutes a check |
| `site_audit_mutes_delete` | `DELETE /sites/{site}/audit/mutes/{id}` | Removes a mute |
| `site_audit_settings_get` | `GET /sites/{site}/audit/settings` | The settings |
| `site_audit_settings_update` | `PATCH /sites/{site}/audit/settings` | Changes the settings |
| `site_audit_notifications_update` | `PATCH /sites/{site}/audit/notifications` | Your summary email for the site |

Reading needs `analytics.read`. Starting, canceling and muting need
`content.write`. The settings need `sites.manage`. The full schemas are
in the [API reference](/docs/api).
