Help & Guides

Step-by-step guides for every tool, recipe, and setting

Connected AccountsWordPress plugin — who fetches your pages

WordPress plugin — who fetches your pages

What this shows you

If your site runs on WordPress, our plugin records visits from the AI crawlers on our list - GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, PerplexityBot, Googlebot, Bingbot and the others - and sends them to your project. They appear in "AI crawlers on your site" on the project's overview: which crawler fetched which page we published, when, and where each page stands on the way from published to cited. It is included on every plan, the trial included. The Cloudflare read stays on Growth and above; the plugin is the way to see crawlers on a WordPress site without Cloudflare.

Installing it

1. Open the project for the site, go to its Settings tab, and find "AI crawlers: Connect WordPress". 2. Press "Download the plugin" to get the zip. 3. Press "Make the key" and copy it. We keep only a fingerprint of the key, so it is shown once; if you lose it, make a new one (the old one stops working). 4. In WordPress: Plugins → Add New → Upload Plugin, choose the zip, install and activate. 5. In WordPress: Settings → AI crawler log, paste the key and press Connect. Connect checks the key and that the plugin is on the site the project is for: a key works for one project's site and nothing else. The plugin sends nothing anywhere until you press Connect.

What it records

Only requests whose user agent names a crawler on our list. The list comes from us when you press Connect and once a day after, so the plugin and our robots.txt advice always use the same crawlers. For each request it records: • the crawler's user agent, • the page path - never the query string, • the status code your site answered with, and the time, • the address the request came from, used only to check it (below). Every 15 minutes, or sooner when 200 visits are waiting, the plugin sends them to us on WP-Cron, never while a page is loading. If we cannot be reached it keeps them and tries again later.

What it never records

Anything about people. Visits from anyone whose user agent is not a crawler on the list are ignored and never stored. No cookies, no form data, no page content, no query strings. The crawler's address is used to check it and then dropped: we do not keep it, and the plugin deletes it from your site as soon as we have taken the visit.

Verified, unverified, cannot verify

A user agent can be faked, so each visit is checked against the addresses the crawler's operator publishes: Google's lists for Googlebot and Google's other crawlers, OpenAI's for GPTBot, OAI-SearchBot and ChatGPT-User, and Perplexity's for PerplexityBot and Perplexity-User. We fetch those lists from the operators and refresh them daily. • Verified - the address is on the operator's list. • Unverified - the operator publishes a list and the address is not on it: usually someone else using the crawler's name. • Cannot verify - we hold no list for that crawler (Anthropic publishes none, so Claude's visits are always here; we do not check Microsoft's, Apple's, Meta's or the other crawlers' yet), or the list could not be read, or the address could not be seen. Unverified visits are shown, not hidden, and still count in the report's figures; the table beside them says how many there were for each crawler.

The page-cache limit: why counts say 'at least'

A page cache can answer a request without starting WordPress at all, and then the plugin never hears of the visit. That includes caching plugins such as WP Rocket, W3 Total Cache, WP Super Cache and LiteSpeed Cache, your host's cache (Varnish or Nginx on many managed hosts), and Cloudflare when it caches your pages. The plugin looks for the caches it knows - by the plugins and hosts it recognises and by the cache headers on your own home page - and tells us what it found. When it finds one, the report says "Your cache may hide some crawler visits" and every count reads "at least". Finding nothing does not prove there is no cache: a cache at your host that the site cannot see from inside is possible. We have not yet measured how much a cache hides on real sites, so we do not promise the plugin sees every crawl. If your site is also on Cloudflare, connecting Cloudflare (Growth and above) gives the fuller count, and the report uses it for every day it covers. Both counts are kept and shown side by side - never added, since they count the same visits from two places - and where Cloudflare's is the larger, the report says by how much: most likely visits answered without WordPress running, usually by a page cache. The Verified and Unverified split and the first and last visit times still come from the plugin.

Stopping it

Press Disconnect on the plugin's settings screen (it stops recording and deletes visits waiting to be sent), or press "Revoke the key" on the project's Settings tab - the plugin's next send is refused and it stops recording. Deleting the plugin removes its table and its settings from your site.