Methodology

How we measure AI visibility

AI answers are not deterministic. Ask an engine the same question twice and a different business can be named. So every number we show states what was measured, how much of it, and how confident that makes it — and we refuse to show a figure we did not measure. This page is the method, written down where it can be checked.

Everything below is implemented and covered by tests. Nothing here is planned.

1. The same questions, asked of every engine

Every generated prompt is asked of every engine in the run — ChatGPT, Gemini, Perplexity, Claude, Grok and DeepSeek, whichever you have connected — so the per-engine results are directly comparable: same questions, same run, same day.

A single blended score is never the headline, because it hides the finding that matters most: you can be strong in one engine and invisible in another, and those have different fixes.

2. Every rate carries a 95% confidence range

Each headline rate — named, recommended, cited — is shown with a Wilson 95% range and the sample size behind it. "Named in 60%" from 10 prompts and from 100 prompts are not the same claim; the range is what tells them apart, and a wide range means few samples, not a poor result.

The sample counts prompt-and-engine pairs, not raw API calls. Asking one prompt five times is still one prompt's worth of coverage — counting it as five would narrow the range on evidence that does not exist, which is the overclaiming the range was added to prevent.

3. Repeat sampling measures stability, and is reported separately

AI answers are not deterministic: ask the same question twice and a different business can be named. That is a property of the engines, not a defect in the measurement — and any tool reporting one percentage from one run is hiding it.

So each prompt can be asked once, three times or five times. Coverage ("how sure are we of this rate across these questions?") comes from the range across the prompt set. Stability ("would this question answer the same way again?") comes from agreement across repeats of one prompt. They answer different questions and are never collapsed into one number — and a single-sample run shows no stability figure at all, because nothing about repeatability was measured.

Repeats multiply calls on your own AI keys, so the cost is stated on the form in calls before you choose, and scheduled runs never raise it on their own.

4. "Where the answer changed" is listed, as counts

With repeats on, the report names the specific questions whose repeats disagreed — "named you 2 of 5" — rather than folding them into a percentage. A headline stability of 80% says something moved; it does not say what. These are the questions the engine has not made its mind up about, and with a single-sample run whichever answer that run happened to get would have been reported as fact.

5. Engine sets are stamped, and trends refuse to join across them

Every run records which engines it measured. Runs on different engine sets are not drawn on one line: changing engines is a change of method, not a change in visibility. Adding an engine starts a fresh trend; the old history stays under the old set, and you are told so at the moment you add one.

A corollary worth stating plainly: more engines gives a fuller picture, not a sharper one. Precision comes from more prompts per engine. "More providers means more accurate" would be false, so we do not say it.

6. Where a number came from is part of the number

Search positions come from two sources — Google Search Console for sites you have connected, and a scraped results page otherwise — and the two measure subtly different things. Each reading is stamped with its source, and a comparison never mixes them silently.

AI-answer figures are labelled as answer-based. Where a probe is search-grounded, that is stated, because a grounded answer and a model's own recall are different evidence about you.

7. What is excluded, and why it is said out loud

Runs stopped early — a provider out of credits, a bad key, a quota — are excluded from trends rather than scored. A partial run is skewed towards whichever engines answered; a distorted point is worse than a missing one, and the exclusion is stated in the report.

A mention whose position could not be located is excluded from the average position, never scored on a fallback. A rate that predates the detector that would have measured it shows "not measured", never 0%. A finding we could not measure — no AI key, say — appears as a locked row stating what it would tell you, because a hidden row makes a site look healthier than the evidence supports.

The same rule runs through every comparison in the product: unknown is never reported as zero; a sample below the floor is called a first look and says what would make it a measurement; and where our own work is compared against doing nothing, the comparison states that it is observational and which way its biases point.

What these numbers do not claim

Stated plainly, because the gap is real and a summary of this page that leaves it out would be wrong.

  • It measures what the engines say when asked, not how often real people see you. There is no impression or traffic data behind these numbers.
  • A clean readiness result is what our checks can see from outside. It is not a promise about how AI answers describe you.
  • We cannot tell "content the engine retrieved and ignored" from "content it never retrieved", because we have no crawler-log verification. Citation measurement is answer-based, and it says so.
  • No tool can promise citations or timescales. Results depend on your market and your competition; we measure, we do not guarantee.
  • Emerging standards (llms.txt and the agent-readiness checks) are scored separately and labelled experimental, because no major engine has confirmed consuming them. They never lift the headline.

Where this shows up in the product

  • • The AI Visibility tracker: per-engine rates with ranges and sample sizes, a repeat-sampling option with its cost stated, and "where the answer changed" listed by question.
  • • The site check and audits: a locked row for anything we could not measure, in place, saying what it would tell you.
  • • The Evidence Engine: every closed gap links to the page that closed it, and the next run re-asks the question — so "it worked" is measured, not assumed.
  • • Admin effectiveness: our own work compared against campaigns that lay dormant, across accounts, with the observational caveats printed beside the result.