Methodology

How we measure AI visibility

AI answers are not deterministic. Ask an engine the same question twice and a different business can be named. So every number we show states what was measured, how much of it, and how confident that makes it — and we refuse to show a figure we did not measure. This page is the method, written down where it can be checked.

Everything below is implemented and covered by tests. Nothing here is planned.

1. The same questions, asked of every engine

Every generated prompt is asked of every engine in the run — ChatGPT, Gemini, Perplexity, Claude, Grok and DeepSeek, whichever you have connected — so the per-engine results are directly comparable: same questions, same run, same day.

A composite score is shown, but never on its own: each engine's own rates sit beside it, because the finding that matters most is that you can be strong in one engine and invisible in another, and those have different fixes. How the composite is made is set out, with a worked example, in section 11.

2. Every rate carries a 95% confidence range

Each headline rate — named, recommended, cited — is shown with a Wilson 95% range and the sample size behind it. "Named in 60%" from 10 prompts and from 100 prompts are not the same claim; the range is what tells them apart, and a wide range means few samples, not a poor result.

The sample counts prompt-and-engine pairs, not raw API calls. Asking one prompt five times is still one prompt's worth of coverage — counting it as five would narrow the range on evidence that does not exist, which is the overclaiming the range was added to prevent.

3. Repeat sampling measures stability, and is reported separately

AI answers are not deterministic: ask the same question twice and a different business can be named. That is a property of the engines, not a defect in the measurement — and any tool reporting one percentage from one run is hiding it.

So each prompt can be asked once, three times or five times. Coverage ("how sure are we of this rate across these questions?") comes from the range across the prompt set. Stability ("would this question answer the same way again?") comes from agreement across repeats of one prompt. They answer different questions and are never collapsed into one number — and a single-sample run shows no stability figure at all, because nothing about repeatability was measured.

Repeats multiply calls on your own AI keys, so the cost is stated on the form in calls before you choose, and scheduled runs never raise it on their own.

4. "Where the answer changed" is listed, as counts

With repeats on, the report names the specific questions whose repeats disagreed — "named you 2 of 5" — rather than folding them into a percentage. A headline stability of 80% says something moved; it does not say what. These are the questions the engine has not made its mind up about, and with a single-sample run whichever answer that run happened to get would have been reported as fact.

5. Engine sets are stamped, and trends refuse to join across them

Every run records which engines it measured. Runs on different engine sets are not drawn on one line: changing engines is a change of method, not a change in visibility. Adding an engine starts a fresh trend; the old history stays under the old set, and you are told so at the moment you add one.

A corollary worth stating plainly: more engines gives a fuller picture, not a sharper one. Precision comes from more prompts per engine. "More providers means more accurate" would be false, so we do not say it.

6. Where a number came from is part of the number

Search positions come from two sources — Google Search Console for sites you have connected, and a scraped results page otherwise — and the two measure subtly different things. Each reading is stamped with its source, and a comparison never mixes them silently.

AI-answer figures are labelled as answer-based. Where a probe is search-grounded, that is stated, because a grounded answer and a model's own recall are different evidence about you.

7. Google AI Overviews are measured — as a separate surface, never blended in

The AI Visibility score above comes from asking the engines directly. Google's AI Overview is a different thing entirely: a live results page, generated at the moment of the search, with its own retrieval and its own sources. Both are worth knowing; averaging them would produce a number describing neither.

So they are kept apart in the code, not merely in the wording. A trend refuses to join readings from different sources, which is why you will never see an AI Overview result inside an AI Visibility score, or a single "AI score" that quietly contains both.

Each watched search is checked weekly, or daily where an account chooses that. An AI Overview is generated, so two looks on the same day can genuinely differ — we saw exactly that on 22 August 2026, when one search returned a detailed answer citing the customer twice and, minutes later, returned nothing useful at all. A rate across several checks says something; a single reading does not, so we report counts with the sample size rather than a percentage.

That an answer varies is an argument for MORE readings, not fewer, so daily gives a firmer rate — and what it costs is list length, since the same budget buys roughly a seventh as many searches. We say that plainly rather than presenting one cadence as the correct one. A search where no overview appears is re-checked monthly whatever cadence is chosen, because there is nothing to sample.

Every watched search keeps a source history: which sites the overview drew on, in how many readings out of how many, and which of them are yours. Nothing is called a site it "keeps returning to" on fewer than three readings, and the churn caveat travels with the list — we put no figure on how fast sources turn over, because the published measurements disagree sharply depending on the window, and most week-to-week movement is the surface rather than the site being measured.

Presence and citation are reported as two facts, because they need different responses: "Google shows no AI Overview for this search" and "it shows one and does not name you" are not the same problem. Where an overview appears, the sources it drew on instead are listed — that is the answer to "why them and not us".

Readings that failed, and readings where Google reinterpreted the search into a different one, are set aside and counted rather than averaged in. An error is not an absence, and a different question is not another sample of the same one.

These checks run on your own DataForSEO account, and only for searches you explicitly chose to watch. There is no default that starts spending on your behalf, and nothing is checked that is not on your list.

What we still do not measure: Google AI Mode, and the ChatGPT app. Those are separate products with their own retrieval, and we do not claim them.

8. Which searches are worth watching is answered from your own data, or not at all

A rigorously measured answer about a search nobody performs is still commercially worthless, so something has to say which searches are worth the money. We buy no keyword-volume database and we estimate no volume — a number invented for a phrase and presented as a finding is exactly the kind of confidence this page exists to refuse.

What we use instead is better, where you have it: your own Search Console. Those are not estimates of what a phrase might get — they are the searches people actually ran to reach YOUR site, with real impressions, click-through and average position.

That gives two situations with two different justifications, and they are never averaged into one score. Where you already appear, Search Console says what the demand is — and a search with real impressions where you sit on page one and almost nobody clicks is the fingerprint of an AI Overview answering on your behalf. That is the strongest recommendation this product makes about anything, and the check itself is the test of it.

Where you are invisible there is no Search Console row at all, because you never appeared. There the standing evidence is that a competitor holds page one: somebody in your market is investing in that search. That demonstrates commercial interest, not demand, and it is labelled as the weaker of the two rather than quietly counted as equal.

You are shown the numbers, not a score. "3,400 impressions, 0.6% clicked, you rank fourth" can be checked and argued with; a priority score of 87 has to be taken on trust. A blended priority number would be the same mistake as a blended AI score, made somewhere smaller.

Marking a topic important does not make it recommended. That is your opinion, and handing it back as a finding would be circular — it earns a mention and it breaks ties, and it never moves the verdict.

Where Search Console is not connected, you are told the evidence could not be reached rather than that none was found. Those read very differently to somebody deciding whether to trust a list, and only one of them is true.

Note what this is NOT: no AI writes these recommendations. The inputs are numbers and the rule states in a sentence, so a model would add cost, latency and the occasional confident justification that is not in the data. Arithmetic is the honest tool for arithmetic.

9. What is excluded, and why it is said out loud

Runs stopped early — a provider out of credits, a bad key, a quota — are excluded from trends rather than scored. A partial run is skewed towards whichever engines answered; a distorted point is worse than a missing one, and the exclusion is stated in the report.

A mention whose position could not be located is excluded from the average position, never scored on a fallback. A rate that predates the detector that would have measured it shows "not measured", never 0%. A finding we could not measure — no AI key, say — appears as a locked row stating what it would tell you, because a hidden row makes a site look healthier than the evidence supports.

The same rule runs through every comparison in the product: unknown is never reported as zero; a sample below the floor is called a first look and says what would make it a measurement; and where our own work is compared against doing nothing, the comparison states that it is observational and which way its biases point.

10. Does each page answer what it is shown for? Graded per page, against a real search

Whether an engine cites a page depends on retrieval and popularity nobody controls, so "we published it and you got cited" was never a claim we could stand behind. What we can measure is the step before that, which we do control: does the page contain material that answers the question it is shown for. Where grading is switched on for your account, every page published to your own site is put to two or three AI models from different companies, with web search switched off, and each is asked to quote the passage that answers. Every quote is checked against the page's own text here. A model that agrees to be agreeable produces no evidence, because a quote that is not in the page does not count.

The question matters more than the graders, and it comes from Google. Where Search Console reports a real search that showed your page, that search is the question — a person typed it, the writer never saw it, and the page can genuinely fail it. That is an observed-question assessment: it carries a verdict (Answers, Does not answer, or Contested when the models disagree), the number of graders behind it, and the Search Console row — "shown 312 times, average position 8.7 — no grader found an answer" is worth more than a tick.

Where there is no such search yet, the page is asked about its own subject instead. That is a topic diagnostic, and it carries no verdict, deliberately: a page is written to answer its own subject, so passing that question says almost nothing. What it does carry is useful — the lines on the page a writer would have reason to quote (each one checked against the page), the kind of source the models think the question expects, and what they say is missing. The card says why there was no search, and "we could not read Search Console" is a different reason from "Google reported no search in the last 90 days", because only one of them is about the page.

The two are never added together. The headline counts pages analysed and then splits them; there is no "answerability rate", and no field the two could be pooled from. A shape badge on a diagnostic means the majority of the models, reading the page afterwards, judged it the right kind of source for its topic. It is their opinion, shown in grey and labelled as one, and never a pass: asked again about the same page they split more often than not.

What is checked and what is opinion is marked on the card. Quotes and evidence lines are matched to the page's text. The expected kind of source, the shape gap and what is missing are the models' own words, and are labelled as such. Asking the graders those shape questions in the same call as the verdict is tested with a paired run whenever the prompt changes: each document goes to each grader three times — with the questions, without them, and with them again as the noise floor. On 2 September 2026, over 37 documents and three graders (Anthropic, Gemini and OpenAI via OpenRouter), removing the questions flipped no verdict: 0 of 30, 0 of 31 and 0 of 35, against a noise floor of 0 flips between identical prompts. The same graders passed none of 15 pages that only told them what to answer and none of 15 that only repeated the question's words.

It costs what it says, which is why it is off until you switch it on for your account (since 15 September 2026). Grading runs on your own AI keys: two or three calls per page, once, nightly, newest pages first. A page is graded again only when something changes — the text extractor, the question, or the page itself — never re-rolled for a better answer. A weekly check that the verified passage is still on the page is a single page read and no AI at all, and it runs whether grading is on or not.

Shape-first reverses the order, and since 15 September 2026 it does so for every project: before anything is written for a topic, the panel is asked what kind of content gets cited for it, what it must contain, and where else that shape belongs. The page is written to that contract and then checked element by element, each proven by a quote found in the page. The card leads with the contract - "written as a Comparison, carries 4 of 5" - and the later opinion of the graders follows it, as an opinion. It costs two or three model calls per topic, once, on your own keys; an account with fewer than two AI providers connected has no panel to ask, and its pages are written as before.

Where the models say a page is not the right shape, you can rebuild it from their brief. The rewrite keeps the verified lines word for word, may add no facts, and may link only sources already on the page or recorded against your business — an invented reference is worse than none. On a connected WordPress site the post is updated in place at the same address; elsewhere the rebuilt page is held for you to copy across. Then it is graded again, so before and after are two readings rather than a hope.

11. What makes up the visibility score - a worked example

Take a made-up business, Harbour Plumbing, checked weekly with 10 questions on 3 engines: 30 answers a run. Four of its topics are each asked two ways - once without its name ("Who can fix a leaking boiler in Leeds?") and once with it ("Is Harbour Plumbing a good choice for boiler repair in Leeds?") - plus two questions about the business itself. So 12 answers are to questions without its name, and 18 to questions with it.

The answer score (0-100) weighs five things: named 35, cited 25, positive tone 10, named when the question left the name out 20, and how early the name comes 10 (first 1.0, second 0.8, third 0.6). Harbour Plumbing was named in 17 of 30 answers (57%), its site cited in 9 (30%), spoken of positively in 15 (50%), named in 1 of the 12 answers to questions without its name (8%), and on average named second (0.8): 19.8 + 7.5 + 5.0 + 1.7 + 8.0 = 42.

The visibility score adds the site: 40% the answer score, 20% whether AI crawlers can read the page, 20% machine legibility (structured data, headings, answer shape), 15% whether robots.txt lets AI crawlers in, and 5% llms.txt. With 80, 70, 100 and no llms.txt: 16.8 + 16.0 + 14.0 + 15.0 + 0 = 62, a C (A from 85, B from 70, C from 55, D from 40).

Now the part a single number hides. The next week the run reaches other topics, and six of its ten questions leave the name out instead of four. Harbour Plumbing is exactly as visible as it was - named in nearly every answer to a question that names it, and in hardly any that do not - but now 12 of 30 answers name it (40%), 7 cite it, and 1 of 18 unnamed answers finds it. The answer score falls to 34 and the visibility score to 59. Nothing changed but the questions. A model update, an engine that fails to answer, or plain luck - each question is asked once - move it the same way.

So read the score as a composite of your site and one run's answers, and read the parts beside it for the facts: per engine, how many answers named you when the question left your name out, out of how many; how many named you when it did; how many cited you - each with its range. A change between two runs means something only when the ranges do not overlap.

Since 30 September 2026 the 20 points for being named without your name in the question replaced 20 points for answers to comparison questions, which no run had asked since 11 August and so could not be earned. A run that asks no question without your name leaves that part out rather than scoring it zero. Runs before the change keep the score they had on the day.

Since 2 October 2026 the latest run of a site can be scored again after you fix the site, without asking the engines again: the site is read again and the 60% that comes from it is worked out afresh, while the answers - the other 40% - are the run's own and stay as they were. The run shows when that was done and the score it had on the day it was made. An earlier run is never re-scored.

What these numbers do not claim

Stated plainly, because the gap is real and a summary of this page that leaves it out would be wrong.

  • Its AI-answer numbers are what the engines say when asked, not how often real people see you. The only impression counts here come from Google: Search Console for search, and the AI Overview and AI Mode figures in the Search Console export you upload. No AI engine reports how often a person saw an answer naming you, and apart from Search Console nothing here counts visits to your site.
  • A clean readiness result is what our checks can see from outside. It is not a promise about how AI answers describe you.
  • We cannot tell "content the engine retrieved and ignored" from "content it never retrieved", because we read crawler logs only where you connect your own Cloudflare zone. Citation measurement is answer-based, and it says so.
  • No tool can promise citations or timescales. Results depend on your market and your competition; we measure, we do not guarantee.
  • A page that answers its question is not a page an engine will surface. Answerability says nothing about retrieval, and it is never averaged with visibility or plotted beside it.
  • Emerging standards (llms.txt and the agent-readiness checks) are scored separately and labelled experimental, because no major engine has confirmed consuming them. They never lift the headline.

Where this shows up in the product

  • • The AI Visibility tracker: per-engine rates with ranges and sample sizes, a repeat-sampling option with its cost stated, and "where the answer changed" listed by question.
  • • The site check and audits: a locked row for anything we could not measure, in place, saying what it would tell you.
  • • The Evidence Engine: every closed gap links to the page that closed it. Whether that page answers the question it is shown for is graded separately, against real Search Console queries; whether an engine later cites it is measured on later visibility runs and reported as its own number. Publishing is never counted as the result.
  • • Admin effectiveness: our own work compared against campaigns that lay dormant, across accounts, with the observational caveats printed beside the result.