1. The same questions, asked of every engine
Every generated prompt is asked of every engine in the run — ChatGPT, Gemini, Perplexity, Claude, Grok and DeepSeek, whichever you have connected — so the per-engine results are directly comparable: same questions, same run, same day.
A single blended score is never the headline, because it hides the finding that matters most: you can be strong in one engine and invisible in another, and those have different fixes.
2. Every rate carries a 95% confidence range
Each headline rate — named, recommended, cited — is shown with a Wilson 95% range and the sample size behind it. "Named in 60%" from 10 prompts and from 100 prompts are not the same claim; the range is what tells them apart, and a wide range means few samples, not a poor result.
The sample counts prompt-and-engine pairs, not raw API calls. Asking one prompt five times is still one prompt's worth of coverage — counting it as five would narrow the range on evidence that does not exist, which is the overclaiming the range was added to prevent.
3. Repeat sampling measures stability, and is reported separately
AI answers are not deterministic: ask the same question twice and a different business can be named. That is a property of the engines, not a defect in the measurement — and any tool reporting one percentage from one run is hiding it.
So each prompt can be asked once, three times or five times. Coverage ("how sure are we of this rate across these questions?") comes from the range across the prompt set. Stability ("would this question answer the same way again?") comes from agreement across repeats of one prompt. They answer different questions and are never collapsed into one number — and a single-sample run shows no stability figure at all, because nothing about repeatability was measured.
Repeats multiply calls on your own AI keys, so the cost is stated on the form in calls before you choose, and scheduled runs never raise it on their own.
4. "Where the answer changed" is listed, as counts
With repeats on, the report names the specific questions whose repeats disagreed — "named you 2 of 5" — rather than folding them into a percentage. A headline stability of 80% says something moved; it does not say what. These are the questions the engine has not made its mind up about, and with a single-sample run whichever answer that run happened to get would have been reported as fact.
5. Engine sets are stamped, and trends refuse to join across them
Every run records which engines it measured. Runs on different engine sets are not drawn on one line: changing engines is a change of method, not a change in visibility. Adding an engine starts a fresh trend; the old history stays under the old set, and you are told so at the moment you add one.
A corollary worth stating plainly: more engines gives a fuller picture, not a sharper one. Precision comes from more prompts per engine. "More providers means more accurate" would be false, so we do not say it.
6. Where a number came from is part of the number
Search positions come from two sources — Google Search Console for sites you have connected, and a scraped results page otherwise — and the two measure subtly different things. Each reading is stamped with its source, and a comparison never mixes them silently.
AI-answer figures are labelled as answer-based. Where a probe is search-grounded, that is stated, because a grounded answer and a model's own recall are different evidence about you.
7. What is excluded, and why it is said out loud
Runs stopped early — a provider out of credits, a bad key, a quota — are excluded from trends rather than scored. A partial run is skewed towards whichever engines answered; a distorted point is worse than a missing one, and the exclusion is stated in the report.
A mention whose position could not be located is excluded from the average position, never scored on a fallback. A rate that predates the detector that would have measured it shows "not measured", never 0%. A finding we could not measure — no AI key, say — appears as a locked row stating what it would tell you, because a hidden row makes a site look healthier than the evidence supports.
The same rule runs through every comparison in the product: unknown is never reported as zero; a sample below the floor is called a first look and says what would make it a measurement; and where our own work is compared against doing nothing, the comparison states that it is observational and which way its biases point.