Why a single answer proves nothing
Ask an assistant the same commercial question ten times and you will not get ten identical answers. The sources cited shift, the brands named change, and the ordering moves. A screenshot of a model naming your client is therefore a single sample from a distribution nobody has measured — it demonstrates possibility, not position.
This is the root of why so much generative visibility marketing is unfalsifiable. A claim built on one favourable response cannot be checked, cannot be reproduced, and cannot be trended. The remedy is not better screenshots but a sampling method, which is the entire subject of this piece.
A useful test when someone presents AI visibility results: ask how many prompts, how often, on which engines, and under what scoring rules. If those four answers are not immediately available, the finding is an anecdote wearing the clothes of research.
Building the panel
A panel is a fixed set of commercial prompts. Ours typically run to between one and two hundred, which is large enough that run-to-run variance averages out and small enough to execute weekly across several engines without the cost becoming the point.
The prompts must come from how buyers actually speak, and this is where most panels go wrong. Keyword lists reworded as questions produce phrasing nobody uses with an assistant. The raw material is sales call recordings, support tickets and customer interviews — sources where you can hear the actual sentence a buyer would type.
Panels should span the buying journey rather than clustering at the bottom. Category orientation prompts, comparison prompts, requirement-specific prompts and vendor-name prompts each behave differently, and a panel weighted entirely toward the last will show flattering numbers that do not correspond to being discovered by anyone new.
Running cold
Personalisation, memory and conversation history all bias assistant answers toward what the account has recently engaged with. Run a panel from the browser you have been researching a client in, and you will measure your own history rather than the market. This is the most common methodological failure we encounter, and it always flatters.
Cold execution means fresh sessions, no signed-in account carrying memory, no prior turns in the conversation, and ideally automation so that a human's browsing does not contaminate anything. Geographic consistency matters too, since answers vary by inferred location.
Frequency should be regular and boring. Weekly works for most categories. Daily generates noise that invites over-interpretation; monthly leaves too much time for a change to be attributed to something other than what actually caused it.
Scoring rules, written before you look
Scoring must be defined in writing before the first run, because after the fact it is remarkably easy to decide that an ambiguous mention counts. We score three things separately: whether the brand is named at all, whether it is cited with a link, and whether the mention is favourable, neutral or a caveat.
Separating those matters because they move independently. A brand can gain mentions while the sentiment shifts from recommendation to qualified mention, which is a real change that a single blended score would hide entirely.
Competitor sets should be fixed at the same time. Share of voice against a moving competitor list is not a measurement, and the temptation to drop a competitor who is doing well is stronger than most people expect.
What the output actually supports
A properly run panel supports statements about direction and proportion: the share of category prompts naming you moved from one level to another over a period, against a fixed competitor set, on named engines. That is a defensible claim and it is genuinely useful for deciding where to invest.
It does not support attribution to a single action, because too many variables move at once. It also does not support precise claims about individual prompts, since those are exactly the level at which variance dominates. Reporting that respects those limits is more credible than reporting that does not, and clients notice.
The reason to publish the method rather than only the numbers is that a study whose methodology is hidden cannot be evaluated. If we are going to tell clients that unfalsifiable claims are the problem with this category, the least we can do is show our own working.
Takeaways
- A single assistant answer is one sample from an unmeasured distribution — it cannot support a claim about position.
- Panels of 100–200 prompts, built from recorded customer language rather than reworded keyword lists.
- Run cold: fresh sessions, no account memory, consistent geography, ideally automated.
- Write scoring rules before the first run, and score mention, citation and sentiment separately.
- Fix the competitor set at the start; a moving comparison set makes share of voice meaningless.
- Report direction and proportion. Do not attribute movement to a single action.
Related services
More reading
- llms.txt is not robots.txt — and treating it that way costs you citations
- The answer-first rewrite: a 50-word pattern that wins snippets
- Why your Shopify feed is invisible to shopping agents
- Best SEO practices for 2026: the ten that still decide rankings
- On-page SEO best practices: what a page must do to be the best answer
- Technical SEO best practices: what has to be true before a page can rank or be quoted
- SEO best practices for AI search: getting cited by ChatGPT, Claude and Perplexity
- Outdated SEO practices: what to stop doing in 2026, and what replaced each one
