Scheduled panel runs
Your two hundred prompts executed cold and logged out on a fixed schedule, so every run is comparable with every other run rather than reflecting somebody's session history.
Synthetic query testing runs a fixed panel of buyer prompts against the assistants on a schedule, scoring whether you were mentioned, cited or absent each time. It converts a category that feels unmeasurable into a trend line you can put in a board pack.
Typically ongoing monitoring, reported weekly.
Through sampling and consistency. Any single assistant answer is unreliable — wording shifts between runs and accounts. Two hundred prompts, run cold on a fixed schedule and scored by fixed rules, produces an aggregate that is stable enough to trend even though every individual response varies.
The variability is real and it is why most agencies avoid measuring this. Ask an assistant the same commercial question twice and you may get different brands named, different sources cited, and different framing. Anyone showing you a screenshot as evidence of visibility is showing you a coin flip that happened to land well.
Sampling solves it the way polling solves the same problem. No individual response tells you anything; the distribution across a large fixed panel does. Once the panel size is adequate, week-over-week movement reflects genuine change rather than noise, and you can start attributing that movement to specific work.
Three conditions make it trustworthy. The panel must be fixed, because changing prompts between runs makes the comparison meaningless. The runs must be cold and logged out, because a personalised session reflects your own history rather than a buyer's. And the scoring rules must be written down, because inconsistent scoring produces exactly the trend the person scoring expects.
What this buys you is the ability to say no. Without measurement, every tactic sounds plausible and nothing can be disproven, which is how agencies sell activity indefinitely. With a panel, a tactic that fails to move the number over a quarter gets dropped — including ours.
The monitoring provides scheduled runs of your fixed prompt panel across all five engines, consistent scoring against written rules, per-engine and aggregate trend lines, competitor share of voice, and alerting for whenever a position you previously held disappears from the answers.
Your two hundred prompts executed cold and logged out on a fixed schedule, so every run is comparable with every other run rather than reflecting somebody's session history.
Mentioned, cited with a link, or absent — defined precisely enough that two people scoring the same response reach the same result.
ChatGPT, Perplexity, Gemini, Copilot and AI Overviews tracked separately as well as in aggregate, because they disagree and the disagreements are diagnostic.
Who gets named when you do not, tracked over time. This is usually the chart that gets circulated internally without anyone being asked to.
Notification when a prompt you previously won stops naming you, so a loss is investigated in days rather than discovered in a quarterly review.
One page showing trend, movement and competitor position, written so someone who has never heard of GEO can read it without a translator.
Setup takes about a week: build or import the prompt panel, agree the written scoring rules, and establish a baseline. After that it runs on schedule with weekly reporting, a quarterly panel review, and alerting whenever a previously held position disappears.
From prompt-space research, from a previous baseline audit, or built fresh from your sales and support language. Whatever the source, it gets fixed in writing before the first run.
→ A fixed, documented prompt panel
What counts as a mention, a citation and an absence, written precisely enough to be reproducible. Ambiguous rules are how monitoring quietly becomes advocacy.
→ Written scoring methodology
A first full run across all five engines, cold, producing the number every later run is compared against. Nothing before this point is a measurement.
→ A scored baseline per engine
Scheduled runs with weekly reporting and regression alerts, plus a quarterly review of whether the panel still reflects how your buyers ask questions.
→ Weekly trend and quarterly panel review
Four properties separate measurement from anecdote: a fixed panel, cold runs, written scoring rules, and enough sample size to survive variability. Removing any one of them produces numbers that look rigorous and move for reasons unrelated to anything you did.
| Property | Done properly | Common failure |
|---|---|---|
| Panel | Fixed and documented before the first run | Prompts edited between runs, so nothing is comparable |
| Session state | Cold and logged out every time | Logged-in runs reflecting the tester's own history |
| Scoring | Written rules, reproducible by two people | Judgement calls that drift toward the desired trend |
| Sample size | Large enough that noise averages out | A screenshot of one favourable answer |
| Cadence | Scheduled and consistent | Run when someone remembers, usually before a review |
You need this when you are spending on generative visibility without any scoreboard at all, when an agency reports AI progress using screenshots, or when leadership asks how AI search is performing and the honest answer is that nobody currently knows.
The immediate change is that arguments about generative visibility become empirical. Within a quarter you can usually see which tactics moved the number and which did not, which is uncomfortable for whoever proposed the ones that did not — including us, by design.
Each answer is written to stand alone in 40 to 60 words — the shape an AI Overview or Perplexity citation lifts. Ships with FAQPage schema.
By sampling. Any single response is unreliable, but two hundred prompts run cold on a fixed schedule and scored by written rules produce an aggregate stable enough to trend. It is the same logic that makes polling work despite individual variability.
Because changing prompts between runs makes comparison meaningless — you cannot tell whether a movement reflects real change or a different question. We review the panel quarterly for continued relevance, but changes are deliberate and documented rather than incidental.
Yes, and some clients do exactly that — they run execution in-house or with another agency and use us purely for the scoreboard. We think that is a reasonable arrangement, and it keeps the measurement honest by separating it from the delivery.
The audit is a one-off diagnostic producing a starting number and a gap analysis. This is the ongoing infrastructure that re-runs the same panel indefinitely. Most clients start with the audit; monitoring is what makes it a trend rather than a snapshot.
Then that appears in the weekly report and we change the plan at the quarterly rescope. Building the scoreboard that can disprove our own work is deliberate — an agency that cannot be held to a number will keep selling activity indefinitely.
The baseline is available within a week of setup. Week-over-week movement becomes interpretable after about a month, once you can distinguish signal from normal variability. Attributing movement to specific tactics reliably takes about a quarter.
Thirty minutes with a senior strategist. We pull your live visibility while we talk and tell you plainly whether a AI visibility monitoring is what you need — or whether your problem sits somewhere else.