GEOGet cited by models

Synthetic query testing

Synthetic query testing runs a fixed panel of buyer prompts against the assistants on a schedule, scoring whether you were mentioned, cited or absent each time. It converts a category that feels unmeasurable into a trend line you can put in a board pack.

At a glance
Engagement
Ongoing, reported weekly
Prerequisite
A fixed panel and competitor set
Cadence
Weekly, run cold
Common recommendation
Never change the panel mid-engagement

How you measure something that answers differently every time

Through sampling and consistency. Any single assistant answer is unreliable — wording shifts between runs and accounts. Two hundred prompts, run cold on a fixed schedule and scored by fixed rules, produces an aggregate that is stable enough to trend even though every individual response varies.

The variability is real and it is why most agencies avoid measuring this. Ask an assistant the same commercial question twice and you may get different brands named, different sources cited, and different framing. Anyone showing you a screenshot as evidence of visibility is showing you a coin flip that happened to land well.

Sampling solves it the way polling solves the same problem. No individual response tells you anything; the distribution across a large fixed panel does. Once the panel size is adequate, week-over-week movement reflects genuine change rather than noise, and you can start attributing that movement to specific work.

Three conditions make it trustworthy. The panel must be fixed, because changing prompts between runs makes the comparison meaningless. The runs must be cold and logged out, because a personalised session reflects your own history rather than a buyer's. And the scoring rules must be written down, because inconsistent scoring produces exactly the trend the person scoring expects.

What this buys you is the ability to say no. Without measurement, every tactic sounds plausible and nothing can be disproven, which is how agencies sell activity indefinitely. With a panel, a tactic that fails to move the number over a quarter gets dropped — including ours.

What the monitoring provides

The monitoring provides scheduled runs of your fixed prompt panel across all five engines, consistent scoring against written rules, per-engine and aggregate trend lines, competitor share of voice, and alerting for whenever a position you previously held disappears from the answers.

Scheduled panel runs

Your two hundred prompts executed cold and logged out on a fixed schedule, so every run is comparable with every other run rather than reflecting somebody's session history.

Written scoring rules

Mentioned, cited with a link, or absent — defined precisely enough that two people scoring the same response reach the same result.

Per-engine trends

ChatGPT, Perplexity, Gemini, Copilot and AI Overviews tracked separately as well as in aggregate, because they disagree and the disagreements are diagnostic.

Competitor share of voice

Who gets named when you do not, tracked over time. This is usually the chart that gets circulated internally without anyone being asked to.

Regression alerting

Notification when a prompt you previously won stops naming you, so a loss is investigated in days rather than discovered in a quarterly review.

Board-ready reporting

One page showing trend, movement and competitor position, written so someone who has never heard of GEO can read it without a translator.

How monitoring is set up and run

Setup takes about a week: build or import the prompt panel, agree the written scoring rules, and establish a baseline. After that it runs on schedule with weekly reporting, a quarterly panel review, and alerting whenever a previously held position disappears.

  1. 01

    Build or import the panel

    From prompt-space research, from a previous baseline audit, or built fresh from your sales and support language. Whatever the source, it gets fixed in writing before the first run.

    A fixed, documented prompt panel

  2. 02

    Agree scoring rules

    What counts as a mention, a citation and an absence, written precisely enough to be reproducible. Ambiguous rules are how monitoring quietly becomes advocacy.

    Written scoring methodology

  3. 03

    Establish the baseline

    A first full run across all five engines, cold, producing the number every later run is compared against. Nothing before this point is a measurement.

    A scored baseline per engine

  4. 04

    Run, report and review

    Scheduled runs with weekly reporting and regression alerts, plus a quarterly review of whether the panel still reflects how your buyers ask questions.

    Weekly trend and quarterly panel review

What makes a measurement trustworthy

Four properties separate measurement from anecdote: a fixed panel, cold runs, written scoring rules, and enough sample size to survive variability. Removing any one of them produces numbers that look rigorous and move for reasons unrelated to anything you did.

Each row on the right is a failure mode we have seen presented as reporting.
PropertyDone properlyCommon failure
PanelFixed and documented before the first runPrompts edited between runs, so nothing is comparable
Session stateCold and logged out every timeLogged-in runs reflecting the tester's own history
ScoringWritten rules, reproducible by two peopleJudgement calls that drift toward the desired trend
Sample sizeLarge enough that noise averages outA screenshot of one favourable answer
CadenceScheduled and consistentRun when someone remembers, usually before a review

Signals you need this now

You need this when you are spending on generative visibility without any scoreboard at all, when an agency reports AI progress using screenshots, or when leadership asks how AI search is performing and the honest answer is that nobody currently knows.

  • You are investing in GEO with no before-and-after number
  • An agency reports AI visibility using screenshots as evidence
  • Leadership asks about AI search and nobody can answer
  • You cannot tell whether last quarter's work changed anything
  • Competitor visibility is discussed anecdotally rather than measured
  • You want to hold an agency — including us — to a number
  • A previous baseline exists but has never been re-run

What the monitoring changes

The immediate change is that arguments about generative visibility become empirical. Within a quarter you can usually see which tactics moved the number and which did not, which is uncomfortable for whoever proposed the ones that did not — including us, by design.

Questions about AI visibility monitoring

Each answer is written to stand alone in 40 to 60 words — the shape an AI Overview or Perplexity citation lifts. Ships with FAQPage schema.

How can you measure something that answers differently every time?

By sampling. Any single response is unreliable, but two hundred prompts run cold on a fixed schedule and scored by written rules produce an aggregate stable enough to trend. It is the same logic that makes polling work despite individual variability.

Why does the panel have to stay fixed?

Because changing prompts between runs makes comparison meaningless — you cannot tell whether a movement reflects real change or a different question. We review the panel quarterly for continued relevance, but changes are deliberate and documented rather than incidental.

Can we buy this without the delivery work?

Yes, and some clients do exactly that — they run execution in-house or with another agency and use us purely for the scoreboard. We think that is a reasonable arrangement, and it keeps the measurement honest by separating it from the delivery.

How is this different from the GEO baseline audit?

The audit is a one-off diagnostic producing a starting number and a gap analysis. This is the ongoing infrastructure that re-runs the same panel indefinitely. Most clients start with the audit; monitoring is what makes it a trend rather than a snapshot.

What if the numbers show your work is not helping?

Then that appears in the weekly report and we change the plan at the quarterly rescope. Building the scoreboard that can disprove our own work is deliberate — an agency that cannot be held to a number will keep selling activity indefinitely.

How quickly can we see meaningful trends?

The baseline is available within a week of setup. Week-over-week movement becomes interpretable after about a month, once you can distinguish signal from normal variability. Attributing movement to specific tactics reliably takes about a quarter.

Find out whether this is your constraint.

Thirty minutes with a senior strategist. We pull your live visibility while we talk and tell you plainly whether a AI visibility monitoring is what you need — or whether your problem sits somewhere else.

Book a discovery call →