The two options
Running a fixed prompt panel cold, on a schedule, across several engines, and scoring the answers by written rules.
- You need to measure generative visibility
- You are reporting GEO progress to a board
- Single assistant answers keep contradicting each other
Checking your position for a keyword set in conventional search results.
- You are measuring classic organic performance
- You need long historical baselines
- Competitive position is the reporting requirement
How they differ
| Synthetic query testing | Rank tracking | |
|---|---|---|
| Result stability | Low — varies between identical runs | High |
| Single check meaningful | No | Yes |
| Method | Repeated sampling, aggregate scored | Direct positional lookup |
| Personalisation effects | Severe — must run cold | Manageable |
| Reports | Citation and mention rates | Positions |
What actually decides it
The reason a screenshot of an assistant naming you proves nothing is variance. Run the same prompt ten times and the answer set changes; run it while signed in and it changes again. Any claim built on a single response is describing one sample from a distribution nobody measured, which is why so much generative-visibility marketing is untestable.
Running cold is the part most often skipped. Personalisation, memory and conversation history all bias results toward whatever the tester has looked at before — including, frequently, the client's own site. Sessions must be clean and preferably automated, or the measurement quietly reports the tester rather than the market.
The output is also different in kind. Rank tracking gives you a position; query testing gives you a rate — the share of panel prompts whose answers name or cite you. Rates need larger samples to move meaningfully, which is why panels run to hundreds of prompts and why weekly noise should not be over-read.
Questions
Why not just ask ChatGPT and screenshot it?
Because the next run may differ, and a signed-in session is biased by your own history. One answer is a single sample from a distribution — it demonstrates possibility, not position.
How many prompts and how often?
Typically one to two hundred prompts, run weekly, across the engines your buyers actually use. Smaller panels are swamped by variance; less frequent runs make it hard to attribute movement to anything.
Does this replace rank tracking?
No. It measures a different surface. Classic organic still drives most pipeline for most businesses, and rank tracking remains the cleanest read on competitive position there.
Related services
- Synthetic Query Testing — GEO
- SEO Analytics — SEO
- GEO Baseline Audit — GEO
- AEO Monitoring — AEO
Still not sure?
Thirty minutes with a senior strategist. We pull your live visibility while we talk and tell you which of these two is actually binding for you — including when the answer is neither. Book a discovery call →
