Spoken query identification
Which questions in your category are actually asked aloud, drawn from conversational phrasing patterns rather than assumed from keyword data.
Voice search optimisation works to have assistants read your content aloud in response to spoken questions. It is the least forgiving answer surface: one result is returned, there is no second place, and no link is offered — which makes it a branding position rather than a traffic channel.
Typically usually scoped inside a wider AEO programme.
Voice returns a single spoken answer, so the winning content is whatever an assistant can read aloud completely and unambiguously. That favours short, direct, conversational answers to specific questions — and it eliminates any content that depends on formatting, comparison or visual structure to make sense.
The single-result constraint is the whole story. On a results page, ranking fifth still earns something. In voice there is no fifth. Either the assistant reads your answer or your existence is not communicated to the user at all, which makes voice unusually binary compared with every other surface.
Spoken questions are also phrased differently. People speak in full sentences with more context than they type, which makes voice queries closer to assistant prompts than to keywords. Content optimised for conversational phrasing tends to perform on both surfaces, which is part of why this work is rarely sold alone.
Where voice genuinely converts is local and transactional. Asking for the nearest supplier, opening hours, or a booking is a real behaviour with a real outcome. Business information accuracy matters more here than content quality, because the assistant is reading facts rather than prose.
The honest scoping point is that for many businesses voice is a small surface. It is worth doing well when local intent matters, when your category attracts genuinely spoken questions, or as a by-product of answer-first work. Pursuing it alone, for a category nobody asks aloud about, is not a good use of a budget.
The work covers identifying which of your queries are genuinely asked aloud, restructuring answers into speakable form, ensuring business information is accurate across the sources assistants read, and honest scoping about whether voice deserves attention in your category at all.
Which questions in your category are actually asked aloud, drawn from conversational phrasing patterns rather than assumed from keyword data.
Answers rewritten so they can be read aloud completely and make sense without any visual structure, formatting or comparison to support them.
Hours, location, contact details and service areas verified across the sources assistants read, since voice queries are frequently factual rather than editorial.
The near-me and transactional questions where voice genuinely converts, covered properly rather than treated as an afterthought.
Content matched to how people speak rather than type, which also benefits assistant prompts and People Also Ask coverage on the same topics.
A recommendation on whether voice warrants investment in your category. Where it does not, we say so instead of building a programme around it.
The work runs inside a wider AEO engagement: establish whether your category attracts genuinely spoken questions at all, verify your business information across the assistant sources, restructure the answers worth winning into speakable form, then measure whatever can honestly be measured.
Whether your category attracts spoken questions at meaningful volume. For plenty of business categories the honest answer is no, and we would rather establish that first.
→ A recommendation on whether to proceed
Hours, location, contact details and service areas across the sources assistants read. Factual accuracy carries most voice queries, and errors here are common and cheap to fix.
→ Corrected business information across sources
Rewrite the answers worth winning so they read aloud completely and unambiguously, without depending on formatting or visual comparison to carry the meaning.
→ Speakable answers shipped
Voice reporting is genuinely limited. We test a fixed set of spoken questions across assistants and record the answers, rather than presenting estimated voice traffic as measurement.
→ A tested spoken-question panel
Voice is the strictest surface and the hardest to measure. One result is returned, no link is offered, and reporting is limited to testing questions yourself. Those constraints make it a branding and local-conversion play rather than a traffic channel.
| Voice | Other answer surfaces | |
|---|---|---|
| Results returned | Exactly one | Several, with positions below the first |
| Link offered | None | Usually yes |
| Query phrasing | Full spoken sentences | Typed, more compressed |
| What wins | Short, direct, unambiguous speech | Format-matched extractable passages |
| Where it converts | Local and transactional questions | Across informational and commercial intent |
| Measurement | Manual testing only | Search Console and position tracking |
You need this when local or transactional intent matters to your business, when your business information is inconsistent across the sources assistants read, or when your category genuinely attracts spoken questions and nobody has ever checked what assistants currently say.
Voice is a small surface for most businesses and a meaningful one for local and transactional categories. Expect improved accuracy in what assistants say about you rather than a traffic line in analytics, because voice answers return no link and no session.
Each answer is written to stand alone in 40 to 60 words — the shape an AI Overview or Perplexity citation lifts. Ships with FAQPage schema.
For local and transactional businesses, frequently yes. For categories nobody asks about aloud, usually not on its own — though the conversational phrasing work benefits assistant prompts and People Also Ask coverage, so it is often worth capturing as a by-product.
By testing a fixed set of spoken questions across assistants and recording the answers. There is no voice equivalent of Search Console. Anyone presenting you with voice traffic figures is showing modelled estimates rather than measurement, and we will not do that.
Rarely, because a spoken answer returns no link and no session. Voice is a branding and local-conversion position — the value is being the answer read aloud, and for local queries the follow-on action of a call or a visit.
Assistants frequently draw from the same answer content, but voice returns exactly one result with no link and no second place. It is also stricter about ambiguity, since a listener cannot skim, re-read or compare against anything else.
Factual accuracy in your business information, then short unambiguous answers to genuinely spoken questions. Hours, location and service areas being consistent across the sources assistants read carries more voice queries than content quality does.
Usually not, and we will tell you that during scoping. It works best inside a wider AEO engagement where the answer-first content is being written anyway. Standalone voice programmes tend to be sold on hype rather than measurable return.
Thirty minutes with a senior strategist. We pull your live visibility while we talk and tell you plainly whether a voice search optimization is what you need — or whether your problem sits somewhere else.