A buyer's framework for evaluating answer engine optimisation agencies, including the answers that should end the conversation.

Answer Engine Optimisation is roughly two years old as a commercial category and already has more agencies than track records. Some of the work is real. A good deal of it is a traditional SEO retainer with the deliverables renamed, sold to buyers who have no established way to tell the difference.
This is the framework we would use if we were the ones buying. It is deliberately awkward in places, because the questions that are comfortable to answer are the ones that do not separate anybody.
Ask for the measurement method before anything else. There is no Search Console for ChatGPT, no rank tracker of record for Perplexity, and Google Analytics barely registers assistant traffic at all. Any agency claiming AI visibility work has had to solve measurement somehow, and how they solved it tells you how seriously they take the problem.
What a real answer sounds like: a named set of prompts, run on a fixed cadence, across named assistants, recorded before any work starts, with competitor share captured alongside yours. What a weak answer sounds like: "we track AI mentions" with no method attached, or a dashboard screenshot with no explanation of what generated the numbers.
The follow-up that matters: what did my baseline look like before you started? If nobody recorded one, nobody can show you movement later. They can only show you a number.
ChatGPT, Perplexity, Google AI Overviews, Claude, Gemini and Microsoft Copilot do not select sources the same way, and a brand can be well cited in one and absent from the rest. An agency covering only the one that is easiest to check is measuring convenience, not visibility.
Frequency matters as much as coverage. These systems are non-deterministic — the same prompt can return different sources on different days. A single check is an anecdote. A monthly check against a fixed prompt set is data.
This one is close to a single-question filter.
Nobody controls what a language model cites. The retrieval layer changes without notice, the ranking of sources inside it is opaque, and no contract with an agency binds OpenAI or Google. An agency guaranteeing citations is either misunderstanding the mechanism or counting on you not to check.
The honest version is narrower and more useful: guarantee the work, not the outcome. What ships, when it ships, and what moved — including the months when the answer is nothing.
Most of what earns citations is unglamorous. Structured data that agrees with the visible copy. An organisation entity a model can resolve without ambiguity. Consistent facts across your site, your business profiles and third-party sources. Content written so a passage can be lifted cleanly rather than inferred from a graphic.
Ask to see this on a page they have worked on. Open the source. The schema is either there and coherent, or it is not — this is one of the few claims in the category you can verify yourself in about ninety seconds.
Here is the part most proposals skip. Models weight sources they already trust, and a business only describing itself is a single unverified source. Citations tend to follow corroboration: independent coverage, credible directories, industry roundups, professional profiles that agree with each other.
That work is slow and it is partly outreach, not engineering. An AEO plan with no answer for how third-party sources will start describing you is a plan to optimise a page nobody vouches for.
A specific trap for businesses in smaller language markets. Terms that feel central to the category can carry almost no search volume in a national language, while the equivalent English terms carry real demand. Optimising hard for a term twenty people a month search is a way to produce a report rather than a pipeline.
An agency should be able to show you the volume behind every term they propose targeting, and should tell you plainly when the honest strategy is entity and citation work rather than ranking.
Schema, rewritten pages, documented process and your own measurement data should be yours the moment they ship, and should keep working if the engagement ends. If the answer involves a proprietary dashboard you lose access to, you are renting visibility rather than building it.
| Question | A real answer | A warning sign |
|---|---|---|
| Measurement | Fixed prompt set, named assistants, recorded baseline | "We track AI mentions" |
| Coverage | Several assistants, monthly cadence | One assistant, checked ad hoc |
| Guarantees | Guarantees the work, not the citation | Guarantees a mention or a position |
| Entity work | Verifiable schema on live pages | Described but never shown |
| Corroboration | Named outreach and citation plan | On-page work only |
| Keyword reality | Volume shown for every target term | Category jargon with no data |
| Ownership | You keep schema, content and data | Access ends with the retainer |
It would be inconsistent to publish this and expect an exemption, so: we have no published client case studies yet, and we would rather say that than borrow someone else's numbers. What we do instead is record where you stand before anything changes, then report citation share across ChatGPT, Perplexity and Google AI Overviews against that baseline every month — so movement, or its absence, is visible from month one.
We guarantee the work and not the citation. We will tell you when the honest read is that a term is not worth chasing. And if we stopped working together tomorrow, you would hold the documentation and the data.
Ask us the seven questions. If the answers do not hold up, that is useful information too.
Get a free audit of where you stand — technical SEO health, AI visibility gaps, and your top 3 fixes, ready in 48 hours.
Get Your Free Audit