How far can you trust the performance figures an AEO agency quotes?
AEO practice series · Part 7 of 7SHORT ANSWER
A figure with no measurement conditions attached cannot be verified. How many questions were asked, in which answer engines, when, and whether the number is measured or illustrative — a citation rate or improvement missing those four is a sentence shaped like a number.
The conclusion first
A figure without measurement conditions cannot be verified. “A citation rate of 40%” or “doubled in three months” is by itself neither true nor false. Only with conditions attached does it become a sentence you can judge.
This is not an article about which agency to choose. It is about how to read a figure. So in the last section we put our own published numbers on the same standard.
What a citation rate counts — the denominator is everything
A citation rate is the proportion of asked questions in which you were cited. Which makes the decisive part not the numerator but the denominator.
The same 40% differs this much:
| Denominator | What it means | |
|---|---|---|
| A | 4 of 10 questions | One question changing makes it 30% or 50% |
| B | 80 of 200 questions | Little wobble |
| C | 10 questions the agency chose | Only favourable questions may be in there |
| D | 200 questions the client set | A verifiable condition |
The smaller the denominator, and the more the questions were chosen by the side doing the measuring, the weaker the number’s meaning. A citation rate with no query count tells you nothing yet.
Beyond citation rate — mention rate, recommendation appearance, share — what gets counted and how is set out metric by metric in how to measure AI visibility.
Five things to check
When you receive a figure, look for these five alongside it.
1. The query set — which questions were asked. You should be able to see the full list. Confirm who set the questions, too.
2. The engine — which answer engine it was measured in. Being cited in ChatGPT does not mean being cited in Perplexity; the indexes they reference differ.
3. The date — when it was measured. Answer engine responses change week to week. A figure with no measurement date is food with no expiry date.
4. The sample size — the number of questions and the number of repetitions. Answer engines do not give the same answer every time to the same question. If each was asked once, the value may not reproduce.
5. Measured or illustrative — whether the number on screen is an actual observation or an example made for explanation.
If any one of these is missing, asking for that condition is legitimate. Anywhere that measured properly can answer immediately, and an inability to answer is itself the answer.
How to read a figure footnoted “illustrative data”
A common shape in industry material. A large number in the middle of the screen, with “this is illustrative data” in small type below.
Adding the footnote is itself honest handling. It beats using the figure with no marking at all. The problem is the layout. People read the large type and skip the small.
The reader has one usable test.
Does the claim on the page survive deleting that number?
If it does, the number is an example aiding the explanation — that is fine. If the number is the claim, an example is not enough and measured data has to be produced.
For a screen showing “here is the kind of metric we report and in what form”, illustrative is enough. For a screen saying “engage us and it rises this much”, illustrative does not hold.
What it takes for a figure about timing to hold
Expressions like “three months is enough” or “cited within six weeks” require a recorded starting point.
With no baseline measured on the same question set before starting, confirming a citation later gives no way to judge whether the work caused it or it was always so. Answer engines change their answers as indexes refresh, whether or not we did anything.
So what to check is not the outcome figure but this:
- Is a baseline measured at the outset?
- Do the baseline and the later measurements use the same query set and the same engines?
- If the query set changed midway, is the change disclosed?
An improvement figure presented without a baseline is a number with nothing to compare against.
Six questions to put to an agency verbatim
Feel free to copy them. The purpose is not hearing the answer so much as seeing whether they can answer.
- How many questions produced that figure, and can we see the question list?
- Who set the questions? Can we add to or change them?
- Which answer engines, and when, was it measured?
- How many times was the same question repeated?
- Do you measure a baseline the same way before starting?
- Over the contract, do you guarantee the result, or the scope of work and reporting method?
If the answer to 6 is “we guarantee citation”, it is worth thinking again. What decides citation is neither the agency nor the site owner but the answer engine. What can be guaranteed extends as far as what will be done, how, and how it will be reported.
Actual overselling seen in the market, and how to judge each sales phrase, are set out separately in can top placement be guaranteed — the agency selection guide.
Hold our figures to the same standard
Applying this standard only to others would make the whole article hypocrisy. Here are our own published numbers on the same scale.
What we published — field measurement of regional search across all 17 Korean provinces, 340 items.
| Item to check | Our value |
|---|---|
| Query set | "<region> marketing agency", "<region> advertising agency" — 17 region names × 2, all published |
| Engine | Naver (web document results via the search API) |
| Date | 29 July 2026 |
| Sample size | Top 10 per query × 34 cells = 340 items. Not repeated — one day, once |
| Measured or illustrative | Measured |
The weak points go in as they are. It was a single day’s measurement with no repeat, so reproducibility was not confirmed, and the top-10 window is narrow enough to move a lot. It measured the composition of search results rather than answer engine citation rate, so it is not an AEO performance metric in itself.
And we got something wrong once and corrected it. We initially labelled this value “Naver search page one”; on 3 August 2026 we opened the actual integrated search screen in a browser and found Power Link and Place blocks sitting above the web document area. What we measured was not the whole of page one but the web document section inside it. We corrected the wording in 75 places to “top 10 Naver web documents”.
The figures were right and the name was wrong. Not disclosing that kind of thing leaves no reason to believe the next number either.
One more. The answer example card on our home page uses a fictional company (Company A) as an illustration, not a result. The screen says so.
Summary
- A citation rate can only be read once you know the denominator
- Five things to check — query set, engine, date, sample size, measured or not
- Judge illustrative data by whether the claim survives removing the number
- An improvement figure only holds with a baseline measured before starting
- What can be guaranteed is not the outcome but scope of work and measurement method
The conditions we measure under are set out in how we work, and the raw results, limits included, are in the regional search measurement report.
Frequently asked questions
Q Can we judge on a citation rate of 40% alone?
A No, because without the denominator the meaning is undetermined. Cited in 4 of 10 questions and cited in 80 of 200 are both 40% with entirely different reliability, and if the agency chose the questions themselves, only favourable ones may be in there. You need the denominator (the number of queries) and who set the query set, and how, before that number can be read.
Q Is a footnote saying "illustrative data" a problem?
A Adding the footnote is itself honest handling. The problem is when the figure is laid out as though it were measured. If the largest number on screen is illustrative and the footnote is small, readers take it as a result. The test is simple: if the claim on the page survives removing that number, it was illustrative; if the number is the claim, measured data has to be produced.
Q Should we rule out an agency that guarantees results?
A Guaranteeing citation does not hold up. Answer engine responses vary with time and session even for the same question, and what decides citation is neither the agency nor the site owner but the engine. What can be guaranteed is scope of work and measurement method, not the outcome. An agency that can put "we will measure this, this way, and report it" into the contract is enough.
If you need this done rather than read
This article belongs to Measurement and judging results. The pages that handle the same subject as work are below.
Related reading
- Can AEO or GEO top placement be guaranteed? What to check before choosing an agency Can top placement in ChatGPT, Gemini or Perplexity answers be guaranteed? Starting from ZDNet Korea's August 2026 report on overselling, here is why it cannot, what to measure instead, and a checklist for choosing an agency.
- How do you measure AEO and GEO performance? Mention rate, citation rate, recommendation share How to measure how far a brand is found in AI search across ChatGPT, Gemini and Perplexity — the definitions and calculations for mention rate, citation rate, recommendation appearance and share of voice, plus question set design and the repeated-measurement principle.
- Is your company ready to appear in AI search? A 20-point self-check A readiness checklist a company can run itself with no external tools. Twenty items across five areas — discoverability, entity, content, external trust, measurement — with the reason each one matters.