How do you measure AEO and GEO performance? Mention rate, citation rate, recommendation share
SHORT ANSWER
AEO and GEO performance cannot be reduced to a position in one answer, because the same question yields different answers by engine and by moment. You run the same question set repeatedly across several answer engines and record mention rate, citation rate, recommendation appearance and share of voice, reading them as a trend. Navirang uses a baseline of 40 questions × 7 engines = 280 cells.
Why it cannot be measured like an SEO ranking
SEO has a relatively fixed metric: what position our page holds for this query. Generative AI answers have no such fixed point. For the same question the answer and the brands that appear can vary with the platform, the moment, the wording, whether search ran, and which sources were referenced at that instant.
So “first place in AI search” is not as fixed a concept as a traditional search ranking. A single answer screen is not evidence of performance but one sample. It is why, when we publish our own observation screens, we state alongside them that it was a single observation with no guarantee of reproduction. The 2026 academic survey covering 45 GEO studies confirmed run-to-run variability and the need for repeated measurement in the same direction — set out separately in our analysis of that research.
Navirang’s measurement principle follows from that. The goal of AI search optimization is not to hold first place in one answer momentarily but to raise the frequency with which the brand is reliably found and recommended across the range of questions customers actually ask. Performance is measured as that frequency.
What to measure
Rather than one answer, count these four separately across repeated runs. Definitions that blur together inflate figures easily, so we separate them first.
Brand mention rate
A mention is the company or brand name appearing in the body of an AI answer. The mention rate is calculated as:
mention rate = answers where the brand appeared ÷ total answers measured × 100
The denominator is the core of it. “A mention rate of 50%” cannot be verified without knowing how many questions, in how many engines, asked how many times.
Source citation rate
A citation is the AI answer using your website as a source or reference link while providing information.
Mention rate and citation rate are not the same metric. AI can recommend Navirang without using navirang-ai.com as a source, and conversely it can use a navirang-ai.com document as a source without recommending the brand. A mention is a signal of brand awareness and a citation a signal of document trust, so the two values moving apart is itself diagnostic information.
Recommendation appearance rate
A recommendation is the brand being presented as an actual candidate when a user asks for a company or service to be recommended. It is calculated as the proportion of vendor-search questions such as “tell me about AEO agencies” in which the brand is among the candidates. It is closer to conversion than a plain mention, which makes it especially important for a B2B service like ours.
AI share of voice
The proportion of appearances between our brand and competing brands across the same question set. For example — the figures below illustrate the calculation and are not measured values — if across 100 vendor-recommendation answers Company A appeared 38 times, Company B 27 and our brand 21, the shares are 38%, 27% and 21%. Your own mention rate can hold steady while your relative position falls as competitors rise, so absolute values and share have to be read together.
Per-platform visibility
The answer engines Navirang measures are ChatGPT, Google AI Overviews, Perplexity, Claude, Naver AI Search, Copilot and Gemini — seven. Their retrieval paths differ (own crawler, search partnership, domestic index), so results differ substantially. Averaging per-platform figures into one hides that difference: a change where citation begins in one engine only does not show in an average. Split them, and use the combined figure only as a reference.
Supporting metrics — when looking deeper
Beyond the four core metrics, three more are useful to record during cause analysis.
- Source diversity — across how many distinct external domains is information supporting the brand confirmable? More independent sources may raise the chance AI trusts the brand’s information; this is a hypothesised metric and the size of the effect remains to be verified
- Mention order — when several brands are listed, in what position does the brand appear? It is not a fixed ranking, but the average position across repeated measurement is a trend signal
- Follow-up persistence — does a brand that appeared in the first answer remain a candidate through follow-up questions such as “which of those is best?” In conversational search, persisting matters as much as first appearing
The questions you measure with change the result
The question set matters as much as the measurement tool. A measurement that asks only “AEO agency” is not a measurement. Real customers ask different questions at different purchase stages, and appearance frequency differs by question type.
| Question type | Example | What this type reveals |
|---|---|---|
| Informational | “What is AEO?”, “What is the difference between AEO and SEO?” | Whether the concept documents get cited |
| Problem-solving | “How do we get our company to appear in ChatGPT?” | Whether the practical guides get cited |
| Vendor search | “Tell me about AEO agencies”, “Recommend a GEO agency in Korea” | Whether you are among the recommendation candidates |
| Comparative | “What is the difference between an AEO agency and an SEO agency?” | Whether the comparison documents get cited |
| Purchase decision | “How much does AEO delivery cost?”, “How do you choose an AEO agency?” | Decision-stage documents and brand trust |
Being cited in informational questions while not being recommended in vendor-search questions means the documents are being read but the brand entity is not connected to “an agency in this field”. That diagnosis is only possible when the types are split.
Measurement design — Navirang’s AI visibility framework
Our measurement design summarises in one line.
Persona × Question × Platform × Repetition → Mention · Citation · Recommendation · SOV
| Step | Content |
|---|---|
| ① Persona | Who is asking — customer and purchase stage defined |
| ② Question | What is being asked — question set designed by type |
| ③ Platform | Where it is asked — each of the seven answer engines |
| ④ Repetition | How often — periodic re-runs under the same conditions |
| ⑤ Recording | Mention, citation and recommendation recorded separately per answer |
| ⑥ Comparison | Share against competing brands calculated |
Applying this framework to ourselves gives the case series’ baseline of 40 questions × 7 answer engines = 280 cells. It is not a newly invented concept but the measurement unit we actually run.
How Navirang measures
The actual sequence matches the audit stage in how we work.
- Define the client and the competitors — brand name, services, category and competing brands settled
- Design the real customer questions — a question set built by persona and purchase stage. That design ability is what we sell, so the method for designing a client’s question set is not published, while the question set for our own measurement is published in full
- Measure per platform — each of the seven answer engines recorded separately
- Repeat the same questions — a baseline at the outset, re-run weekly
- Count mentions and citations apart — brand mention, site citation and recommendation inclusion recorded separately
- Analyse share against competitors — recommendation appearance and SOV compared
- Connect results to execution — feeding into citable content, structured data and entity work, with wrong mentions handled as corrections
Measurement is read-only. Inflating visibility through automated repeat searching, or manipulating results, is abuse and we do not use it.
Why your brand does not appear often in AI
Run the measurement and the causes usually split into three layers.
- Technical layer — search-side AI crawlers blocked, missing from the index, body text drawn only by JavaScript. The order of checks is in why ChatGPT does not mention your company
- Content layer — only a company introduction with no documents answering real customer questions; conclusions that do not complete inside a paragraph and so cannot be cut out; too few documents connecting the category (“AEO agency”) to the brand
- Trust layer — the company written differently per channel (entity inconsistency), too little independent external material backing your claims, no primary data to use as evidence
The prescription differs entirely by layer, so measurement comes before improvement. If you want to know which layer your brand is stuck at, a free AEO audit will tell you.
When choosing an agency, look at measurement ability first
The metrics in this article are themselves the verification tools — ask whether they publish the question set, whether it is repeated measurement or a single screenshot, whether mentions and citations are counted separately, and whether you get competitor share. The full selection checklist and how to judge each sales phrase are in the agency selection guide; how to verify a denominator is in the performance figure verification guide.
Where the measurement records accumulate
Navirang runs an eight-week improvement series measuring the same question set repeatedly against itself. The 280-cell baseline measurement is in preparation, and we do not quote figures before the measured data exists — when the baseline is ready it goes out with the full question set, the measurement protocol and the raw data, and the rounds go in as they are even when they are worse. A measurement we have already published is the field measurement of Naver web documents across all 17 Korean provinces, raw CSV included.
References
The argument that AI visibility should be read as repeated measurement and appearance frequency rather than a single ranking is being made elsewhere in the industry too. In Korea, Team HAI’s article on GEO measurement methodology is worth reading, and on why the terminology keeps shifting, Passionfruit’s analysis. Those perspectives and our methodology were developed independently, and the definitions and calculations in this article are Navirang’s official methodology.
Frequently asked questions
Q How is AEO performance measured?
A You settle a core question set, put it to several answer engines repeatedly, and record for each answer whether the brand was mentioned, whether your site was cited, and whether you were among the recommendation candidates. Rather than a single value, you read mention rate, citation rate, recommendation appearance and share against competitors as a trend over time. A measurement that does not publish the question count and engine count (the denominator) cannot be verified.
Q Can we just look at a GEO score?
A Judging on a single composite score is not something we recommend. Combining several metrics into one number makes the result depend on the formula and makes it hard to trace back what improved and what got worse. If a score is used, check that the raw metrics — mention rate, citation rate, recommendation appearance — are published with it.
Q Can you measure what position we hold in ChatGPT?
A There is no fixed position. Generative answers can vary in composition and order by session and by moment even for the same question, so saying "we are Nth" the way a search ranking does is not possible. What is measurable is the proportion of repeated measurements in which the brand appeared, and the context of that appearance — a recommendation candidate, or a passing mention.
Q What is the difference between a mention and a citation in AI search?
A A mention is the brand name appearing in the body of the answer; a citation is the answer using your website as a source. The two move independently. AI can recommend the brand without using your site as a source, and it can use your site as a source without recommending the brand, so the two have to be counted separately.
Q Why does ChatGPT's answer change for the same question?
A Because answer generation is probabilistic. The wording and composition the model chooses can vary for the same question, and where search runs alongside, the referenced documents depend on the search results at that moment. Session history or user settings are sometimes reflected too. So a single answer screen is one sample rather than evidence, and repeated measurement is required.
Q How many times should AEO measurement be run?
A Repetition under the same conditions matters more than a set number of runs. At minimum, build a baseline at the outset and re-measure the same question set weekly to read the trend. Navirang's own measurement takes 40 questions × 7 answer engines = 280 cells as the baseline and repeats it weekly.
Q Can we combine ChatGPT and Gemini results?
A Looking only at a combined average is not something we recommend. Engines differ substantially in retrieval path and answer generation, and averaging hides that difference. If citation begins in one engine only, for instance, the average barely moves and the change gets missed. Split by platform first, then use a combined figure as a reference if needed.
Q Why does our company not appear in AI search?
A There are three layers of cause. The technical layer (crawler blocking, missing from the index, JavaScript dependence), the content layer (no documents answering customer questions, no sentences that can be cut out), and the trust layer (brand entity inconsistency, no external sources backing your claims). The order of checks for the technical layer is set out in the article on why ChatGPT does not mention your company.
Q What should we check when choosing an AEO agency?
A From a measurement point of view, six things: which AI platforms they measure, whether they publish and explain the question set, whether they measure repeatedly, whether they count mentions and citations separately, whether competitor comparison (share) is possible, and whether they go beyond analysis into content, structured data and entity work.
Q Which AI search engines does Navirang measure?
A Seven: ChatGPT, Google AI Overviews, Perplexity, Claude, Naver AI Search, Copilot and Gemini. It is the same list as the measurement targets of our own case series, and we do not change the list mid-measurement so the time series stays intact.
If you need this done rather than read
This article belongs to Measurement and judging results. The pages that handle the same subject as work are below.
Related reading
- Can AEO or GEO top placement be guaranteed? What to check before choosing an agency Can top placement in ChatGPT, Gemini or Perplexity answers be guaranteed? Starting from ZDNet Korea's August 2026 report on overselling, here is why it cannot, what to measure instead, and a checklist for choosing an agency.
- 45 GEO studies from 2023 to 2026: what actually matters for AI search visibility? Reading the 2026 GEO critical survey (45 studies reviewed) from a practitioner's point of view — the difference between discoverability and citation, what "40% visibility improvement" precisely means, the factors that reproduce, and the repeated-measurement protocol.
- AEO and GEO verified against Google's official guide — what actually works? Taking Google's 2026 documentation on optimizing for generative AI search as the primary source, mapped one-to-one against Navirang's own practice. Includes one actual observation from an answer engine outside Google.