Skip to content

MEASUREMENT PROTOCOL

Measurement methodology —
we fix the conditions first

Conditions come before results. Here is the protocol Navirang fixes and holds to when measuring AI search visibility — and what that protocol still cannot tell us.

IN ONE PARAGRAPH

Navirang measures AI search visibility with one question × one answer engine = one cell as the unit of observation. Our own baseline is 40 questions × 7 answer engines = 280 cells, re-asked weekly in a logged-out session on the same day and at the same time of day, in the same words. Mentions, citations and recommendation appearances are recorded separately, and the denominator for the citation rate is fixed at the start of measurement and never changed afterwards.

This page is the protocol we actually hold to, not a general explanation of the metrics.

PROTOCOL

What we fix,
and how

If any value below changes, everything after it is a different time series. So when we change one, we record that it changed and why.

Unit of observation
One question × one answer engine = one cell. Our own baseline is 40 questions × 7 engines = 280 cells
Engines
ChatGPT · Google AI Overviews · Perplexity · Claude · Naver AI Search · Copilot · Gemini
Question composition
10 conceptual · 10 agency-search · 8 brand-named · 7 local · 5 data (for our own baseline)
Session
Logged-out sessions. We ask in a fresh session so personalization and prior conversation context do not mix in
Question wording
Fixed for the whole measurement period. Changing the phrasing starts a different time series from that point
Cadence
Weekly, on the same day and at the same time of day
Recorded items
Mentioned · cited · position in the source list · factual accuracy of the description · competitors named alongside · anything notable
Denominator
Fixed at the start of measurement (280 for our baseline). Engines are not added or removed afterwards

METRICS

Three metrics,
counted separately

Being named, being linked as a source, and making a candidate list are different outcomes. Adding them together erases what has to be fixed.

Mention

Did the brand name appear in the body of the answer. It counts as a mention even with no source link.

Cells with a mention ÷ 280 × 100

A name appearing without a source link is still recorded as a mention, and is never added together with citations.

Citation

Did our domain enter the answer's source list. It does not necessarily travel with a mention.

Cells with a citation ÷ 280 × 100

Position within the source list is recorded alongside. Drop the position and change becomes invisible between rounds.

Recommendation

Did we make the candidate list on questions asking for options. Counted only on vendor-search questions.

Cells where we appeared as a candidate ÷ cells of that question type × 100

The denominator here is that question type, not the whole cell count — and we always say so.

LIMITS

What this method
cannot tell us

A measurement with no stated limits cannot be verified. Below is what we have not solved either.

Responses are probabilistic

Asking the same question again under the same conditions can produce different sentences and different sources. So we do not conclude from a single observation; we measure repeatedly and look at the change.

The sample is one company

Our baseline is a record of Navirang alone. It is not a controlled experiment, and there is no guarantee that what is observed here reproduces at another company.

Changes on the engine side mix in

Values move in stretches where we touched nothing, because index refreshes and model swaps happen in the same period. Items that moved without intervention are marked separately in the report.

A crawler visit is not a citation

A crawler having been here (A), being cited in an answer (B), and a visit arriving via AI (C) are different metrics. We never argue B from a rise in A.

We cannot remove the effect of region, language and account state

Logged-out sessions reduce personalization but the connecting region and language settings remain. We fix the conditions; we do not claim the effect is zero.

REPRODUCIBILITY

So you can measure it
the same way

There is one reason we publish the conditions, the question set and the aggregation method — a figure that cannot be verified is only a claim.

Conditions published

The full question set, engine list, session conditions and aggregation formula are published with the results.

Raw data published

Published reports include the address of the raw data file.

Snapshots frozen

A published data file is never overwritten by a re-measurement; a new file is issued instead.

Change log

  1. Answer engines 8 → 7, denominator 320 → 280

    We removed SearchGPT from the list. OpenAI folded it into ChatGPT Search on 2024-10-31 and retired the name, so counting it as a separate engine would count ChatGPT twice. The correction was possible because measurement had not yet begun; once the first cell is filled the denominator does not change.

  2. The protocol was consolidated onto this page

    The same protocol was written slightly differently across several articles. Experiment write-ups now cite this page rather than restating the conditions.

FAQ

Questions about the protocol

Why measure by question rather than by keyword?

Because answer engines answer questions, not keywords. On the same topic, "what is AEO" and "which AEO agency is any good" draw on different documents as evidence. Keyword rankings do not show that difference, so we make the question the unit of observation and record it by type.

Why fix the denominator?

Because when the denominator moves, the same performance becomes a different number. Cutting the question count by 20 and saying the citation rate rose is a change in arithmetic, not an improvement. We fix the question set and the engine set at the start of measurement, and if something new is worth measuring we start a separate set rather than touching the existing one.

Why count mentions and citations separately?

Because the place to fix is different. A mention without a citation means the engine knows the brand but could not find a document to use as evidence; neither means it is blocked at the discovery stage. Counting them as one erases what needs fixing.

Will measuring this way raise our ranking?

Measurement is not improvement. This protocol exists to confirm the present state repeatedly under the same conditions; what to fix comes out of the results. What each engine selects on is not public, so no work can guarantee visibility.

Do client measurements use the same protocol?

The same frame, but the question set is designed fresh for the industry. The number of questions, the engines and the cadence vary with the engagement, and those conditions are fixed in writing before the start and kept by the client. The method for designing industry question sets is not published.

Can we see the raw data?

Published reports come with the original files. Our measurement across all 17 Korean provinces has a public CSV address, and because it is a frozen snapshot it is not overwritten by re-measurement. Our own baseline will be published the same way once measurement completes.

The research this protocol draws on is collected in the paper library, with each study's sample, conditions and stated limitations.

Want this protocol run against your brand?

Send us a URL. We agree the question set and the conditions before we start, and the record comes back with the raw observations attached.

We reply within one business day.

Free audit Call Email Blog