45 GEO studies from 2023 to 2026: what actually matters for AI search visibility?
SHORT ANSWER
The survey of 45 GEO studies (arXiv:2607.14035) concludes three things. GEO is a probabilistic pipeline from search activation to citation, not one ranking problem. The often-quoted 40% visibility gain came from a setup supplying already-retrieved documents to the model, so it does not show open-web discoverability. And the factors that reproduce are topical relevance and context position.
| Study | |
|---|---|
| Paper | Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026) |
| Author | Olivier Martinez |
| Published | 15 July 2026, arXiv:2607.14035 (Information Retrieval) |
| Scope | 45 studies published November 2023 – July 2026, plus related RAG and evaluation research |
| Analysis | Navirang AEO Research |
This article is Navirang’s reading of a preprint posted to arXiv in 2026 (not yet peer reviewed). We mark the paper’s claims and our practical interpretation separately, and recommend reading the original.
What this study examined
The survey places the 45 GEO studies that appeared rapidly after the first GEO paper in 2023 side by side and reviews them critically. The author’s starting concern is that terminology, metrics and standards of evidence differ from study to study — the same word “visibility” means citation frequency in one paper and share of the answer in another.
The core conclusions in seven lines:
- GEO is not a single “raise the ranking” problem
- AI search visibility has to be read as a multi-stage pipeline
- The most reproducible factors are topical relevance between question and document and position within the context
- Increasing citation for an already-retrieved document and making a document discoverable in the first place are different problems
- Performance cannot be judged from one ChatGPT test — results differ run to run
- Measurement needs repeated runs, paraphrases, several engines and several dates
- No single technique has yet been demonstrated to raise organic discoverability across platforms over a sustained period
GEO is not one ranking problem
The core model the paper offers is to view the stages a brand passes through before appearing in an AI answer as a probabilistic, only partially observable pipeline. The paper’s stages:
| Stage | The question |
|---|---|
| ① Search activation | Does the engine run a web search for this question at all? |
| ② Crawling / indexing | Is our document collected and indexed? |
| ③ Retrieval | Is our document retrieved as a candidate for the question? |
| ④ Reranking / context allocation | Among the candidates, is it actually supplied to the model, and in what position? |
| ⑤ Citation | Among the supplied documents, is it chosen as the answer’s source? |
| ⑥ Prominence | In what position and at what weight does it appear in the answer? |
| ⑦ Absorption | Is the document’s information absorbed into the answer’s content? |
| ⑧ Fidelity | Is the brand and service information expressed accurately? |
| ⑨ User action | Does the answer lead to a visit or an enquiry? |
Why this division matters in practice is simple: the prescription differs entirely according to which stage you are stuck at. Polishing citation phrasing for a site blocked at stage ② is the wrong prescription in the wrong order.
Discoverability and citation are different
The most important distinction for a practitioner to take from this paper.
- Discoverability — before AI generates an answer, does our page enter the retrieval and collection candidate set? (pipeline ①–③)
- Citation — among the retrieved material, does the generated answer choose our page as a source? (pipeline ④–⑤)
Making AI cite an already-retrieved document more often, and making that document retrievable in the first place, are different optimization problems. A good deal of GEO advice in the market does not distinguish them, bundling both under the single phrase “AI visibility”. That produces the misunderstanding in the next section.
“GEO raised AI visibility 40%” may be only half right
The roughly 40% visibility improvement reported by the first GEO paper (Aggarwal et al., 2023) is the most widely quoted figure in this field. This survey does not deny it — it attaches the precise conditions.
In the paper’s terms, the improvement is valid within that experimental environment but depends on the condition that documents had already been retrieved and supplied to the model as fixed context. And that result does not demonstrate organic discoverability or sustained traffic effects.
In practitioner language: there is evidence that “for a document already in the candidate set, polishing it can raise the probability of citation”. There is no evidence that “the probability of being discovered across the web rises by 40%.” Citation optimization and discoverability optimization must not be treated as the same thing.
Factors confirmed repeatedly across the studies
The factors the survey found reproducible across 45 papers. The left two columns carry the paper’s terms; the right column carries Navirang’s practical reading.
| Factor | Meaning in the research | AEO practical reading |
|---|---|---|
| Topical relevance | The most reproducible lever | One page has to answer one question clearly |
| Context position | Position within the supplied context affects citation | Competition continues after retrieval — the direct answer has to be near the front |
| Limits of generic heuristics | Generic heuristics transfer poorly when the environment changes | There is no all-purpose GEO checklist |
| Source competition | Competition erodes an individual document’s gain | One good article does not guarantee continued citation |
| Run-to-run variability | Substantial variability confirmed in audits of commercial tools | Do not judge from one search screen |
| Low source overlap between engines | Referenced documents differ by engine for the same question | ChatGPT, Gemini and Perplexity have to be measured separately |
| Fidelity gaps | Cases where the mention occurs but the information is inaccurate persist | Track accuracy of content, not just whether a mention occurred |
Translating the paper’s concepts into our measurement
The paper proposes treating visibility not as a single score but as a vector separating discoverability, citation, absorption and economic outcome. Below is the six-lens version Navirang reconstructed for practical measurement, adding the pipeline’s prominence and fidelity stages to that vector — the paper does not present this form.
- Discoverability — does our document enter the retrieval candidate set for the question?
- Citation — is it actually chosen as the answer’s source?
- Prominence — at what position and weight does it appear in the answer (first-appearance position, recommendation inclusion)?
- Absorption — do our document’s facts get reflected in the answer’s content?
- Fidelity — is the company and service information expressed accurately?
- Business outcome — does it lead to actual visits and enquiries?
How those six lenses get measured as concrete metrics (mention rate, citation rate, recommendation appearance, share) is set out in how to measure AI visibility, and the measurement unit we applied to ourselves is the baseline of 40 questions × 7 answer engines = 280 cells in the eight-week improvement series.
Why one test is not enough to judge by
Users looking for the same information phrase the question differently. “Tell me about AEO specialist firms”, “recommend a Korean AEO agency”, “which agency is good for ChatGPT visibility?”, “tell me about GEO specialists”, “which company is good at answer engine optimization” — all paraphrases of the same intent, and when the phrasing changes, the answer and the brands that appear can change too.
So the measurement protocol this survey proposes has repeated runs, paraphrases, a control group and human verification. As a measurement unit:
Engine × Query × Paraphrase × Run × Date
For instance, 20 core questions in several phrasings each, across several engines, repeated on several dates (the specific counts are a matter of your own resources; the paper does not prescribe numbers). “Always first” is not an achievable goal in a probabilistic system, so the KPI should be a frequency metric such as the proportion of repeated measurements in which the brand was among the candidates — for example “candidate inclusion rate across 100 repeated runs of non-brand AEO queries” (an illustration of measurement design, not an industry-standard metric).
Why citation-friendly prose should not be trusted blindly
The part of this survey most likely to be uncomfortable for practitioners. An observation is reported that content rewritten to be citation-oriented can harm the retrieval stage.
That is, an article AI finds easy to cite and an article a retrieval system finds easy to surface are not always the same. An approach that only polishes citation phrasing looks at one stage of the pipeline. So AEO has to be approached not as content rewriting alone but as a compound problem covering crawlability, topical relevance, retrieval, citation and external trust signals together. It is also why our audit checks indexing and accessibility before content.
Navirang’s conclusion
Navirang does not regard AEO as “how to write sentences AI likes”. We approach it as a search optimization problem that has to be observed stage by stage, from search activation through citation and fidelity, with the same question set measured repeatedly. This paper adds academic weight to that approach — and it is equally a basis for demanding measurement conditions from every claim an agency like ours makes. Please apply the same standard to our figures too.
If you want to know which stage of the pipeline your brand is stuck at, a free AEO audit will tell you.
References
- Martinez, Olivier. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026).” arXiv:2607.14035 (2026). DOI: 10.48550/arXiv.2607.14035
- Aggarwal, Pranjal, et al. “GEO: Generative Engine Optimization.” arXiv:2311.09735 (2023) — the first GEO paper, the survey’s starting point
Frequently asked questions
Q What is GEO?
A GEO (generative engine optimization) is optimization work that raises the chance content appears, is cited or has influence in answers generative AI produces. This survey proposes treating GEO not as a single ranking task but as a probabilistic pipeline running through search activation, crawling and indexing, retrieval, reranking and context allocation, citation, prominence, absorption, fidelity and user action.
Q Are AEO and GEO the same concept?
A Their execution mostly overlaps but they are terms with different origins. AEO began as a practitioner term for a search engine's direct-answer areas; GEO was proposed in a 2023 academic paper. The relationship and the differences are covered in a separate comparison article.
Q Does doing GEO guarantee ChatGPT visibility?
A It does not. One of the survey's conclusions is that none of the techniques reviewed demonstrated a causal effect of reliably raising organic discoverability across multiple platforms over a sustained period. If a service claims to guarantee visibility, start by asking for the basis and the measurement conditions.
Q What is the 40% improvement the GEO literature talks about?
A It is the visibility improvement reported by the first GEO paper in 2023 and widely quoted since. This survey points out that the result is valid inside that experimental setup but depends on the condition that documents had already been retrieved and supplied to the model as fixed context. In other words it is evidence that citation can be increased for documents already in the candidate set — not that the chance of being discovered on the open web rises by 40%.
Q What is the difference between discoverability and citation in AI search?
A Discoverability is whether our document enters the retrieval candidate set before AI builds an answer; citation is whether, among documents already in that set, ours is chosen as the source of the answer. They are different stages and their optimization differs. Polishing citation-friendly phrasing acts on the latter; the former belongs to indexing, crawling and topical relevance.
Q How should ChatGPT citation be measured?
A Run the same question set repeatedly and record the proportion of answers whose sources include your domain. The survey likewise proposes a protocol with repeated runs, paraphrases, a control group and human verification. Results differing between runs is normal, so a single screenshot is not a measurement.
Q Which metrics should AEO performance use?
A Proportion metrics from repeated measurement, rather than a fixed position such as "always first". Brand mention rate, source citation rate, retrieval inclusion rate, recommendation inclusion, mention accuracy and variance between engines are the main ones. Navirang's metric definitions and formulas are set out in the AI visibility measurement article.
Q Will adding structured data alone raise AI search visibility?
A No. The factors that reproduced best in this survey were not a single technique such as schema but topical relevance between question and document, and position within the context. Generic heuristics were found to transfer poorly when the environment changed. Structured data belongs as a supporting measure that helps machines read facts.
Q Do we have to measure ChatGPT, Gemini and Perplexity separately?
A Yes. Audits of commercial tools confirmed low source overlap between engines — the set of documents referenced differs considerably by engine even for the same question. One engine's result cannot be used to infer another's, so measure per platform and treat a combined figure only as a reference.
If you need this done rather than read
This article belongs to Research and primary sources. The pages that handle the same subject as work are below.
Related reading
- How do you measure AEO and GEO performance? Mention rate, citation rate, recommendation share How to measure how far a brand is found in AI search across ChatGPT, Gemini and Perplexity — the definitions and calculations for mention rate, citation rate, recommendation appearance and share of voice, plus question set design and the repeated-measurement principle.
- Can AEO or GEO top placement be guaranteed? What to check before choosing an agency Can top placement in ChatGPT, Gemini or Perplexity answers be guaranteed? Starting from ZDNet Korea's August 2026 report on overselling, here is why it cannot, what to measure instead, and a checklist for choosing an agency.
- AEO and GEO verified against Google's official guide — what actually works? Taking Google's 2026 documentation on optimizing for generative AI search as the primary source, mapped one-to-one against Navirang's own practice. Includes one actual observation from an answer engine outside Google.