Skip to content

45 GEO studies from 2023 to 2026: what actually matters for AI search visibility?

Navirang Published AEO ResearchGEOPaper analysisMeasurement

SHORT ANSWER

The survey of 45 GEO studies (arXiv:2607.14035) concludes three things. GEO is a probabilistic pipeline from search activation to citation, not one ranking problem. The often-quoted 40% visibility gain came from a setup supplying already-retrieved documents to the model, so it does not show open-web discoverability. And the factors that reproduce are topical relevance and context position.

Study
PaperOptimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)
AuthorOlivier Martinez
Published15 July 2026, arXiv:2607.14035 (Information Retrieval)
Scope45 studies published November 2023 – July 2026, plus related RAG and evaluation research
AnalysisNavirang AEO Research

This article is Navirang’s reading of a preprint posted to arXiv in 2026 (not yet peer reviewed). We mark the paper’s claims and our practical interpretation separately, and recommend reading the original.

What this study examined

The survey places the 45 GEO studies that appeared rapidly after the first GEO paper in 2023 side by side and reviews them critically. The author’s starting concern is that terminology, metrics and standards of evidence differ from study to study — the same word “visibility” means citation frequency in one paper and share of the answer in another.

The core conclusions in seven lines:

  • GEO is not a single “raise the ranking” problem
  • AI search visibility has to be read as a multi-stage pipeline
  • The most reproducible factors are topical relevance between question and document and position within the context
  • Increasing citation for an already-retrieved document and making a document discoverable in the first place are different problems
  • Performance cannot be judged from one ChatGPT test — results differ run to run
  • Measurement needs repeated runs, paraphrases, several engines and several dates
  • No single technique has yet been demonstrated to raise organic discoverability across platforms over a sustained period

GEO is not one ranking problem

The core model the paper offers is to view the stages a brand passes through before appearing in an AI answer as a probabilistic, only partially observable pipeline. The paper’s stages:

StageThe question
① Search activationDoes the engine run a web search for this question at all?
② Crawling / indexingIs our document collected and indexed?
③ RetrievalIs our document retrieved as a candidate for the question?
④ Reranking / context allocationAmong the candidates, is it actually supplied to the model, and in what position?
⑤ CitationAmong the supplied documents, is it chosen as the answer’s source?
⑥ ProminenceIn what position and at what weight does it appear in the answer?
⑦ AbsorptionIs the document’s information absorbed into the answer’s content?
⑧ FidelityIs the brand and service information expressed accurately?
⑨ User actionDoes the answer lead to a visit or an enquiry?

Why this division matters in practice is simple: the prescription differs entirely according to which stage you are stuck at. Polishing citation phrasing for a site blocked at stage ② is the wrong prescription in the wrong order.

Discoverability and citation are different

The most important distinction for a practitioner to take from this paper.

  • Discoverability — before AI generates an answer, does our page enter the retrieval and collection candidate set? (pipeline ①–③)
  • Citation — among the retrieved material, does the generated answer choose our page as a source? (pipeline ④–⑤)

Making AI cite an already-retrieved document more often, and making that document retrievable in the first place, are different optimization problems. A good deal of GEO advice in the market does not distinguish them, bundling both under the single phrase “AI visibility”. That produces the misunderstanding in the next section.

“GEO raised AI visibility 40%” may be only half right

The roughly 40% visibility improvement reported by the first GEO paper (Aggarwal et al., 2023) is the most widely quoted figure in this field. This survey does not deny it — it attaches the precise conditions.

In the paper’s terms, the improvement is valid within that experimental environment but depends on the condition that documents had already been retrieved and supplied to the model as fixed context. And that result does not demonstrate organic discoverability or sustained traffic effects.

In practitioner language: there is evidence that “for a document already in the candidate set, polishing it can raise the probability of citation”. There is no evidence that “the probability of being discovered across the web rises by 40%.” Citation optimization and discoverability optimization must not be treated as the same thing.

Factors confirmed repeatedly across the studies

The factors the survey found reproducible across 45 papers. The left two columns carry the paper’s terms; the right column carries Navirang’s practical reading.

FactorMeaning in the researchAEO practical reading
Topical relevanceThe most reproducible leverOne page has to answer one question clearly
Context positionPosition within the supplied context affects citationCompetition continues after retrieval — the direct answer has to be near the front
Limits of generic heuristicsGeneric heuristics transfer poorly when the environment changesThere is no all-purpose GEO checklist
Source competitionCompetition erodes an individual document’s gainOne good article does not guarantee continued citation
Run-to-run variabilitySubstantial variability confirmed in audits of commercial toolsDo not judge from one search screen
Low source overlap between enginesReferenced documents differ by engine for the same questionChatGPT, Gemini and Perplexity have to be measured separately
Fidelity gapsCases where the mention occurs but the information is inaccurate persistTrack accuracy of content, not just whether a mention occurred

Translating the paper’s concepts into our measurement

The paper proposes treating visibility not as a single score but as a vector separating discoverability, citation, absorption and economic outcome. Below is the six-lens version Navirang reconstructed for practical measurement, adding the pipeline’s prominence and fidelity stages to that vector — the paper does not present this form.

  1. Discoverability — does our document enter the retrieval candidate set for the question?
  2. Citation — is it actually chosen as the answer’s source?
  3. Prominence — at what position and weight does it appear in the answer (first-appearance position, recommendation inclusion)?
  4. Absorption — do our document’s facts get reflected in the answer’s content?
  5. Fidelity — is the company and service information expressed accurately?
  6. Business outcome — does it lead to actual visits and enquiries?

How those six lenses get measured as concrete metrics (mention rate, citation rate, recommendation appearance, share) is set out in how to measure AI visibility, and the measurement unit we applied to ourselves is the baseline of 40 questions × 7 answer engines = 280 cells in the eight-week improvement series.

Why one test is not enough to judge by

Users looking for the same information phrase the question differently. “Tell me about AEO specialist firms”, “recommend a Korean AEO agency”, “which agency is good for ChatGPT visibility?”, “tell me about GEO specialists”, “which company is good at answer engine optimization” — all paraphrases of the same intent, and when the phrasing changes, the answer and the brands that appear can change too.

So the measurement protocol this survey proposes has repeated runs, paraphrases, a control group and human verification. As a measurement unit:

Engine × Query × Paraphrase × Run × Date

For instance, 20 core questions in several phrasings each, across several engines, repeated on several dates (the specific counts are a matter of your own resources; the paper does not prescribe numbers). “Always first” is not an achievable goal in a probabilistic system, so the KPI should be a frequency metric such as the proportion of repeated measurements in which the brand was among the candidates — for example “candidate inclusion rate across 100 repeated runs of non-brand AEO queries” (an illustration of measurement design, not an industry-standard metric).

Why citation-friendly prose should not be trusted blindly

The part of this survey most likely to be uncomfortable for practitioners. An observation is reported that content rewritten to be citation-oriented can harm the retrieval stage.

That is, an article AI finds easy to cite and an article a retrieval system finds easy to surface are not always the same. An approach that only polishes citation phrasing looks at one stage of the pipeline. So AEO has to be approached not as content rewriting alone but as a compound problem covering crawlability, topical relevance, retrieval, citation and external trust signals together. It is also why our audit checks indexing and accessibility before content.

Navirang does not regard AEO as “how to write sentences AI likes”. We approach it as a search optimization problem that has to be observed stage by stage, from search activation through citation and fidelity, with the same question set measured repeatedly. This paper adds academic weight to that approach — and it is equally a basis for demanding measurement conditions from every claim an agency like ours makes. Please apply the same standard to our figures too.

If you want to know which stage of the pipeline your brand is stuck at, a free AEO audit will tell you.

References

  • Martinez, Olivier. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026).” arXiv:2607.14035 (2026). DOI: 10.48550/arXiv.2607.14035
  • Aggarwal, Pranjal, et al. “GEO: Generative Engine Optimization.” arXiv:2311.09735 (2023) — the first GEO paper, the survey’s starting point

Frequently asked questions

Q What is GEO?

A GEO (generative engine optimization) is optimization work that raises the chance content appears, is cited or has influence in answers generative AI produces. This survey proposes treating GEO not as a single ranking task but as a probabilistic pipeline running through search activation, crawling and indexing, retrieval, reranking and context allocation, citation, prominence, absorption, fidelity and user action.

Q Are AEO and GEO the same concept?

A Their execution mostly overlaps but they are terms with different origins. AEO began as a practitioner term for a search engine's direct-answer areas; GEO was proposed in a 2023 academic paper. The relationship and the differences are covered in a separate comparison article.

Q Does doing GEO guarantee ChatGPT visibility?

A It does not. One of the survey's conclusions is that none of the techniques reviewed demonstrated a causal effect of reliably raising organic discoverability across multiple platforms over a sustained period. If a service claims to guarantee visibility, start by asking for the basis and the measurement conditions.

Q What is the 40% improvement the GEO literature talks about?

A It is the visibility improvement reported by the first GEO paper in 2023 and widely quoted since. This survey points out that the result is valid inside that experimental setup but depends on the condition that documents had already been retrieved and supplied to the model as fixed context. In other words it is evidence that citation can be increased for documents already in the candidate set — not that the chance of being discovered on the open web rises by 40%.

Q What is the difference between discoverability and citation in AI search?

A Discoverability is whether our document enters the retrieval candidate set before AI builds an answer; citation is whether, among documents already in that set, ours is chosen as the source of the answer. They are different stages and their optimization differs. Polishing citation-friendly phrasing acts on the latter; the former belongs to indexing, crawling and topical relevance.

Q How should ChatGPT citation be measured?

A Run the same question set repeatedly and record the proportion of answers whose sources include your domain. The survey likewise proposes a protocol with repeated runs, paraphrases, a control group and human verification. Results differing between runs is normal, so a single screenshot is not a measurement.

Q Which metrics should AEO performance use?

A Proportion metrics from repeated measurement, rather than a fixed position such as "always first". Brand mention rate, source citation rate, retrieval inclusion rate, recommendation inclusion, mention accuracy and variance between engines are the main ones. Navirang's metric definitions and formulas are set out in the AI visibility measurement article.

Q Will adding structured data alone raise AI search visibility?

A No. The factors that reproduced best in this survey were not a single technique such as schema but topical relevance between question and document, and position within the context. Generic heuristics were found to transfer poorly when the environment changed. Structured data belongs as a supporting measure that helps machines read facts.

Q Do we have to measure ChatGPT, Gemini and Perplexity separately?

A Yes. Audits of commercial tools confirmed low source overlap between engines — the set of documents referenced differs considerably by engine even for the same question. One engine's result cannot be used to infer another's, so measure per platform and treat a combined figure only as a reference.

If you need this done rather than read

This article belongs to Research and primary sources. The pages that handle the same subject as work are below.

Related reading

Free audit Call Email Blog