Skip to content

PAPERS

13 papers on AEO and GEO —
what is established
and where they diverge

Summarised so you do not have to read the originals. Every number is stated with its conditions, and the limitations and conflicts of interest the authors declare are carried over as written.

WHAT THIS PAGE IS

A summary of research on AEO, GEO and citation in AI search. For each paper we record, in the same frame, what was measured on what sample · which numbers came out · what limitations the paper states itself.

The claims table at the bottom is the point of this page. A list of summaries leaves you with no way to know which to believe. We collect the places where the papers give different answers to the same question.

13 papersNumbers with conditionsLimits and conflicts includedOur reading marked separately

LIBRARY

13 papers

Grouped by SEO, AEO and GEO. The literature does not divide cleanly into three, so papers are assigned by the question they answer rather than the label they give themselves. The date on each card is the day we verified the original — arXiv papers get revised, so we record the verification date rather than the submission date. Papers keep being added.

SEO — results and traffic

What happened to clicks and traffic once AI entered search

2
arXiv:2608.18352 arXiv preprint (preregistered field experiment)

When AI enters search, referrals fall — and so does user satisfaction

AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence

Stephanie T. Wang (corresponding author) · Jeffrey Gleason · Yakov Bart · Christo Wilson · Danaé Metaxa · University of Pennsylvania · Northeastern University (supported by NSF IIS-2442711)

They fell, and satisfaction fell with them. In a preregistered field experiment (N=1,100), users assigned to AI Mode had an 18.8 percentage-point lower publisher click-through rate, while removing AI Overviews from the screen raised it by 8.8 points. Under the same conditions, trust, usefulness, satisfaction and sense of control all dropped significantly — the common claim that clicks fall but the user experience improves did not hold in this experiment.

arXiv:2605.14021 arXiv preprint

Google AI Overviews at scale — when it appears, what it cites, and whether the citation holds

Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact

Haofei Xu · Umar Iqbal · Jacob M. Montgomery · Washington University in St. Louis

They appeared especially often on question-shaped queries; the sources cited were more trustworthy on average than ordinary search results; but a meaningful share of the generated claims was not supported by the pages cited. Across 55,393 queries observed over 40 days, the activation rate was 13.7% overall and 64.7% on question-shaped queries. Of 98,020 atomic claims, 11.0% lacked support from the cited page, and the authors state that crawl exclusions may have inflated this, putting the lower bound at about 5.3%.

AEO — citation inside the answer

When a citation is attached, is it actually accurate

3
arXiv:2604.25707 arXiv preprint (no conference publication listed)

From citation selection to citation absorption — what counting citations misses

From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms

Zhang Kai · He Xinyue · Yao Jingang · Academic preprint (cs.IR)

Separately, the authors argue. The paper splits 'citation selection' from 'citation absorption'. The first is the engine choosing your page as a source; the second is your page's wording, evidence and structure actually contributing to the final answer. Counting citations alone misses the second.

arXiv:2605.28565 arXiv preprint (CC-BY-SA 4.0)

The source is real but the content does not match — the structure of citation failure

Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs

Yongsik Seo · Wooseok Jeong · Eunyoung Kim · Hyeonseo Jang · Dongha Lee · ParamitaAI · Konkuk University, Dept. of Computer Engineering · Ewha Womans University · Yonsei University, Dept. of Artificial Intelligence · Incheon International Airport Corporation

In a substantial share of cases, no. The authors define the phenomenon of a model citing a real, accessible source while failing along several dimensions, and verified 761,495 citation pairs across ten models from five providers. 30.6% of citations misrepresented their source and 27.1% came from sources unsuited to the domain. Up to 96% of users encountered at least one structurally misleading citation per response.

arXiv:2509.10762 arXiv preprint CONFLICT DISCLOSED

What kind of page do answer engines cite — scoring it across 16 pillars

AI Answer Engine Citation Behavior: Bringing the GEO-16 Framework in B2B SaaS

Arlen Kumar · Leanid Palkhouski · UC Berkeley · Wrodium Research

They do. The authors split measurable page characteristics into 16 pillars (metadata, semantic HTML, structured data, content quality, links), scored each 0–3 and converted them to a composite score between 0 and 1. Across 1,702 citations gathered from 16 English-language B2B SaaS verticals, a higher composite score meant 4.2× the odds of being cited, with metadata and recency, semantic HTML and structured data correlating most strongly in that order. It is an observational study, so this is not causal.

GEO — becoming the material

What generative engines actually select and use

7
arXiv:2311.09735 Published at KDD 2024

GEO — the paper that named generative engine optimization

GEO: Generative Engine Optimization

Pranjal Aggarwal · Vishvak Murahari · Tanmay Rajpurohit · Ashwin Kalyan · Karthik Narasimhan · Ameet Deshpande · Academic research (Princeton and others)

The authors concluded you can. This paper introduced the term GEO (Generative Engine Optimization) and defined it as a black-box optimization framework for raising content visibility in generative engine responses. On their own benchmark, GEO-bench, they reported visibility gains of up to 40%.

arXiv:2602.12187 Accepted at KDD 2026

SAGEO Arena — re-testing GEO across retrieval, reranking and generation

SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization

Sunghwan Kim · Wooseok Jeong · Serin Kim · Sangam Lee · Dongha Lee (corresponding author) · Yonsei University, Dept. of Artificial Intelligence · Konkuk University, Dept. of Computer Engineering

They did not, and in places they hurt. When methods that optimize body text alone were reproduced in an environment with retrieval → reranking → generation, performance at the retrieval and reranking stages fell noticeably. The authors concluded that optimization has to be matched to each pipeline stage.

arXiv:2603.20213 arXiv preprint (code released)

AgenticGEO — what happens when strategy varies per document instead of following fixed rules

AgenticGEO: A Self-Evolving Agentic System for Generative Engine Optimization

Jiaqi Yuan · Jialu Wang · Zihan Wang · Qingyun Sun · Ruijie Wang (corresponding author) · Jianxin Li · Beihang University, School of Computer Science · with independent contributors

It did. The authors argue that the static heuristics proposed in earlier GEO work — add keywords, add citations, add statistics — fail to optimize roughly half of samples, and propose an agentic system that evolves a strategy per document. In their own evaluation they report gains of 26–28% in-domain and 37–70% out-of-domain over the previous best (AutoGEO).

arXiv:2606.20065 arXiv preprint (single author) CONFLICT DISCLOSED

Measurement at scale — where the citations in AI answers actually come from

Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines

Pratyush Kumar · Ranqo (a commercial GEO tracking platform)

Mostly outside the brand's own site. Across 149,912 citations, the brand's own domain accounted for just 2.9%, while company and third-party brand pages made up 75.2%. Appearance rates diverged sharply by brand size. Note that this data comes from a company selling GEO tracking tools, produced with its own platform.

arXiv:2607.14035 arXiv preprint

A critical survey of 45 GEO studies

Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)

Martinez · Academic preprint

It proposes modelling GEO as a nine-stage probabilistic pipeline and argues that discoverability and citation must be treated separately. It also points out that figures like 'a 40% visibility increase' are conditional values that presuppose the document was already retrieved.

arXiv:2605.25517 arXiv preprint CONFLICT DISCLOSED

When only one of two documents is cited, what decided it

What Gets Cited: Competitive GEO in AI Answer Engines

Rahul Vishwakarma · Shushant Kumar · Ratnesh Jamidar · Sprinklr (Gurugram, India · Dubai, UAE) — a commercial AI/CX platform

Topical match and list position dominated, with pricing information and a recent timestamp next. The authors ran 252,000 pairwise comparisons injecting two candidate documents that differed in **exactly one factor**. Topical mismatch, missing pricing, an old timestamp and second position in a list were conditions that lost essentially always, and the gap between assertive and hedged phrasing was also large.

arXiv:2605.00012 arXiv preprint (14 pages, no conference publication listed)

Can AI search overviews be manipulated?

Exploring LLM biases to manipulate AI search overview

Roman Smirnov · Single author (no affiliation stated)

The author reports that it can. He trained a small language model with reinforcement learning to rewrite search snippets, and writes that the rewritten snippets drew the LLM overview's selection in most cases. He also shows that selection depends on **relative comparison between candidates** rather than absolute quality, and that context poisoning attacks can produce inaccurate or harmful results. No success rate is reported numerically.

Shared conditions

Premises underlying all three — language and culture, axes we do not control

1
arXiv:2407.05502 Published at NAACL 2025

Documents in the language you asked in get cited — information disparity in multilingual AI

Faux Polyglot: A Study on Information Disparity in Multilingual Large Language Models

Nikhil Sharma · Kenton Murray · Ziang Xiao · Center for Speech and Language Processing, Johns Hopkins University

Yes. At both the retrieval and generation stages, models systematically preferred documents in the language of the query. 68% of the top 10 retrieval results shared the query's language, and same-language documents were cited 42% of the time in generation versus 29.58% for foreign-language ones. When only foreign-language documents were used, high-resource languages (English, German) were preferred over low-resource ones (Hindi, Arabic). The authors warn that this can reinforce language-specific information cocoons and further marginalise low-resource-language perspectives.

WHERE THEY DISAGREE

The claims table

What each paper answered to the same question. Entries where they diverge are marked.

01

Does rewriting content to be AI-friendly raise citation?

THE PAPERS DIVERGE

NAVIRANG'S JUDGEMENT — NOT THE PAPERS' CONCLUSION

We do not treat 'rewriting copy raises citation' as a general proposition. What matters is where the three papers diverge — both papers reporting a gain measured in **environments with retrieval and reranking removed** (a benchmark, or the generation stage of open models), and SAGEO Arena, which reproduced the work with those stages present, saw performance fall. So the right reading is not 'copy optimization is pointless' but 'it only means something once you are past retrieval'. Our order of work — index and crawler access first, paragraph structure after — follows that reading. One more thing: AgenticGEO's observation that static heuristics fail on about half of samples means the 'add citations, add statistics' checklist circulating in the industry should not be applied identically to every document.

02

So what should you fix?

NAVIRANG'S JUDGEMENT — NOT THE PAPERS' CONCLUSION

Both papers point the same way — structure and evidence, not adjectives. It is the same direction as our requirement that citable content carry paragraphs that still make sense when cut, with sources and dates stated.

03

Is fixing your own site enough?

NAVIRANG'S JUDGEMENT — NOT THE PAPERS' CONCLUSION

We treat your own pages as necessary but not sufficient. That said, this figure came from a vendor using its own platform, and the sample skews to SaaS and fintech, so we do not carry it over to industries here. We take the direction and measure our own vertical.

04

What should performance actually be counted in?

NAVIRANG'S JUDGEMENT — NOT THE PAPERS' CONCLUSION

Navirang counts mentions and citations separately and publishes the denominator (questions × engines) alongside. Splitting out 'absorption' is not yet part of our protocol — adding it would require fixing the adjudication criteria in writing first.

05

If a source is attached, is the content accurate?

NAVIRANG'S JUDGEMENT — NOT THE PAPERS' CONCLUSION

Two studies on different material — technical Q&A queries and Google AI Overviews — point the same way: a source being attached does not mean the sentence came from that source. For a brand that means two things. You cannot relax because you were cited (words that are not in your document may be attributed to your domain), and conversely an inaccurate description of a competitor can circulate carrying that company's source. It is why we record the **factual accuracy of the description** separately from whether a citation occurred.

06

When AI answers enter search, what happens to site traffic?

NAVIRANG'S JUDGEMENT — NOT THE PAPERS' CONCLUSION

The traffic decline is real, and it is the **only item in this library confirmed causally** — the rest are observational. Both studies are US Google, though, and neither touched Naver or Korean, so we do not carry the magnitudes over as expectations here. What matters practically is the activation rate: at 64.7% on question-shaped queries, a brand without documents that answer questions is not even a candidate on that screen.

07

Can findings from English-language research be applied directly to Korean?

NAVIRANG'S JUDGEMENT — NOT THE PAPERS' CONCLUSION

They cannot. Every other paper here rests on English-language data, and if documents in the query's language are systematically preferred, results for Korean queries are bound to differ. Read favourably, that means Korean-language assets are worth having; read unfavourably, it means that on topics with few Korean documents, AI may carry over an English-language perspective wholesale. **But Korean was not in that experiment** — take the direction and measure your own brand's Korean figures. It is also why we keep Naver AI Search in the engine set.

08

Can visibility be guaranteed?

NAVIRANG'S JUDGEMENT — NOT THE PAPERS' CONCLUSION

We do not guarantee it. No paper has been able to guarantee visibility in a commercial engine, and engine internals are not public. One point is worth making, though — we do not decline because manipulation is **technically impossible**. There is research showing a rewritten snippet can draw selection, and we do not build that. It is abuse. And if selection depends on relative comparison between candidates, your document can stay the same while the result changes because a competitor's document changed. That is why we measure repeatedly under the same conditions rather than once.

HOW TO READ

Reading this library

  1. 01

    Read numbers with their conditions

    The same '40%' means something entirely different depending on which benchmark, how many queries and which engine produced it. That is why every figure on this page carries its conditions.

  2. 02

    Look at who produced it

    Research from an academic institution and research from a company selling a tool do not carry the same weight. Papers with a disclosed conflict of interest are marked, and we quote the authors' own sentence.

  3. 03

    Read the limitations the paper states

    The better the paper, the more carefully it states its limits. That is where you find what the sample skews toward, whether causality is claimed, and what the study could not see.

  4. 04

    Keep our judgement separate from the paper's claim

    The blue box in the claims table is Navirang's reading, not the paper's conclusion. We separate them on screen so that we never borrow a paper's name to make our own argument.

  5. 05

    You still have to measure your own numbers

    No figure here predicts the performance of a particular brand or industry. Check how far the sample and conditions overlap with your situation — in the end the answer is to measure with your own question set.

If you find a missing paper or something summarised incorrectly, tell us. We re-read the original and fix it.

If you want your own brand's numbers rather than a paper's

Every figure here came from someone else's sample. Send us a URL and we ask all 7 answer engines directly and measure yours.

We reply within one business day.

Free audit Call Email Blog