PAPER
From citation selection to citation absorption — what counting citations misses
Is being cited enough, or does whether the page was actually used in the answer have to be measured separately?
From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms
- Authors
- Zhang Kai · He Xinyue · Yao Jingang
- Affiliation
- Academic preprint (cs.IR)
- Venue
- arXiv preprint (no conference publication listed)
- Submitted
- 2026-04-28 · revised 2026-04-29 (v2)
- arXiv
- arXiv:2604.25707
- We verified
- 2026-08-22
WHAT THE PAPER SAYS
Separately, the authors argue. The paper splits 'citation selection' from 'citation absorption'. The first is the engine choosing your page as a source; the second is your page's wording, evidence and structure actually contributing to the final answer. Counting citations alone misses the second.
Read the original on arXiv ↗METHOD
How it was measured
Read the conditions before the numbers. The same figure means something different under a different sample or environment.
- Prompts
- 602 controlled prompts
- Citations
- 21,143 retrieval-layer citations · 23,745 citation-level features
- Pages collected
- 18,151 · 72 extracted features
- Engines
- ChatGPT · Google AI Overviews / Gemini · Perplexity
FINDINGS
What came out
- Citation behaviour by engine
- Perplexity and Google cite more sources on average, while ChatGPT cites fewer but shows markedly higher average citation influence per page
- What absorbed pages have in common
- Longer · more structured · semantically aligned · and rich in extractable evidence such as definitions, numeric facts, comparisons and procedural steps
- Proposal
- GEO performance must be measured beyond citation counts, treating answer-level absorption as a separate outcome
LIMITATIONS
Limitations the authors state themselves
Not our criticism — this is what the authors wrote in the paper.
- A preprint with no conference publication listed.
- Feature extraction and absorption judgement are model-based, so classifier error is mixed into the results.
- It observes three engines at one point in time; results may differ as engines change.
NAVIRANG'S READING — NOT THE PAPER'S CONCLUSION
This runs in the same direction as our practice of counting mentions and citations separately in an audit, but goes one step further. We split 'did the name appear (mention) / did the domain enter the sources (citation)'; this paper adds 'once in the sources, was it actually used in the answer (absorption)'. The practically useful part is what absorbed pages look like — documents containing definitions, numbers, comparisons and procedures. That is very nearly the same thing we mean when we require a paragraph that still makes sense after being cut.
IN COMPARISON
Where it diverges from other papers
The claims-table entries this paper appears in.
RELATED
Related reading
If you want your own brand's numbers rather than a paper's
Every figure here came from someone else's sample. Send us a URL and we ask all 7 answer engines directly and measure yours.
We reply within one business day.