Skip to content

PAPER

From citation selection to citation absorption — what counting citations misses

Is being cited enough, or does whether the page was actually used in the answer have to be measured separately?

From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms

Authors
Zhang Kai · He Xinyue · Yao Jingang
Affiliation
Academic preprint (cs.IR)
Venue
arXiv preprint (no conference publication listed)
Submitted
2026-04-28 · revised 2026-04-29 (v2)
arXiv
arXiv:2604.25707
We verified
2026-08-22

WHAT THE PAPER SAYS

Separately, the authors argue. The paper splits 'citation selection' from 'citation absorption'. The first is the engine choosing your page as a source; the second is your page's wording, evidence and structure actually contributing to the final answer. Counting citations alone misses the second.

Read the original on arXiv ↗

METHOD

How it was measured

Read the conditions before the numbers. The same figure means something different under a different sample or environment.

Prompts
602 controlled prompts
Citations
21,143 retrieval-layer citations · 23,745 citation-level features
Pages collected
18,151 · 72 extracted features
Engines
ChatGPT · Google AI Overviews / Gemini · Perplexity

FINDINGS

What came out

Citation behaviour by engine
Perplexity and Google cite more sources on average, while ChatGPT cites fewer but shows markedly higher average citation influence per page
What absorbed pages have in common
Longer · more structured · semantically aligned · and rich in extractable evidence such as definitions, numeric facts, comparisons and procedural steps
Proposal
GEO performance must be measured beyond citation counts, treating answer-level absorption as a separate outcome

LIMITATIONS

Limitations the authors state themselves

Not our criticism — this is what the authors wrote in the paper.

  • A preprint with no conference publication listed.
  • Feature extraction and absorption judgement are model-based, so classifier error is mixed into the results.
  • It observes three engines at one point in time; results may differ as engines change.

NAVIRANG'S READING — NOT THE PAPER'S CONCLUSION

This runs in the same direction as our practice of counting mentions and citations separately in an audit, but goes one step further. We split 'did the name appear (mention) / did the domain enter the sources (citation)'; this paper adds 'once in the sources, was it actually used in the answer (absorption)'. The practically useful part is what absorbed pages look like — documents containing definitions, numbers, comparisons and procedures. That is very nearly the same thing we mean when we require a paragraph that still makes sense after being cut.

If you want your own brand's numbers rather than a paper's

Every figure here came from someone else's sample. Send us a URL and we ask all 7 answer engines directly and measure yours.

We reply within one business day.

Free audit Call Email Blog