PAPER
What kind of page do answer engines cite — scoring it across 16 pillars
Do the pages answer engines cite share measurable technical characteristics? If so, which ones correlate most strongly?
AI Answer Engine Citation Behavior: Bringing the GEO-16 Framework in B2B SaaS
- Authors
- Arlen Kumar · Leanid Palkhouski
- Affiliation
- UC Berkeley · Wrodium Research
- Venue
- arXiv preprint
- Submitted
- 2026-09-10
- arXiv
- arXiv:2509.10762
- We verified
- 2026-08-24
WHAT THE PAPER SAYS
They do. The authors split measurable page characteristics into 16 pillars (metadata, semantic HTML, structured data, content quality, links), scored each 0–3 and converted them to a composite score between 0 and 1. Across 1,702 citations gathered from 16 English-language B2B SaaS verticals, a higher composite score meant 4.2× the odds of being cited, with metadata and recency, semantic HTML and structured data correlating most strongly in that order. It is an observational study, so this is not causal.
Read the original on arXiv ↗⚠️ THIS PAPER HAS A DISCLOSED CONFLICT OF INTEREST
Wrodium Research is listed alongside the authors' affiliations, but the paper carries no conflict-of-interest declaration. Whether a commercial interest exists cannot be determined from the paper alone, so we state only what is on record.
METHOD
How it was measured
Read the conditions before the numbers. The same figure means something different under a different sample or environment.
- Framework
- GEO-16 — 16 pillars scored on a 0–3 scale, converted to a composite score G between 0 and 1
- Queries
- 70 prompts aimed at 16 B2B SaaS verticals
- Collection
- 1,702 citations · 1,100 unique URLs audited
- Engines
- Brave · Google AI Overviews · Perplexity (three)
FINDINGS
What came out
- Composite score and citation
- Logistic regression odds ratio 4.2 (95% CI 3.1, 5.7) — higher-scoring pages had that much higher odds of being cited
- Strongest correlations
- Metadata and recency r=0.68 · semantic HTML r=0.65 · structured data r=0.63
- Mean score by engine
- Brave 0.727 · Google AI Overviews 0.687 · Perplexity 0.300 — engines differ sharply in the character of the pages they cite
- Citation rate by engine
- Brave 78% · Google AI Overviews 72% · Perplexity 45%
- ⚠️ The conditions on these numbers
- The material is English-language B2B SaaS content and the design is observational. These are correlations, not causes, and the authors acknowledge the risk of unobserved confounding
LIMITATIONS
Limitations the authors state themselves
Not our criticism — this is what the authors wrote in the paper.
- Limited to English-language B2B SaaS content — the authors state explicitly that results may differ in other languages or industries.
- An observational design carries a risk of unobserved confounding — this is not a causal claim.
- Single-point collection limits generalization.
- No conflict-of-interest declaration.
NAVIRANG'S READING — NOT THE PAPER'S CONCLUSION
The order matches what we look at first in practice — metadata and recency, semantic HTML, structured data are the top three. It adds outside evidence for why an audit starts with crawler access and structured data. Two cautions, though. First, **it is observational.** The finding is not 'add structured data and you get cited' but 'pages that were cited tend to have structured data'. A well-maintained site may simply have both. Second, Perplexity's mean score of 0.300 is less than half the other engines — that means it also cites lower-scoring pages, not that it is a worse engine. We read it as evidence that engines select on different criteria.
IN COMPARISON
Where it diverges from other papers
The claims-table entries this paper appears in.
RELATED
Related reading
If you want your own brand's numbers rather than a paper's
Every figure here came from someone else's sample. Send us a URL and we ask all 7 answer engines directly and measure yours.
We reply within one business day.