Skip to content

PAPER

What kind of page do answer engines cite — scoring it across 16 pillars

Do the pages answer engines cite share measurable technical characteristics? If so, which ones correlate most strongly?

AI Answer Engine Citation Behavior: Bringing the GEO-16 Framework in B2B SaaS

Authors
Arlen Kumar · Leanid Palkhouski
Affiliation
UC Berkeley · Wrodium Research
Venue
arXiv preprint
Submitted
2026-09-10
arXiv
arXiv:2509.10762
We verified
2026-08-24

WHAT THE PAPER SAYS

They do. The authors split measurable page characteristics into 16 pillars (metadata, semantic HTML, structured data, content quality, links), scored each 0–3 and converted them to a composite score between 0 and 1. Across 1,702 citations gathered from 16 English-language B2B SaaS verticals, a higher composite score meant 4.2× the odds of being cited, with metadata and recency, semantic HTML and structured data correlating most strongly in that order. It is an observational study, so this is not causal.

Read the original on arXiv ↗

⚠️ THIS PAPER HAS A DISCLOSED CONFLICT OF INTEREST

Wrodium Research is listed alongside the authors' affiliations, but the paper carries no conflict-of-interest declaration. Whether a commercial interest exists cannot be determined from the paper alone, so we state only what is on record.

METHOD

How it was measured

Read the conditions before the numbers. The same figure means something different under a different sample or environment.

Framework
GEO-16 — 16 pillars scored on a 0–3 scale, converted to a composite score G between 0 and 1
Queries
70 prompts aimed at 16 B2B SaaS verticals
Collection
1,702 citations · 1,100 unique URLs audited
Engines
Brave · Google AI Overviews · Perplexity (three)

FINDINGS

What came out

Composite score and citation
Logistic regression odds ratio 4.2 (95% CI 3.1, 5.7) — higher-scoring pages had that much higher odds of being cited
Strongest correlations
Metadata and recency r=0.68 · semantic HTML r=0.65 · structured data r=0.63
Mean score by engine
Brave 0.727 · Google AI Overviews 0.687 · Perplexity 0.300 — engines differ sharply in the character of the pages they cite
Citation rate by engine
Brave 78% · Google AI Overviews 72% · Perplexity 45%
⚠️ The conditions on these numbers
The material is English-language B2B SaaS content and the design is observational. These are correlations, not causes, and the authors acknowledge the risk of unobserved confounding

LIMITATIONS

Limitations the authors state themselves

Not our criticism — this is what the authors wrote in the paper.

  • Limited to English-language B2B SaaS content — the authors state explicitly that results may differ in other languages or industries.
  • An observational design carries a risk of unobserved confounding — this is not a causal claim.
  • Single-point collection limits generalization.
  • No conflict-of-interest declaration.

NAVIRANG'S READING — NOT THE PAPER'S CONCLUSION

The order matches what we look at first in practice — metadata and recency, semantic HTML, structured data are the top three. It adds outside evidence for why an audit starts with crawler access and structured data. Two cautions, though. First, **it is observational.** The finding is not 'add structured data and you get cited' but 'pages that were cited tend to have structured data'. A well-maintained site may simply have both. Second, Perplexity's mean score of 0.300 is less than half the other engines — that means it also cites lower-scoring pages, not that it is a worse engine. We read it as evidence that engines select on different criteria.

IN COMPARISON

Where it diverges from other papers

The claims-table entries this paper appears in.

If you want your own brand's numbers rather than a paper's

Every figure here came from someone else's sample. Send us a URL and we ask all 7 answer engines directly and measure yours.

We reply within one business day.

Free audit Call Email Blog