Skip to content

PAPER

The source is real but the content does not match — the structure of citation failure

If an AI cites a source that genuinely exists, can that citation be trusted?

Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs

Authors
Yongsik Seo · Wooseok Jeong · Eunyoung Kim · Hyeonseo Jang · Dongha Lee
Affiliation
ParamitaAI · Konkuk University, Dept. of Computer Engineering · Ewha Womans University · Yonsei University, Dept. of Artificial Intelligence · Incheon International Airport Corporation
Venue
arXiv preprint (CC-BY-SA 4.0)
Submitted
2026-05-28
arXiv
arXiv:2605.28565
We verified
2026-08-24

WHAT THE PAPER SAYS

In a substantial share of cases, no. The authors define the phenomenon of a model citing a real, accessible source while failing along several dimensions, and verified 761,495 citation pairs across ten models from five providers. 30.6% of citations misrepresented their source and 27.1% came from sources unsuited to the domain. Up to 96% of users encountered at least one structurally misleading citation per response.

Read the original on arXiv ↗

METHOD

How it was measured

Read the conditions before the numbers. The same figure means something different under a different sample or environment.

Queries
11,200 real questions from 28 Stack Exchange communities
Systems
10 LLMs from five providers (OpenAI, Anthropic, Google, xAI, Perplexity)
Collection
112,000 responses · 761,495 evaluable citation pairs
Data cutoff
Stack Exchange archive as of 2025-12-31

FINDINGS

What came out

Fidelity failure
30.6% of citations misrepresent the source (Fidelity Failure Rate)
Suitability failure
27.1% came from a source unsuited to the domain (Suitability Failure Rate)
User exposure
Up to 96% of users meet at least one structurally misleading citation per response
Intent–purpose misalignment
5.1%
Differences between providers
Provider-level differences explain 88–96% of the variance in citation quality — which engine you use matters a great deal
⚠️ Which way these numbers point
The authors state that crawl-failure bias, particularly for forum and Q&A sources, makes the reported failure rates a conservative lower bound. The real figures may be higher

LIMITATIONS

Limitations the authors state themselves

Not our criticism — this is what the authors wrote in the paper.

  • Reported failure rates are a conservative lower bound because of crawl-failure bias, particularly for forum and Q&A sources.
  • An LLM judge was used for adjudication, so the judge model's error is mixed into the results.
  • Failure types were split by a predefined matrix, and that division involves judgement.
  • Language and cultural coverage is limited.
  • A point-in-time snapshot; results will shift as models change.

NAVIRANG'S READING — NOT THE PAPER'S CONCLUSION

This is the most uncomfortable number in the library. We count 'mentions' and 'citations' separately in an audit; this paper asks the question after that — **does being cited mean the citation is accurate?** If 30.6% misrepresent their source, it does not. Two things follow in practice. First, you cannot relax because the brand was cited: it may have been cited with words that are not in your document attributed to your domain, and that is a correction task. Second, the finding that provider differences explain 88–96% of the variance strengthens the case for measuring engines separately — it is the reason we count each engine as its own cell. That said, the data is English-language technical questions from Stack Exchange. There is no guarantee the same proportions hold for Korean queries in other industries.

IN COMPARISON

Where it diverges from other papers

The claims-table entries this paper appears in.

If you want your own brand's numbers rather than a paper's

Every figure here came from someone else's sample. Send us a URL and we ask all 7 answer engines directly and measure yours.

We reply within one business day.

Free audit Call Email Blog