PAPER
The source is real but the content does not match — the structure of citation failure
If an AI cites a source that genuinely exists, can that citation be trusted?
Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs
- Authors
- Yongsik Seo · Wooseok Jeong · Eunyoung Kim · Hyeonseo Jang · Dongha Lee
- Affiliation
- ParamitaAI · Konkuk University, Dept. of Computer Engineering · Ewha Womans University · Yonsei University, Dept. of Artificial Intelligence · Incheon International Airport Corporation
- Venue
- arXiv preprint (CC-BY-SA 4.0)
- Submitted
- 2026-05-28
- arXiv
- arXiv:2605.28565
- We verified
- 2026-08-24
WHAT THE PAPER SAYS
In a substantial share of cases, no. The authors define the phenomenon of a model citing a real, accessible source while failing along several dimensions, and verified 761,495 citation pairs across ten models from five providers. 30.6% of citations misrepresented their source and 27.1% came from sources unsuited to the domain. Up to 96% of users encountered at least one structurally misleading citation per response.
Read the original on arXiv ↗METHOD
How it was measured
Read the conditions before the numbers. The same figure means something different under a different sample or environment.
- Queries
- 11,200 real questions from 28 Stack Exchange communities
- Systems
- 10 LLMs from five providers (OpenAI, Anthropic, Google, xAI, Perplexity)
- Collection
- 112,000 responses · 761,495 evaluable citation pairs
- Data cutoff
- Stack Exchange archive as of 2025-12-31
FINDINGS
What came out
- Fidelity failure
- 30.6% of citations misrepresent the source (Fidelity Failure Rate)
- Suitability failure
- 27.1% came from a source unsuited to the domain (Suitability Failure Rate)
- User exposure
- Up to 96% of users meet at least one structurally misleading citation per response
- Intent–purpose misalignment
- 5.1%
- Differences between providers
- Provider-level differences explain 88–96% of the variance in citation quality — which engine you use matters a great deal
- ⚠️ Which way these numbers point
- The authors state that crawl-failure bias, particularly for forum and Q&A sources, makes the reported failure rates a conservative lower bound. The real figures may be higher
LIMITATIONS
Limitations the authors state themselves
Not our criticism — this is what the authors wrote in the paper.
- Reported failure rates are a conservative lower bound because of crawl-failure bias, particularly for forum and Q&A sources.
- An LLM judge was used for adjudication, so the judge model's error is mixed into the results.
- Failure types were split by a predefined matrix, and that division involves judgement.
- Language and cultural coverage is limited.
- A point-in-time snapshot; results will shift as models change.
NAVIRANG'S READING — NOT THE PAPER'S CONCLUSION
This is the most uncomfortable number in the library. We count 'mentions' and 'citations' separately in an audit; this paper asks the question after that — **does being cited mean the citation is accurate?** If 30.6% misrepresent their source, it does not. Two things follow in practice. First, you cannot relax because the brand was cited: it may have been cited with words that are not in your document attributed to your domain, and that is a correction task. Second, the finding that provider differences explain 88–96% of the variance strengthens the case for measuring engines separately — it is the reason we count each engine as its own cell. That said, the data is English-language technical questions from Stack Exchange. There is no guarantee the same proportions hold for Korean queries in other industries.
IN COMPARISON
Where it diverges from other papers
The claims-table entries this paper appears in.
RELATED
Related reading
If you want your own brand's numbers rather than a paper's
Every figure here came from someone else's sample. Send us a URL and we ask all 7 answer engines directly and measure yours.
We reply within one business day.