PAPER
Documents in the language you asked in get cited — information disparity in multilingual AI
If you ask the same question in a different language, does the AI find different documents and give a different answer?
Faux Polyglot: A Study on Information Disparity in Multilingual Large Language Models
- Authors
- Nikhil Sharma · Kenton Murray · Ziang Xiao
- Affiliation
- Center for Speech and Language Processing, Johns Hopkins University
- Venue
- Published at NAACL 2025
- Submitted
- 2024-07-07 · revised 2025-02-11 (v3)
- arXiv
- arXiv:2407.05502
- We verified
- 2026-08-24
WHAT THE PAPER SAYS
Yes. At both the retrieval and generation stages, models systematically preferred documents in the language of the query. 68% of the top 10 retrieval results shared the query's language, and same-language documents were cited 42% of the time in generation versus 29.58% for foreign-language ones. When only foreign-language documents were used, high-resource languages (English, German) were preferred over low-resource ones (Hindi, Arabic). The authors warn that this can reinforce language-specific information cocoons and further marginalise low-resource-language perspectives.
Read the original on arXiv ↗METHOD
How it was measured
Read the conditions before the numbers. The same figure means something different under a different sample or environment.
- Languages
- English (en) · Hindi (hi) · German (de) · Arabic (ar) · Chinese (zh) — five, spanning four scripts and mixing high-resource (en, de, zh) with low-resource (ar, hi). Korean is not included
- Retrieval models
- 8 — OpenAI (ada-002, embedding-3-small/large) · Cohere (multilingual-light-v3.0, v3.0) · Voyage (voyage-2, large-2, large-2-instruct)
- Generation models
- 5 — gpt-3.5-turbo · gpt-4o · aya-23-8B · aya-23-35B · claude-3-opus
- Queries
- 27 factual and 16 opinion queries, each translated into all five languages
- Documents
- 170 synthetic documents — ten core facts written manually, then expanded and machine-translated
FINDINGS
What came out
- Retrieval stage
- 68% of the top 10 were documents in the query's language. Mean z-score of same-language documents 1.03 vs -0.25 for foreign-language
- Generation stage
- Same-language documents cited 42% vs foreign-language 29.58%. Both cited 8.55%; neither used 19.8%
- Ranking among foreign languages
- When only foreign-language documents were used: English 49.28% > German 44.48% > Chinese 42.05% > Hindi 40.77% > Arabic 39.77% — tilted toward high-resource languages
- The authors' warning
- Multilingual capability may in fact reinforce language-specific information cocoons and filter bubbles, further marginalising low-resource-language perspectives
- ⚠️ On Korean
- Korean was not included in this study. Nor does it address whether Korean would count as high- or low-resource. Take the direction, but do not derive Korean figures from it
LIMITATIONS
Limitations the authors state themselves
Not our criticism — this is what the authors wrote in the paper.
- Only five languages were covered — seeing interactions between scripts and language families would need a much larger pool. Very low-resource languages such as Swahili were excluded because building a document set was difficult.
- It used a controlled synthetic task rather than empirical data. That choice isolates a confounder (the model's parametric knowledge) but trades off against realism.
- The influence of pretraining was not examined in depth.
- It looked at language differences and not cultural ones — narrative differences between cultures sharing a language may be mixed into the results.
- Only one RAG architecture (cosine-similarity retrieval) was evaluated. Others, such as summarisation or reranking, were not covered.
NAVIRANG'S READING — NOT THE PAPER'S CONCLUSION
Every other paper here rests on English-language data, while our clients ask in Korean, in Korea. We had been attaching the caveat that 'there is no guarantee this holds in Korean' to each entry, but that was our supposition; this paper gives the supposition a basis — documents in the query's language are systematically preferred at both retrieval and generation. Practically it cuts two ways. In our favour: **Korean queries pull Korean documents**, which is what makes Korean-language assets on your own domain worth having. Against: a brand holding only English assets tends to lose on Korean queries, and on topics where Korean documents are scarce, AI may carry over an English-language perspective wholesale. But **Korean is not in this experiment.** Take the direction only; the numbers have to be replaced by our own measurements.
IN COMPARISON
Where it diverges from other papers
The claims-table entries this paper appears in.
RELATED
Related reading
If you want your own brand's numbers rather than a paper's
Every figure here came from someone else's sample. Send us a URL and we ask all 7 answer engines directly and measure yours.
We reply within one business day.