Skip to content

PAPER

Documents in the language you asked in get cited — information disparity in multilingual AI

If you ask the same question in a different language, does the AI find different documents and give a different answer?

Faux Polyglot: A Study on Information Disparity in Multilingual Large Language Models

Authors
Nikhil Sharma · Kenton Murray · Ziang Xiao
Affiliation
Center for Speech and Language Processing, Johns Hopkins University
Venue
Published at NAACL 2025
Submitted
2024-07-07 · revised 2025-02-11 (v3)
arXiv
arXiv:2407.05502
We verified
2026-08-24

WHAT THE PAPER SAYS

Yes. At both the retrieval and generation stages, models systematically preferred documents in the language of the query. 68% of the top 10 retrieval results shared the query's language, and same-language documents were cited 42% of the time in generation versus 29.58% for foreign-language ones. When only foreign-language documents were used, high-resource languages (English, German) were preferred over low-resource ones (Hindi, Arabic). The authors warn that this can reinforce language-specific information cocoons and further marginalise low-resource-language perspectives.

Read the original on arXiv ↗

METHOD

How it was measured

Read the conditions before the numbers. The same figure means something different under a different sample or environment.

Languages
English (en) · Hindi (hi) · German (de) · Arabic (ar) · Chinese (zh) — five, spanning four scripts and mixing high-resource (en, de, zh) with low-resource (ar, hi). Korean is not included
Retrieval models
8 — OpenAI (ada-002, embedding-3-small/large) · Cohere (multilingual-light-v3.0, v3.0) · Voyage (voyage-2, large-2, large-2-instruct)
Generation models
5 — gpt-3.5-turbo · gpt-4o · aya-23-8B · aya-23-35B · claude-3-opus
Queries
27 factual and 16 opinion queries, each translated into all five languages
Documents
170 synthetic documents — ten core facts written manually, then expanded and machine-translated

FINDINGS

What came out

Retrieval stage
68% of the top 10 were documents in the query's language. Mean z-score of same-language documents 1.03 vs -0.25 for foreign-language
Generation stage
Same-language documents cited 42% vs foreign-language 29.58%. Both cited 8.55%; neither used 19.8%
Ranking among foreign languages
When only foreign-language documents were used: English 49.28% > German 44.48% > Chinese 42.05% > Hindi 40.77% > Arabic 39.77% — tilted toward high-resource languages
The authors' warning
Multilingual capability may in fact reinforce language-specific information cocoons and filter bubbles, further marginalising low-resource-language perspectives
⚠️ On Korean
Korean was not included in this study. Nor does it address whether Korean would count as high- or low-resource. Take the direction, but do not derive Korean figures from it

LIMITATIONS

Limitations the authors state themselves

Not our criticism — this is what the authors wrote in the paper.

  • Only five languages were covered — seeing interactions between scripts and language families would need a much larger pool. Very low-resource languages such as Swahili were excluded because building a document set was difficult.
  • It used a controlled synthetic task rather than empirical data. That choice isolates a confounder (the model's parametric knowledge) but trades off against realism.
  • The influence of pretraining was not examined in depth.
  • It looked at language differences and not cultural ones — narrative differences between cultures sharing a language may be mixed into the results.
  • Only one RAG architecture (cosine-similarity retrieval) was evaluated. Others, such as summarisation or reranking, were not covered.

NAVIRANG'S READING — NOT THE PAPER'S CONCLUSION

Every other paper here rests on English-language data, while our clients ask in Korean, in Korea. We had been attaching the caveat that 'there is no guarantee this holds in Korean' to each entry, but that was our supposition; this paper gives the supposition a basis — documents in the query's language are systematically preferred at both retrieval and generation. Practically it cuts two ways. In our favour: **Korean queries pull Korean documents**, which is what makes Korean-language assets on your own domain worth having. Against: a brand holding only English assets tends to lose on Korean queries, and on topics where Korean documents are scarce, AI may carry over an English-language perspective wholesale. But **Korean is not in this experiment.** Take the direction only; the numbers have to be replaced by our own measurements.

IN COMPARISON

Where it diverges from other papers

The claims-table entries this paper appears in.

If you want your own brand's numbers rather than a paper's

Every figure here came from someone else's sample. Send us a URL and we ask all 7 answer engines directly and measure yours.

We reply within one business day.

Free audit Call Email Blog