Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models
Jirui QiRaquel FernándezArianna Bisazza
Introduces a ranking-based consistency metric and a multi-parallel benchmark to evaluate factual knowledge alignment across languages in multilingual language models, showing that scaling model size fails to improve cross-lingual consistency and that consistency scores predict whether edited facts transfer across languages.
Multilingual artificial intelligence language models frequently return inconsistent answers to the same factual question when queried in different languages. This lack of reliability poses operational and fairness risks for global systems that serve diverse linguistic populations. The article investigates how consistently multilingual models represent factual knowledge across languages and examines what factors drive this consistency, independent of raw factual accuracy.
To evaluate this, the authors introduced a metric called Ranking-based Consistency (RankC), which measures whether a model ranks potential answers identically across language pairs regardless of whether the answers are correct. The team constructed two fully translated, balanced benchmarks—BMLAMA-17 (spanning 17 languages and roughly 6,800 queries) and BMLAMA-53 (covering 53 languages and over 3,000 queries)—and assessed multiple model families, including encoder-only, encoder-decoder, and modern decoder-only architectures up to 7 billion parameters.
The findings reveal that cross-lingual knowledge consistency is generally low across all tested models, averaging between 23% and 33%. While expanding model size consistently improved factual accuracy, it failed to meaningfully improve consistency; for instance, scaling parameters more than fivefold in the BLOOM model series increased average consistency by only about two percentage points. Furthermore, the analysis demonstrated that high consistency between languages is driven primarily by subword vocabulary overlap—sharing identical word fragments across common scripts—rather than genetic linguistic similarity, geographic proximity, or grammatical structure. Finally, a targeted model editing experiment showed that inserting a new fact in English transferred successfully only to languages sharing high consistency and significant vocabulary overlap with English, while failing to update in linguistically distant languages.
These results indicate that current multilingual models rely heavily on shallow, token-level overlap to share facts rather than maintaining deeper, unified concepts across languages. Organizations deploying these models should not assume that scaling model parameters or updating knowledge in English will naturally resolve discrepancies in other languages. While this shallow transfer creates deployment risks, understanding vocabulary-driven consistency also provides a predictable framework for targeting multilingual model updates.
Decision-makers should exercise caution when deploying multilingual models in mission-critical applications across varied languages and scripts without localized verification. For system updates and model editing, teams should map language dependencies using vocabulary overlap and consistency metrics to identify where targeted interventions in non-English languages are required. Future research must expand evaluation beyond Western-centric fact sets and explore how these consistency patterns behave in very large models beyond 7 billion parameters.
- Paper: Language Models as Knowledge Bases?, Fabio Petroni et al. (2019). Its LAMA probes establish the factual-knowledge ranking approach that helps frame how this paper measures facts stored in multilingual models.
- Paper: Unsupervised Cross-lingual Representation Learning at Scale, Alexis Conneau et al. (2019). XLM-R provides a key multilingual pretrained-model baseline and scaling context for understanding the models whose factual knowledge the source compares across languages.
- Paper: XNLI: Evaluating Cross-lingual Sentence Representations, Alexis Conneau et al. (2018). XNLI introduces a standard way to evaluate cross-lingual transfer, clarifying the multilingual evaluation setting behind the source’s consistency comparisons.
No sufficiently relevant recommendations were found.
