Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models

Jirui QiRaquel FernándezArianna Bisazza

article2023EMNLP114 citationsOutstanding Paper (Multilinguality and Linguistic Diversity Track)

Introduces a ranking-based consistency metric and a multi-parallel benchmark to evaluate factual knowledge alignment across languages in multilingual language models, showing that scaling model size fails to improve cross-lingual consistency and that consistency scores predict whether edited facts transfer across languages.

Listen

Multilingual artificial intelligence language models frequently return inconsistent answers to the same factual question when queried in different languages. This lack of reliability poses operational and fairness risks for global systems that serve diverse linguistic populations. The article investigates how consistently multilingual models represent factual knowledge across languages and examines what factors drive this consistency, independent of raw factual accuracy.

To evaluate this, the authors introduced a metric called Ranking-based Consistency (RankC), which measures whether a model ranks potential answers identically across language pairs regardless of whether the answers are correct. The team constructed two fully translated, balanced benchmarks—BMLAMA-17 (spanning 17 languages and roughly 6,800 queries) and BMLAMA-53 (covering 53 languages and over 3,000 queries)—and assessed multiple model families, including encoder-only, encoder-decoder, and modern decoder-only architectures up to 7 billion parameters.

The findings reveal that cross-lingual knowledge consistency is generally low across all tested models, averaging between 23% and 33%. While expanding model size consistently improved factual accuracy, it failed to meaningfully improve consistency; for instance, scaling parameters more than fivefold in the BLOOM model series increased average consistency by only about two percentage points. Furthermore, the analysis demonstrated that high consistency between languages is driven primarily by subword vocabulary overlap—sharing identical word fragments across common scripts—rather than genetic linguistic similarity, geographic proximity, or grammatical structure. Finally, a targeted model editing experiment showed that inserting a new fact in English transferred successfully only to languages sharing high consistency and significant vocabulary overlap with English, while failing to update in linguistically distant languages.

These results indicate that current multilingual models rely heavily on shallow, token-level overlap to share facts rather than maintaining deeper, unified concepts across languages. Organizations deploying these models should not assume that scaling model parameters or updating knowledge in English will naturally resolve discrepancies in other languages. While this shallow transfer creates deployment risks, understanding vocabulary-driven consistency also provides a predictable framework for targeting multilingual model updates.

Decision-makers should exercise caution when deploying multilingual models in mission-critical applications across varied languages and scripts without localized verification. For system updates and model editing, teams should map language dependencies using vocabulary overlap and consistency metrics to identify where targeted interventions in non-English languages are required. Future research must expand evaluation beyond Western-centric fact sets and explore how these consistency patterns behave in very large models beyond 7 billion parameters.

No sufficiently relevant recommendations were found.

Cover for Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models

Abstract

Multilingual large-scale Pretrained Language Models (PLMs) have been shown to store considerable amounts of factual knowledge, but large variations are observed across languages. With the ultimate goal of ensuring that users with different language backgrounds obtain consistent feedback from the same model, we study the cross-lingual consistency (CLC) of factual knowledge in various multilingual PLMs. To this end, we propose a Ranking-based Consistency (RankC) metric to evaluate knowledge consistency across languages independently from accuracy. Using this metric, we conduct an in-depth analysis of the determining factors for CLC, both at model level and at language-pair level. Among other results, we find that increasing model size leads to higher factual probing accuracy in most languages, but does not improve cross-lingual consistency. Finally, we conduct a case study on CLC when new factual associations are inserted in the PLMs via model editing. Results on a small sample of facts inserted in English reveal a clear pattern whereby the new piece of knowledge transfers only to languages with which English has a high RankC score.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Measuring Cross-Lingual Consistency
  • 3.1 Prior Work: Correct Predictions Overlap
  • 3.2 This Work: RankC Metric
  • 4 Experimental Setup
  • 5 Main Consistency Results
  • 5.1 Consistency in Different PLMs
  • 5.2 Effect of Model Size
  • 6 Factors of Cross-Lingual Consistency
  • 6.1 Typological Similarity
  • 6.2 Subword Vocabulary Overlap
  • 7 Case Study: Cross-Lingual Consistency and Knowledge Incorporation
  • 8 Conclusion
  • Acknowledgements
  • Limitations
  • Ethics Statement
  • References
  • A GenBench Evaluation Card
  • B Prompt-based Probing
  • C Ranking-Based Weight
  • D RankC Computation Example
  • E Range of RankC Values
  • F Comparison of RankC and COverlap
  • G Effect of Non-Supported Languages in BLOOM
  • H Additional Results on BMLAMA-53
  • I Knowledge Consistency on More Models
  • J Additional Cases of Counterfactual Knowledge Incorporation
  • K Raw Logits Before/After Model Editing

Knowls

  1. Knowl 1 — Ranking-based Consistency measures agreement across complete candidate rankings

    definition

    Ranking-based Consistency (RankC) measures whether a multilingual language model ranks corresponding candidate answers similarly for translated versions of the same query; it does not require that the top-ranked answer be correct. Assume each query qiq_i and its translated query qi′q_i' have the same NiN_i candidates, aligned by answer identity. Write their candidates in descending model-score order as ci1,…,ciNic_i^1,\ldots,c_i^{N_i} and ci′1,…,ci′Nic_i'^1,\ldots,c_i'^{N_i}. For each rank cutoff jj, define P@jP@j as the fraction of the two top-jj sets that overlap, and weight smaller cutoffs more heavily: P@j=∣{ci1,…,cij}∩{ci′1,…,ci′j}∣jP@j=\frac{|\{c_i^1,\ldots,c_i^j\}\cap\{c_i'^1,\ldots,c_i'^j\}|}{j} and wj=eNi−j∑k=1NieNi−kw_j=\frac{e^{N_i-j}}{\sum_{k=1}^{N_i}e^{N_i-k}}. The query-pair consistency is consist⁡(qi,qi′)=∑j=1NiwjP@j\operatorname{consist}(q_i,q_i')=\sum_{j=1}^{N_i}w_jP@j; for languages l,l′l,l', RankC is the mean of this score over their ∣Ql∣|Q_l| aligned translated query pairs: RankC⁡(l,l′)=1∣Ql∣∑i=1∣Ql∣consist⁡(qi,qi′)\operatorname{RankC}(l,l')=\frac{1}{|Q_l|}\sum_{i=1}^{|Q_l|}\operatorname{consist}(q_i,q_i'). RankC ranges from 0 to 1 and can treat matching incorrect answers as consistent. The balanced, parallel query-and-candidate requirement makes language-pair scores comparable; unlike a metric based only on overlapping correct predictions, RankC is designed to separate consistency from factual accuracy.

  2. Knowl 2 — BMLAMA provides parallel factual-probing benchmarks at two language scales

    data/table

    The paper constructs Balanced Multilingual LAnguage Model Analysis (BMLAMA) by retaining queries and candidates from X-FACTR and MLAMA that are available in every language in a benchmark version. This parallel structure is required for comparing translated query pairs with a common candidate set. BMLAMA-17 contains 41 relations and 6,792 queries in each of 17 languages (6,792 × 17 query-language instances), with an average of 9.71 candidates per query. BMLAMA-53 contains 30 relations and 3,070 queries in each of 53 languages (3,070 × 53 instances), with an average of 9.56 candidates per query. The 17-language version retains more queries per language, while the 53-language version broadens language coverage.

  3. Knowl 3 — Candidate scoring accommodates encoder-only, encoder-decoder, and decoder-only models

    model/method

    For factual probing, each prompt supplies a subject and a relation-specific answer slot, and the model scores each candidate answer; candidates are ranked by those scores. The scoring procedure accounts for each architecture and for candidates that tokenize into multiple subwords. For an encoder-only model, the candidate's score is the mean log probability of its subwords at their corresponding masked positions, conditioning on the prompt and earlier filled masks. For an encoder-decoder model, the score is the mean autoregressive log probability of the candidate's subword sequence given the prompt. For a decoder-only model, the candidate is inserted into the prompt's answer slot, and its score is the mean autoregressive log probability of the resulting subword sequence. These architecture-specific scores produce the candidate rankings used for factual accuracy and RankC.

  4. Knowl 4 — Accuracy rises with BLOOM scale while measured consistency changes little

    empirical result

    On BMLAMA-17, average RankC and factual probing accuracy (%) were respectively 23.1 and 21.1 for BLOOM-560m, 23.9 and 22.9 for BLOOM-1.1b, 24.8 and 24.8 for BLOOM-1.7b, and 25.2 and 26.0 for BLOOM-3b. Thus, across this fivefold increase in parameters, accuracy improved consistently while average RankC rose only about two percentage points. The broader model comparison reported 32.0 RankC and 32.5 accuracy for XLM-RoBERTa-large, and 32.3 RankC and 32.4 accuracy for mT5-large, compared with 25.2 and 26.0 for BLOOM-3b. The authors also report 26.3 RankC and 27.8 accuracy for BLOOM-7.1b, and 28.5 for both metrics for LLaMA-7b. Average consistency excludes self-pairs, which trivially score 100%. These results show that greater size can improve factual retrieval without a comparable increase in cross-language agreement; the paper cautions that the size trend should not automatically be generalized beyond the tested models.

  5. Knowl 5 — Subword-vocabulary overlap correlates more strongly with consistency than typological similarity

    empirical result

    The study computes Pearson correlations across language pairs between RankC and language similarity measures, including overlap of model-tokenized text. For BMLAMA-17, the correlations of RankC with subword overlap measured on BMLAMA and Flores-200 were, respectively, 0.70 (p = 2e-21) and 0.49 (p = 2e-09) for BLOOM-3b; 0.71 (2e-22) and 0.68 (5e-20) for mT5-large; and 0.80 (8e-32) and 0.79 (1e-30) for XLM-RoBERTa-large. For BMLAMA-53, the corresponding correlations were 0.69 (3e-195) and 0.61 (2e-140), 0.74 (8e-243) and 0.65 (2e-168), and 0.70 (5e-205) and 0.63 (7e-156). By comparison, correlations with genetic similarity were 0.42–0.52 on BMLAMA-17 and 0.28–0.40 on BMLAMA-53; correlations with geographical similarity were 0.34–0.43 and 0.15–0.23, respectively. Syntactic similarity had no significant correlation on BMLAMA-17 and only weak positive correlations on BMLAMA-53; phonological similarity had no meaningful association. The strong positive association with vocabulary overlap also appears when overlap is measured on the separate Flores-200 corpus, although the correlations are lower than those based on BMLAMA. These are correlations, not evidence that overlap causes consistency. The authors interpret the pattern as consistent with knowledge transferring partly through shared subword representations.

  6. Knowl 6 — Consistency varies substantially across language pairs

    empirical result

    The pairwise RankC analyses of BMLAMA-17 show that average cross-language consistency is low, while particular language pairs agree more strongly. English, French, Dutch, Spanish, and Catalan share considerable consistency in mT5-large and XLM-RoBERTa-large; BLOOM-3b shows a similar pattern except for Dutch, which was not included in that model's training corpus. Vietnamese and Turkish also show notable consistency with these European languages across the tested models. Russian and Ukrainian are another high-consistency pair, sharing Cyrillic script and close linguistic relatedness. The study further notes that average RankC on BMLAMA-53 is broadly in line with BMLAMA-17, with a two-percentage-point decrease for XLM-RoBERTa-large. These patterns motivate investigating both linguistic properties and model-specific subword overlap rather than assuming that all language pairs behave alike.

  7. Knowl 7 — English model edits transfer more to languages with high English RankC

    empirical result

    In a case study using BLOOM-3b and Rank-One Model Editing (ROME), the researchers inserted counterfactual facts in English and measured normalized scores for one correct and one edited, incorrect answer. The target languages were Spanish and Vietnamese, which had English RankC scores of 52 and 49, and Hungarian and Greek, with scores of 26 and 24. Six facts were selected so that the original answer was highly probable before editing and the subject and object entities were the same tokens across languages. For the query “Steve Jobs worked for Apple/Microsoft,” the normalized correct/wrong scores before and after the English edit were: English, 0.95/0.05 to 0.19/0.81; Spanish, 0.93/0.07 to 0.12/0.88; Vietnamese, 0.99/0.01 to 0.24/0.76; Hungarian, 0.95/0.05 to 0.81/0.19; and Greek, 0.99/0.01 to 0.91/0.09. The three other displayed main-text facts and three additional facts in the appendix show the same general pattern: high-RankC languages change more in the direction of the English edit, while low-RankC languages are less affected. This small case study provides preliminary evidence that pre-existing CLC predicts cross-language effects of knowledge editing, not a broad statistical demonstration.

  8. Knowl 8 — RankC identifies agreement missed by correct-prediction overlap

    empirical result

    On BMLAMA-17, the paper compares RankC with COverlap, a measure based on overlap of correct top-ranked predictions. Nearly all language pairs classified as highly consistent by COverlap were also identified as highly consistent by RankC. RankC additionally identified seven high-consistency pairs that COverlap underestimated because one language had low probing accuracy; three high-consistency pairs identified by COverlap were not recovered by RankC. This comparison supports the intended distinction: two languages can rank the same answers similarly even when those answers are incorrect, so a consistency measure restricted to correct predictions can miss agreement.

  9. Knowl 9 — Evaluation scope is limited by model scale and benchmark coverage

    limitation

    The experiments could not test models larger than BLOOM-7.1b because of GPU-resource constraints, so the reported size findings do not establish how consistency behaves in still larger models. BMLAMA is inherited from X-FACTR and MLAMA, whose facts are drawn from Wikidata; the paper notes that these facts may be more relevant to the Western world and therefore bias evaluation coverage. It identifies broader-scale testing and more geographically representative factual benchmarks as areas needing further work.

Coverage note — The full pairwise heatmaps, the alternative RankC weighting-scheme comparison, and raw model-editing logits are omitted because they are supporting diagnostics rather than independent load-bearing contributions.

References

  1. 1.Israa Alghanmi, Luis Espinosa Anke, and Steven Schockaert. 2021. Probing pre-trained language models for disease knowledge. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3023–3033, Online. Association for Computational Linguistics.
  2. 2.Mikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer, Ramakanth Pasunuru, Giridharan Anantharaman, Xian Li, Shuohui Chen, Halil Akin, Mandeep Baines, Louis Martin, Xing Zhou, Punit Singh Koura, Brian O’Horo, Jeffrey Wang, Luke Zettlemoyer, Mona Diab, Zornitsa Kozareva, and Veselin Stoyanov. 2022. Efficient large scale language modeling with mixtures of experts. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11699–11732, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  3. 3.Terra Blevins and Luke Zettlemoyer. 2022. Language contamination helps explains the cross-lingual capabilities of English pretrained models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3563–3574, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  4. 4.Zied Bouraoui, Jose Camacho-Collados, and Steven Schockaert. 2020. Inducing relational knowledge from bert. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 7456–7463.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  6. 6.Israel Cohen, Yiteng Huang, Jingdong Chen, Jacob Benesty, Jacob Benesty, Jingdong Chen, Yiteng Huang, and Israel Cohen. 2009. Pearson correlation coefficient. Noise reduction in speech processing, pages 1–4.
  7. 7.Marta R Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al. 2022. No language left behind: Scaling human-centered machine translation. arXiv preprint arXiv:2207.04672.
  8. 8.Joe Davison, Joshua Feldman, and Alexander Rush. 2019. Commonsense knowledge mining from pre-trained models. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1173–1178, Hong Kong, China. Association for Computational Linguistics.
  9. 9.Nicola De Cao, Wilker Aziz, and Ivan Titov. 2021. Editing factual knowledge in language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6491–6506, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  11. 11.Ning Ding, Yujia Qin, Guang Yang, Fuchao Wei, Zonghan Yang, Yusheng Su, Shengding Hu, Yulin Chen, Chi-Min Chan, Weize Chen, et al. 2022. Delta tuning: A comprehensive study of parameter efficient methods for pre-trained language models. arXiv preprint arXiv:2203.06904.
  12. 12.Constanza Fierro and Anders Søgaard. 2022. Factual consistency of multilingual pretrained language models. In Findings of the Association for Computational Linguistics: ACL 2022, pages 3046–3052, Dublin, Ireland. Association for Computational Linguistics.
  13. 13.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3816–3830, Online. Association for Computational Linguistics.
  14. 14.Harald Hammarström, Robert Forkel, and Martin Haspelmath. 2017. Glottolog 3.0. Max Planck Institute for the Science of Human History.
  15. 15.Yifan Hou, Wenxiang Jiao, Meizhen Liu, Carl Allen, Zhaopeng Tu, and Mrinmaya Sachan. 2022. Adapters for enhanced modeling of multilingual knowledge and text. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 3902–3917, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  16. 16.Shengding Hu, Ning Ding, Huadong Wang, Zhiyuan Liu, Jingang Wang, Juanzi Li, Wei Wu, and Maosong Sun. 2022. Knowledgeable prompt-tuning: Incorporating knowledge into prompt verbalizer for text classification. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2225–2240, Dublin, Ireland. Association for Computational Linguistics.
  17. 17.Dieuwke Hupkes, Mario Giulianelli, Verna Dankers, Mikel Artetxe, Yanai Elazar, Tiago Pimentel, Christos Christodoulopoulos, Karim Lasri, Naomi Saphra, Arabella Sinclair, Dennis Ulmer, Florian Schottmann, Khuyagbaatar Batsuren, Kaiser Sun, Koustuv Sinha, Leila Khalatbari, Maria Ryskina, Rita Frieske, Ryan Cotterell, and Zhijing Jin. 2022. State-of-the-art generalisation research in NLP: a taxonomy and review. CoRR.
  18. 18.Zhengbao Jiang, Antonios Anastasopoulos, Jun Araki, Haibo Ding, and Graham Neubig. 2020. X-FACTR: Multilingual factual knowledge retrieval from pre-trained language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5943–5959, Online. Association for Computational Linguistics.
  19. 19.Nora Kassner, Philipp Dufter, and Hinrich Schütze. 2021. Multilingual LAMA: Investigating knowledge in multilingual pretrained language models. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 3250–3258, Online. Association for Computational Linguistics.
  20. 20.Tao Li, Vivek Gupta, Maitrey Mehta, and Vivek Srikumar. 2019. A logic-driven framework for consistency of neural models. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3924–3935, Hong Kong, China. Association for Computational Linguistics.
  21. 21.Patrick Littell, David R. Mortensen, Ke Lin, Katherine Kairis, Carlisle Turner, and Lori Levin. 2017. URIEL and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 8–14, Valencia, Spain. Association for Computational Linguistics.
  22. 22.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  23. 23.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in gpt. Advances in Neural Information Processing Systems, 35:17359–17372.
  24. 24.Eric Mitchell, Joseph Noh, Siyan Li, Will Armstrong, Ananth Agarwal, Patrick Liu, Chelsea Finn, and Christopher Manning. 2022. Enhancing self-consistency and performance of pre-trained language models through natural language inference. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 1754–1768, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  25. 25.Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. 2023. Crosslingual generalization through multitask finetuning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15991–16111, Toronto, Canada. Association for Computational Linguistics.
  26. 26.Xenia Ohmer, Elia Bruni, and Dieuwke Hupkes. 2023. Separating form and meaning: Using self-consistency to quantify task understanding across multiple senses. CoRR.
  27. 27.Hao Peng, Xiaozhi Wang, Shengding Hu, Hailong Jin, Lei Hou, Juanzi Li, Zhiyuan Liu, and Qun Liu. 2022. COPEN: Probing conceptual knowledge in pre-trained language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 5015–5035, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  28. 28.Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2463–2473, Hong Kong, China. Association for Computational Linguistics.
  29. 29.Edoardo Maria Ponti, Helen O’Horan, Yevgeni Berzak, Ivan Vulic, Roi Reichart, Thierry Poibeau, Ekaterina Shutova, and Anna Korhonen. 2019. Modeling Language Variation and Universals: A Survey on Typological Linguistics for Natural Language Processing. Computational Linguistics, 45(3):559–601.
  30. 30.Yujia Qin, Yankai Lin, Jing Yi, Jiajie Zhang, Xu Han, Zhengyan Zhang, Yusheng Su, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou. 2022. Knowledge inheritance for pre-trained language models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3921–3937, Seattle, United States. Association for Computational Linguistics.
  31. 31.Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5418–5426, Online. Association for Computational Linguistics.
  32. 32.Gabriele Sarti, Nils Feldhus, Ludwig Sickert, and Oskar van der Wal. 2023. Inseq: An interpretability toolkit for sequence generation models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), pages 421–435, Toronto, Canada. Association for Computational Linguistics.
  33. 33.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  34. 34.Hinrich Schutze, Christopher D Manning, and Prabhakar Raghavan. 2008. Introduction to information retrieval. Cambridge University Press.
  35. 35.Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4222–4235, Online. Association for Computational Linguistics.
  36. 36.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  37. 37.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations.
  38. 38.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online. Association for Computational Linguistics.
  39. 39.Da Yin, Hritik Bansal, Masoud Monajatipoor, Liunian Harold Li, and Kai-Wei Chang. 2022. GeoMLAMA: Geo-diverse commonsense probing on multilingual pre-trained language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 2039–2055, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  40. 40.Ahmet Üstün, Arianna Bisazza, Gosse Bouma, and Gertjan van Noord. 2022. UDapter: Typology-based Language Adapters for Multilingual Dependency Parsing and Sequence Labeling. Computational Linguistics, 48(3):555–592.

Citation

MLA
Qi, J., et al. “Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 10650–66, https://doi.org/10.18653/v1/2023.emnlp-main.658.
APA
Qi, J., Fernández, R., & Bisazza, A. (2023). Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 10650–10666. https://doi.org/10.18653/v1/2023.emnlp-main.658
Chicago
Qi, J., R. Fernández, and A. Bisazza. 2023. “Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 10650–66. https://doi.org/10.18653/v1/2023.emnlp-main.658.
Harvard
Qi, J., Fernández, R. and Bisazza, A. (2023) “Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 10650–10666. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.658.
Vancouver
1. Qi J, Fernández R, Bisazza A (2023) Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 10650–10666

BibTeX

@inproceedings{qi-etal-2023-cross,
    title = "Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models",
    author = "Qi, Jirui  and
      Fern{\'a}ndez, Raquel  and
      Bisazza, Arianna",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.658/",
    doi = "10.18653/v1/2023.emnlp-main.658",
    pages = "10650--10666"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/