Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multilingual language models

Multilingual language models are artificial intelligence models trained on text corpora spanning multiple natural languages, enabling them to understand, process, and generate text across diverse linguistic systems within a single shared framework. Built typically on deep transformer architectures, these models map different languages into unified representation spaces to facilitate cross-lingual transfer, which allows patterns, concepts, and task capabilities learned from data-rich languages to generalize to low-resource languages. They support a broad range of natural language processing applications, including classification, question answering, text generation, and translation, without requiring separate specialized architectures for each individual language.

11 items

Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models

Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models

Jirui Qi, Raquel Fernández, Arianna Bisazza

OrganizationsUniversity of AmsterdamUniversity of Groningen

Why you should read this

Introduces a ranking-based consistency metric and a multi-parallel benchmark to evaluate factual knowledge alignment across languages in multilingual language models, showing that scaling model size fails to improve cross-lingual consistency and that consistency scores predict whether edited facts transfer across languages.

Multilingual large-scale Pretrained Language Models (PLMs) have been shown to store considerable amounts of factual knowledge, but large variations are observed across languages. With the ultimate goal of ensuring that users with different language backgrounds obtain consistent feedback from the same model, we study the cross-lingual consistency (CLC) of factual knowledge in various multilingual PLMs. To this end, we propose a Ranking-based Consistency (RankC) metric to evaluate knowledge consistency across languages independently from accuracy. Using this metric, we conduct an in-depth analysis of the determining factors for CLC, both at model level and at language-pair level. Among other results, we find that increasing model size leads to higher factual probing accuracy in most languages, but does not improve cross-lingual consistency. Finally, we conduct a case study on CLC when new factual associations are inserted in the PLMs via model editing. Results on a small sample of facts inserted in English reveal a clear pattern whereby the new piece of knowledge transfers only to languages with which English has a high RankC score.

Added

2026-10-01

Towards a Common Understanding of Contributing Factors for Cross-Lingual Transfer in Multilingual Language Models: A Review

Towards a Common Understanding of Contributing Factors for Cross-Lingual Transfer in Multilingual Language Models: A Review

Fred Philippy, Siwen Guo, Shohreh Haddadan

OrganizationsUniversity of LuxembourgZortify

Why you should read this

Categorizes and synthesizes empirical findings across five key drivers of cross-lingual transfer in multilingual language models to reconcile conflicting literature and guide more effective zero-shot cross-lingual adaptation.

In recent years, pre-trained Multilingual Language Models (MLLMs) have shown a strong ability to transfer knowledge across different languages. However, given that the aspiration for such an ability has not been explicitly incorporated in the design of the majority of MLLMs, it is challenging to obtain a unique and straightforward explanation for its emergence. In this review paper, we survey literature that investigates different factors contributing to the capacity of MLLMs to perform zero-shot cross-lingual transfer and subsequently outline and discuss these factors in detail. To enhance the structure of this review and to facilitate consolidation with future studies, we identify five categories of such factors. In addition to providing a summary of empirical evidence from past studies, we identify consensuses among studies with consistent findings and resolve conflicts among contradictory ones. Our work contextualizes and unifies existing research streams which aim at explaining the cross-lingual potential of MLLMs. This review provides, first, an aligned reference point for future research and, second, guidance for a better-informed and more efficient way of leveraging the cross-lingual capacity of MLLMs.

Added

2026-09-26

NoisyTune: A Little Noise Can Help You Finetune Pretrained Language Models Better

NoisyTune: A Little Noise Can Help You Finetune Pretrained Language Models Better

Chuhan Wu, Fangzhao Wu, Tao Qi, Yongfeng Huang

OrganizationsMicrosoftTsinghua University

Why you should read this

Proposes NoisyTune, a lightweight and easily implementable method that improves downstream fine-tuning by adding matrix-wise uniform noise scaled by parameter standard deviations to prevent pretrained language models from overfitting their pretraining tasks.

Effectively finetuning pretrained language models (PLMs) is critical for their success in downstream tasks. However, PLMs may have risks in overfitting the pretraining tasks and data, which usually have gap with the target downstream tasks. Such gap may be difficult for existing PLM finetuning methods to overcome and lead to suboptimal performance. In this paper, we propose a very simple yet effective method named NoisyTune to help better finetune PLMs on downstream tasks by adding some noise to the parameters of PLMs before finetuning. More specifically, we propose a matrix-wise perturbing method which adds different uniform noises to different parameter matrices based on their standard deviations. In this way, the varied characteristics of different types of parameters in PLMs can be considered. Extensive experiments on both GLUE English benchmark and XTREME multilingual benchmark show NoisyTune can consistently empower the finetuning of different PLMs on different downstream tasks.

Added

2026-09-26

Expanding Pretrained Models to Thousands More Languages via Lexicon-based Adaptation

Expanding Pretrained Models to Thousands More Languages via Lexicon-based Adaptation

Xinyi Wang, Sebastian Ruder, Graham Neubig

OrganizationsCarnegie Mellon UniversityGoogle

Why you should read this

Proposes a data augmentation framework using widely available bilingual lexicons to adapt multilingual pretrained models to under-represented languages with little to no existing text corpora.

The performance of multilingual pretrained models is highly dependent on the availability of monolingual or parallel text present in a target language. Thus, the majority of the world’s languages cannot benefit from recent progress in NLP as they have no or limited textual data. To expand possibilities of using NLP technology in these under-represented languages, we systematically study strategies that relax the reliance on conventional language resources through the use of bilingual lexicons, an alternative resource with much better language coverage. We analyze different strategies to synthesize textual or labeled data using lexicons, and how this data can be combined with monolingual or parallel text when available. For 19 under-represented languages across 3 tasks, our methods lead to consistent improvements of up to 5 and 15 points with and without extra monolingual text respectively. Overall, our study highlights how NLP methods can be adapted to thousands more languages that are under-served by current technology.1

Added

2026-09-26

Do Llamas Work in English? On the Latent Language of Multilingual Transformers

Do Llamas Work in English? On the Latent Language of Multilingual Transformers

Chris Wendler, Veniamin Veselovsky, Giovanni Monea, Robert West

OrganizationsÉcole Polytechnique Fédérale de Lausanne

Why you should read this

Reveals through logit lens analysis that multilingual transformer models route non-English inputs through an internal, English-aligned concept space in intermediate layers before decoding them into the target language.

We ask whether multilingual language models trained on unbalanced, English-dominated corpora use English as an internal pivot language—a question of key importance for understanding how language models function and the origins of linguistic bias. Focusing on the Llama-2 family of transformer models, our study uses carefully constructed non-English prompts with a unique correct single-token continuation. From layer to layer, transformers gradually map an input embedding of the final prompt token to an output embedding from which next-token probabilities are computed. Tracking intermediate embeddings through their high-dimensional space reveals three distinct phases, whereby intermediate embeddings (1) start far away from output token embeddings; (2) already allow for decoding a semantically correct next token in middle layers, but give higher probability to its version in English than in the input language; (3) finally move into an input-language-specific region of the embedding space. We cast these results into a conceptual model where the three phases operate in “input space”, “concept space”, and “output space”, respectively. Crucially, our evidence suggests that the abstract “concept space” lies closer to English than to other languages, which may have important consequences regarding the biases held by multilingual language models. Code and data is made available here: https://github.com/epfl-dlab/llm-latent-language.

Added

2026-09-26

Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model

Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model

Ahmet Üstün, Viraat Aryabumi, Zheng Xin Yong, Wei-Yin Ko, Daniel D'souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, Freddie Vargus, Phil Blunsom, Shayne Longpre, Niklas Muennighoff, Marzieh Fadaee, Julia Kreutzer, Sara Hooker

OrganizationsBrown UniversityCarnegie Mellon UniversityCohereMassachusetts Institute of Technology

Why you should read this

Presents Aya, an open-access instruction-finetuned language model spanning 101 languages—over half of which are lower-resourced—that consistently outperforms mT0 and BLOOMZ across both discriminative and generative benchmarks.

Recent breakthroughs in large language models (LLMs) have centered around a handful of data-rich languages. What does it take to broaden access to breakthroughs beyond first-class citizen languages? Our work introduces Aya, a massively multilingual generative language model that follows instructions in 101 languages of which over 50% are considered lower-resourced. Aya outperforms mT0 and BLOOMZ on the majority of tasks while covering double the number of languages. We introduce extensive new evaluation suites that broaden the state-of-art for multilingual eval across 99 languages — including discriminative and generative tasks, human evaluation, and simulated win rates that cover both held-out tasks and in-distribution performance. Furthermore, we conduct detailed investigations on the optimal finetuning mixture composition, data pruning, as well as the toxicity, bias, and safety of our models.

Added

2026-09-26

Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages

Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages

Ayyoob Imani, Peiqin Lin, Amir Hossein Kargaran, Silvia Severini, Masoud Jalili Sabet, Nora Kassner, Chunlan Ma, Helmut Schmid, André F. T. Martins, François Yvon, Hinrich Schütze

OrganizationsCNRSInstituto de TelecomunicaçõesInstituto Superior TécnicoISIRLudwig Maximilian University of MunichMunich Center for Machine LearningSorbonne UniversitéUnbabel

Why you should read this

Presents an open-source multilingual corpus and language model spanning over 500 predominantly low-resource languages, outperforming XLM-R across multiple natural language understanding benchmarks while identifying the key drivers of multilingual representation quality.

The NLP community has mainly focused on scaling Large Language Models (LLMs) vertically, i.e., making them better for about 100 languages. We instead scale LLMs horizontally: we create, through continued pretraining, Glot500-m, an LLM that covers 511 predominantly low-resource languages. An important part of this effort is to collect and clean Glot500-c, a corpus that covers these 511 languages and allows us to train Glot500-m. We evaluate Glot500-m on five diverse tasks across these languages. We observe large improvements for both high-resource and low-resource languages compared to an XLM-R baseline. Our analysis shows that no single factor explains the quality of multilingual LLM representations. Rather, a combination of factors determines quality including corpus size, script, “help” from related languages and the total capacity of the model. Our work addresses an important goal of NLP research: we should not limit NLP to a small fraction of the world’s languages and instead strive to support as many languages as possible to bring the benefits of NLP technology to all languages and cultures. Code, data and models are available at https://github.com/cisnlp/Glot500.

Added

2026-09-26

Unsupervised Cross-lingual Representation Learning at Scale

Unsupervised Cross-lingual Representation Learning at Scale

Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco (Paco) Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, Veselin Stoyanov

OrganizationsMeta

Why you should read this

Demonstrates that scaling multilingual masked language modeling across one hundred languages with XLM-R significantly outperforms multilingual BERT on cross-lingual transfer while matching monolingual baselines on standard benchmarks.

This paper shows that pretraining multilingual language models at scale leads to significant performance gains for a wide range of cross-lingual transfer tasks. We train a Transformer-based masked language model on one hundred languages, using more than two terabytes of filtered CommonCrawl data. Our model, dubbed XLM-R, significantly outperforms multilingual BERT (mBERT) on a variety of cross-lingual benchmarks, including +14.6% average accuracy on XNLI, +13% average F1 score on MLQA, and +2.4% F1 score on NER. XLM-R performs particularly well on low-resource languages, improving 15.7% in XNLI accuracy for Swahili and 11.4% for Urdu over previous XLM models. We also present a detailed empirical analysis of the key factors that are required to achieve these gains, including the trade-offs between (1) positive transfer and capacity dilution and (2) the performance of high and low resource languages at scale. Finally, we show, for the first time, the possibility of multilingual modeling without sacrificing per-language performance; XLM-R is very competitive with strong monolingual models on the GLUE and XNLI benchmarks. We will make our code, data and models publicly available.

Added

2026-09-07

Creative Commons License