Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multilingual transformers

Multilingual transformers are neural network models based on the transformer architecture that are trained on text from multiple natural languages to process, understand, and generate multilingual content within a single system. Utilizing self-attention mechanisms, these architectures map tokens from diverse languages into shared or aligned representation spaces, facilitating cross-lingual transfer so that knowledge gained from high-resource languages can benefit low-resource language tasks. They support various natural language processing applications, including machine translation, cross-lingual text classification, and multilingual question answering, without requiring separate models for each language. To manage the linguistic diversity across different scripts, vocabularies, and grammars, these models share underlying parameters across languages and can incorporate modular or language-specific components to mitigate interference between competing language representations.

2 items

Lifting the Curse of Multilinguality by Pre-training Modular Transformers

Lifting the Curse of Multilinguality by Pre-training Modular Transformers

Jonas Pfeiffer, Naman Goyal, Xi Victoria Lin, Xian Li, James Cross, Sebastian Riedel, Mikel Artetxe

OrganizationsMetaNew York UniversityTechnische Universität Darmstadt

Why you should read this

Proposes pre-training multilingual transformers with dedicated language-specific modules from the start, preventing capacity dilution across languages and enabling post-hoc extension to new languages without performance degradation.

Multilingual pre-trained models are known to suffer from the curse of multilinguality, which causes per-language performance to drop as they cover more languages. We address this issue by introducing language-specific modules, which allows us to grow the total capacity of the model, while keeping the total number of trainable parameters per language constant. In contrast with prior work that learns language-specific components post-hoc, we pre-train the modules of our Cross-lingual Modular (X-Mod) models from the start. Our experiments on natural language inference, named entity recognition and question answering show that our approach not only mitigates the negative interference between languages, but also enables positive transfer, resulting in improved monolingual and cross-lingual performance. Furthermore, our approach enables adding languages post-hoc with no measurable drop in performance, no longer limiting the model usage to the set of pre-trained languages.

Added

2026-10-05

Do Llamas Work in English? On the Latent Language of Multilingual Transformers

Do Llamas Work in English? On the Latent Language of Multilingual Transformers

Chris Wendler, Veniamin Veselovsky, Giovanni Monea, Robert West

OrganizationsÉcole Polytechnique Fédérale de Lausanne

Why you should read this

Reveals through logit lens analysis that multilingual transformer models route non-English inputs through an internal, English-aligned concept space in intermediate layers before decoding them into the target language.

We ask whether multilingual language models trained on unbalanced, English-dominated corpora use English as an internal pivot language—a question of key importance for understanding how language models function and the origins of linguistic bias. Focusing on the Llama-2 family of transformer models, our study uses carefully constructed non-English prompts with a unique correct single-token continuation. From layer to layer, transformers gradually map an input embedding of the final prompt token to an output embedding from which next-token probabilities are computed. Tracking intermediate embeddings through their high-dimensional space reveals three distinct phases, whereby intermediate embeddings (1) start far away from output token embeddings; (2) already allow for decoding a semantically correct next token in middle layers, but give higher probability to its version in English than in the input language; (3) finally move into an input-language-specific region of the embedding space. We cast these results into a conceptual model where the three phases operate in “input space”, “concept space”, and “output space”, respectively. Crucially, our evidence suggests that the abstract “concept space” lies closer to English than to other languages, which may have important consequences regarding the biases held by multilingual language models. Code and data is made available here: https://github.com/epfl-dlab/llm-latent-language.

Added

2026-09-26