Do Llamas Work in English? On the Latent Language of Multilingual Transformers

Chris WendlerVeniamin VeselovskyGiovanni MoneaRobert West

article2024ACL350 citations

Reveals through logit lens analysis that multilingual transformer models route non-English inputs through an internal, English-aligned concept space in intermediate layers before decoding them into the target language.

Listen

Modern multilingual artificial intelligence models achieve strong performance across numerous global languages despite being trained predominantly on English text. This disparity raises a critical question regarding whether these systems internally translate inputs into English to process information before generating output in the requested language. Understanding whether models rely on an internal English bridge is essential for identifying subtle linguistic biases, cultural skew, and performance limitations when deploying artificial intelligence globally.

The main objective of the article is to evaluate empirically whether multilingual transformer models use English as an internal pivot language during text generation and to demonstrate the geometrical structure of internal representations across processing layers.

To evaluate this question, the researchers applied mechanistic interpretability methods across multiple sizes of the open-weight Llama-2 family (7B, 13B, and 70B parameters) and validated key findings on Mistral-7B. They designed controlled text-completion benchmarks covering Chinese, German, French, Russian, and Estonian across translation, word repetition, and fill-in-the-blank cloze tasks. The analysis used the logit lens technique—which prematurely projects hidden numerical states in intermediate layers into human-readable word probabilities—alongside geometric analysis measuring the mathematical alignment and energy between hidden vectors and output token embeddings.

The investigation produced three main findings. First, internal processing moves consistently through three distinct operational phases across model scales: an initial input-processing phase in early layers where representations show high entropy and no language emerges; a middle concept phase where entropy drops sharply and semantically correct English words dominate intermediate predictions; and a final target-language phase in the last layers where output probability abruptly shifts to the requested language. Second, geometric analysis revealed that intermediate latent states do not represent literal English text; instead, they operate in an abstract concept space that carries non-linguistic context but sits geometrically closer to English due to training data dominance. Third, tokenization heavily impacts this trajectory: languages with dedicated single-token vocabulary entries (such as Chinese) bypass or reduce the English detour during simple repetition tasks, whereas languages split into multi-token fragments (such as Estonian) are forced deeper through English-aligned concept spaces.

These findings indicate that multilingual models do not perform explicit step-by-step translation, but rather navigate a semantic concept space that is structurally biased toward English. For organizations deploying multilingual systems, this introduces risks of Anglocentric bias, subtle shifts in emotional framing, and distorted reasoning on non-Western cultural topics. Furthermore, underrepresented languages face increased computational inefficiencies and error rates due to vocabulary fragmentation.

To mitigate these risks, developers and policymakers should prioritize rebalancing pretraining corpora, expanding multilingual vocabularies, and adopting culturally balanced tokenizers. Before deploying systems in sensitive multilingual environments, organizations should conduct targeted audits for Anglocentric cognitive drift. Further research should extend these evaluations to complex, multi-token reasoning tasks, explore training interventions on balanced datasets, and examine closed-source commercial models where internal weights cannot currently be accessed.

arXiv: 2402.10588
Cover for Do Llamas Work in English? On the Latent Language of Multilingual Transformers

Abstract

We ask whether multilingual language models trained on unbalanced, English-dominated corpora use English as an internal pivot language—a question of key importance for understanding how language models function and the origins of linguistic bias. Focusing on the Llama-2 family of transformer models, our study uses carefully constructed non-English prompts with a unique correct single-token continuation. From layer to layer, transformers gradually map an input embedding of the final prompt token to an output embedding from which next-token probabilities are computed. Tracking intermediate embeddings through their high-dimensional space reveals three distinct phases, whereby intermediate embeddings (1) start far away from output token embeddings; (2) already allow for decoding a semantically correct next token in middle layers, but give higher probability to its version in English than in the input language; (3) finally move into an input-language-specific region of the embedding space. We cast these results into a conceptual model where the three phases operate in “input space”, “concept space”, and “output space”, respectively. Crucially, our evidence suggests that the abstract “concept space” lies closer to English than to other languages, which may have important consequences regarding the biases held by multilingual language models. Code and data is made available here: https://github.com/epfl-dlab/llm-latent-language.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Materials and methods
  • 3.1 Language models: Llama-2
  • 3.2 Interpreting latent embeddings: Logit lens
  • 3.3 Data: Tasks for eliciting latent language
  • 3.4 Measuring latent language probabilities
  • 4 Results
  • 4.1 Probabilistic view: Logit lens
  • 4.2 Geometric view: An 8192D space Odyssey
  • 5 Conceptual model
  • 6 Discussion
  • Limitations
  • Acknowledgements
  • References
  • A Additional methodological details
  • A.1 Word translation
  • A.2 Computing language probabilities
  • B Additional results
  • B.1 Low-resource language Estonian
  • B.2 Other models: Mistral

Knowls

  1. Knowl 1 — Three-Phase Forward Pass in Multilingual Autoregressive Transformers

    empirical result

    When processing non-English prompts with unambiguous single-token target continuations, the forward pass of autoregressive transformers (specifically evaluated on Llama-2-7B, 13B, 70B, and Mistral-7B) traverses three distinct operational phases across its layers:

    1. Phase 1 (Input/Feature Formation Space; layers 1--40 in Llama-2-70B): The next-token predictive entropy decoded via the logit lens remains near maximum (hicksim14 hicksim 14 bits for a vocabulary of v=32,000v = 32{,}000, close to the uniform distribution entropy of hicksim15 hicksim 15 bits). The squared token energy remains low (hicksim20% hicksim 20\%), and neither the target language token nor its English analog receives meaningful probability mass. Intermediate latents focus on low-level token integration and feature representation.

    2. Phase 2 (Concept Space; layers 41--70 in Llama-2-70B): Predictive entropy drops sharply to 1--2 bits while token energy remains low (hicksim20% hicksim 20\%). The logit-lens-decoded probability mass shifts overwhelmingly onto the English equivalent of the correct target word (P(lang=EN)>0.6P(\text{lang}=\text{EN}) > 0.6--0.80.8), while the target language probability remains low. The model operates in an abstract concept space where semantic representations are geometrically closer to English output embeddings.

    3. Phase 3 (Output/Token Space; layers 71--80 in Llama-2-70B): Squared token energy spikes from hicksim20% hicksim 20\% up to hicksim30%−−35% hicksim 30\%--35\%, and probability mass rapidly transitions from the English token to the correct token in the target language (P(lang=target)>0.8P(\text{lang}=\text{target}) > 0.8). The model projects abstract concept representations onto language-specific output tokens.

  2. Knowl 2 — Squared Token Energy Metric for Latent Representations

    equation

    To quantify the degree to which an intermediate transformer latent vector h∈Rdh \in \mathbb{R}^d aligns with the output token embedding subspace rather than orthogonal feature directions, the squared token energy E(h)2E(h)^2 is defined as the mean squared cosine similarity between hh and all output token vectors (rows of the unembedding matrix U∈Rv×dU \in \mathbb{R}^{v \times d}), normalized by the mean squared cosine similarity among the token embeddings themselves:

    E(h)2=1v∥U^h^∥221v2∥U^U^⊤∥F2=vd∥U^h^∥22∥U^U^⊤∥F2E(h)^2 = \frac{\frac{1}{v} \|\hat{U}\hat{h}\|_2^2}{\frac{1}{v^2} \|\hat{U}\hat{U}^\top\|_F^2} = \frac{v}{d} \frac{\|\hat{U}\hat{h}\|_2^2}{\|\hat{U}\hat{U}^\top\|_F^2}

    where:

    • h∈Rdh \in \mathbb{R}^d is the latent hidden state at layer jj, residing on a hypersphere of radius d\sqrt{d} due to root-mean-square (RMS) normalization.
    • h^=h∥h∥2∈Rd\hat{h} = \frac{h}{\|h\|_2} \in \mathbb{R}^d is the unit-normalized latent vector.
    • U∈Rv×dU \in \mathbb{R}^{v \times d} is the unembedding weight matrix mapping latent dimension dd to vocabulary size vv.
    • U^∈Rv×d\hat{U} \in \mathbb{R}^{v \times d} is the unembedding matrix with each row normalized to unit L2L_2 norm.
    • ∥⋅∥2\|\cdot\|_2 denotes the Euclidean vector norm and ∥⋅∥F\|\cdot\|_F denotes the Frobenius matrix norm.

    In computational practice, the denominator ∥U^U^⊤∥F2\|\hat{U}\hat{U}^\top\|_F^2 is computed equivalently as ∥U^⊤U^∥F2\|\hat{U}^\top\hat{U}\|_F^2 for computational efficiency in dimension d≪vd \ll v.

  3. Knowl 3 — Latent Language Probability Estimation via Premature Logit Lens

    model/method

    To track intermediate language activations without training auxiliary parameters (such as a tuned lens, which would optimize away intermediate multilingual discrepancies), the language modeling unembedding matrix U∈Rv×dU \in \mathbb{R}^{v \times d} is applied prematurely to the latent embedding hn(j)h_n^{(j)} at the final prompt token position nn at intermediate layer jj:

    zn(j)=Uhn(j)∈Rvz_n^{(j)} = U h_n^{(j)} \in \mathbb{R}^v

    Next-token probabilities P(xn+1=t∣hn(j))∝exp⁡(zn,t(j))P(x_{n+1} = t \mid h_n^{(j)}) \propto \exp(z_{n,t}^{(j)}) are computed via the softmax operation over vocabulary VV. For a target language ℓ\ell and a concept whose canonical realization is word ww, the aggregate probability allocated to language ℓ\ell at layer jj is computed by summing the probabilities over all valid starting tokens of ww in language ℓ\ell:

    P(lang=ℓ∣hn(j)):=∑tℓ∈Start(w)P(xn+1=tℓ∣hn(j))P(\text{lang} = \ell \mid h_n^{(j)}) := \sum_{t_\ell \in \text{Start}(w)} P(x_{n+1} = t_\ell \mid h_n^{(j)})

    where Start(w)\text{Start}(w) denotes the set of vocabulary tokens that form prefixes of ww, including bare token prefixes, leading-space token prefixes (e.g., _flower), and individual UTF-8 byte tokens (e.g., <0xE8> for Chinese characters or Cyrillic).

  4. Knowl 4 — Conceptual Model of Concept Space and Semantic English Pivoting

    theoretical result

    The internal forward pass of English-dominated multilingual language models operates via an abstract concept space rather than an explicit lexical English translation pipeline:

    1. Hypersphere Geometry: Latent states and output token embeddings cohabit a dd-dimensional space on a hypersphere. Token embeddings occupy a subspace, leaving remaining dimensions as orthogonal degrees of freedom to store context, syntactic structure, and language targets.
    2. Concept Alignment: In intermediate layers (Phase 2), the model rotates latents into an abstract concept space. Because pretraining data is overwhelmingly English (e.g., 89.70%89.70\% in Llama-2), the output token embeddings of English words lie systematically closer to these abstract concept directions than non-English equivalents.
    3. Semantic vs. Lexical Pivoting: The model does not explicitly translate the prompt to English text and restart execution; instead, its internal lingua franca consists of abstract concept embeddings that are inherently biased toward English token representations in Euclidean angle. In final layers (Phase 3), the latent is rotated into the target language's hemisphere along the output subspace, triggering a sharp increase in token energy.
  5. Knowl 5 — Task-Dependent Intermediate Language Trajectories

    empirical result

    Across multilingual probing tasks, intermediate layer decoding exhibits consistent dynamics with task-specific nuances:

    • Translation Task: In translation prompts (e.g., French-to-Chinese, German-to-Chinese), neither input nor target language tokens receive probability mass in layers 1--40 (70B model). At layer ∼45\thicksim 45, the English translation rises sharply to peak above 0.700.70, before declining in layers 70--80 as the target language token rapidly peaks to >0.80>0.80.
    • Cloze Task: In masked fill-in-the-blank sentence completion prompts, the identical English-first pattern emerges across German, French, Russian, and Chinese: English probability surges in middle layers before target language resolution in the final layers.
    • Repetition (Copying) Task: While French, German, and Russian copy prompts still exhibit an intermediate English probability peak, the Chinese repetition task exhibits Chinese probability rising simultaneously with, or earlier than, English. This occurs because Chinese words in the evaluation were 100%100\% single-token words in the tokenizer vocabulary, whereas Russian (13%13\% single-token), German (43%43\%), and French (55%55\%) suffered from heavy subword splitting that forced representations through the English-biased concept space.
  6. Knowl 6 — Multilingual Probing Task and Prompt Design

    experimental setup

    To evaluate internal language representation without language ambiguity, three few-shot prompt templates were constructed with unambiguous single-token target words:

    1. Translation Task (4 demonstrations + 1 query): Prompts the model to translate a word from a source language (e.g., French, German, Russian) into a target language (e.g., Chinese, German, French, Russian, English), structured as: Français: "fleur" - 中文: "
    2. Repetition / Copy Task (4 demonstrations + 1 query): Prompts the model to repeat the input word in the same language: 中文: "花" - 中文: "
    3. Cloze Task (2 demonstrations + 1 query): Prompts the model to complete a masked sentence describing a target concept, generated via GPT-4 and translated into the target languages: A "___" is often given as a gift and can be found in gardens. Answer: "

    Words were selected by scanning Llama-2's vocabulary for single-token Chinese nouns paired with single-token English translations, then translating to German, French, and Russian via DeepL. Cognates or translations sharing prefix tokens across English and the target language (e.g., English photograph vs. French photographier) were discarded.

  7. Knowl 7 — Dataset Statistics for Multilingual Probing Tasks

    data/table

    The datasets used to probe latent multilingual representations comprise controlled subsets of nouns across German (DE), English (EN), French (FR), Russian (RU), and Chinese (ZH), filtered to remove cross-lingual token prefix overlaps.

    Language Translation (Total/Single) Repetition (Total/Single) Cloze (Total/Single)
    German (DE) 287 / 126 104 / 45 104 / 45
    English (EN) – 132 / 132 132 / 132
    French (FR) 162 / 88 56 / 31 56 / 31
    Russian (RU) 324 / 45 115 / 15 115 / 15
    Chinese (ZH) 353 / 353 139 / 139 139 / 139

    For individual pairwise translation sets, sample counts (and single-token subsets in parentheses) are:

    • DE →\to EN: 120 (120); DE →\to FR: 56 (31); DE →\to RU: 105 (15); DE →\to ZH: 120 (120)
    • EN →\to DE: 104 (45); EN →\to FR: 57 (31); EN →\to RU: 114 (15); EN →\to ZH: 132 (132)
    • FR →\to DE: 93 (40); FR →\to EN: 118 (118); FR →\to RU: 104 (15); FR →\to ZH: 118 (118)
    • RU →\to DE: 90 (41); RU →\to EN: 114 (114); RU →\to FR: 49 (26); RU →\to ZH: 115 (115)
    • ZH →\to DE: 104 (45); ZH →\to EN: 132 (132); ZH →\to FR: 57 (31); ZH →\to RU: 115 (15)
  8. Knowl 8 — Universality of Concept-Space English Bias Across Scales and Architectures

    empirical result

    The three-phase forward pass trajectory and intermediate English logit-lens peak are consistent across model sizes and architectural families:

    • Model Scale Invariance: Llama-2-7B (32 layers, d=4096d = 4096), Llama-2-13B (40 layers, d=5120d = 5120), and Llama-2-70B (80 layers, d=8192d = 8192) all display the same proportional trajectory: high-entropy input processing in the first 50%50\% of layers, peak English probability in layers ∼50%−−85%\thicksim 50\%--85\%, and target language probability spike with token energy surge in the final 15%15\% of layers.
    • Cross-Model Generalization: Probing Mistral-7B (a non-Llama architecture trained independently on English-dominated corpora) on the copy, translation, and cloze tasks in Chinese reveals the identical phenomenon: intermediate layers decode English before transitioning to Chinese at the final layers, accompanied by a final-layer spike in token energy.
  9. Knowl 9 — Latent Dynamics on Low-Resource Languages: Case of Estonian

    empirical result

    Probing Llama-2-7B on Estonian, a low-resource language where only 1 out of 99 test words is a single token in the model vocabulary, reveals that the intermediate English concept representation persists even when the model fails on downstream generation:

    • Translation Task: Although the final-layer accuracy for translating into Estonian is significantly lower than for high-resource languages, intermediate next-token distributions decoded via the logit lens concentrate probability mass on the correct English tokens before attempting transition to Estonian in the final layers.
    • Cloze Task: Llama-2-7B achieves a 0%0\% final-layer success rate on Estonian cloze prompts. Despite this total output failure, logit lens decoding in intermediate layers still produces non-zero probability for the correct English concept word.
    • Copy Task: Estonian copy behavior mirrors Chinese, with Estonian token probability rising earlier in intermediate layers.
  10. Knowl 10 — Methodological Limitations of Logit-Lens Latent Language Probing

    limitation

    The methodology for analyzing latent language dynamics has several fundamental constraints:

    1. Unembedding Orthogonality Blindness: The logit lens projects latents directly via the output unembedding matrix UU. Any intermediate information stored in orthogonal subspaces (such as attention key/query/value computations, long-term contextual tracking, or non-output reasoning) is invisible to the logit lens and registers as entropy or noise.
    2. Task and Single-Token Simplicity: Experiments are restricted to tightly controlled, prompt-driven single-token noun completions and short cloze tests. They do not evaluate multi-token free generation, complex multi-step reasoning, or culturally sensitive narrative tasks.
    3. High-Dimensional Geometry Unresolved: While the empirical evidence establishes the functional existence of an English-biased concept space, its exact high-dimensional topological structure in Rd\mathbb{R}^d (d=4096d = 4096--81928192) remains unmapped.
    4. White-Box Parameter Requirement: The technique strictly requires direct access to hidden state vectors and the unembedding projection matrix, precluding its application to closed-source API-only language models.

Coverage note — None was omitted; all primary conceptual, mathematical, algorithmic, empirical, and dataset contributions have been captured.

References

  1. 1.Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Nambi, Tanuja Ganu, Sameer Segal, Maxamed Axmed, Kalika Bali, and Sunayana Sitaram. 2023. Mega: Multilingual evaluation of generative ai.
  2. 2.Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020. On the cross-lingual transferability of monolingual representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics.
  3. 3.Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, et al. 2023. A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity. arXiv preprint arXiv:2302.04023.
  4. 4.Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt. 2023. Eliciting latent predictions from transformers with the tuned lens. arXiv preprint arXiv:2303.08112.
  5. 5.Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. 2023. Pythia: A suite for analyzing large language models across training and scaling. In International Conference on Machine Learning, pages 2397–2430. PMLR.
  6. 6.Lera Boroditsky, Lauren A. Schmidt, and Webb Phillips. 2003. Sex, syntax, and semantics. In Dedre Gentner and Susan Goldin-Meadow, editors, Language in Mind: Advances in the Study of Language and Thought, pages 61–79. MIT Press, Cambridge, MA.
  7. 7.Nicola Cancedda. 2024. Spectral filters, dark signals, and attention sinks. arXiv preprint arXiv:2402.09221.
  8. 8.Steven Cao, Nikita Kitaev, and Dan Klein. 2020. Multilingual alignment of contextual word representations.
  9. 9.Yihong Chen, Kelly Marchisio, Roberta Raileanu, David Ifeoluwa Adelani, Pontus Stenetorp, Sebastian Riedel, and Mikel Artetxe. 2023. Improving language plasticity via pretraining with active forgetting.
  10. 10.Rochelle Choenni and Ekaterina Shutova. 2020. What does it mean to be language-agnostic? probing multilingual sentence encoders for typological properties. arXiv preprint arXiv:2009.12862.
  11. 11.Arthur Conmy, Augustine N Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso. 2023. Towards automated circuit discovery for mechanistic interpretability. arXiv preprint arXiv:2304.14997.
  12. 12.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020a. Unsupervised cross-lingual representation learning at scale.
  13. 13.Alexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer, and Veselin Stoyanov. 2020b. Emerging cross-lingual structure in pretrained language models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6022–6034.
  14. 14.Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022. LLM.int8(): 8-bit matrix multiplication for transformers at scale. arXiv preprint arXiv:2208.07339.
  15. 15.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding.
  16. 16.Anna Di Natale, Max Pellert, and David Garcia. 2021. Colexification networks encode affective meaning. Affective Science, 2(2):99–111.
  17. 17.Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al. 2021. A mathematical framework for transformer circuits. Transformer Circuits Thread, 1.
  18. 18.Julen Etxaniz, Gorka Azkune, Aitor Soroa, Oier Lopez de Lacalle, and Mikel Artetxe. 2023. Do multilingual language models think better in english?
  19. 19.Mor Geva, Avi Caciularu, Kevin Ro Wang, and Yoav Goldberg. 2022. Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space.
  20. 20.Charles Goddard. 2023. Llama-polyglot-13b. https://huggingface.co/chargoddard/llama-polyglot-13b. Accessed: 2024-01-22.
  21. 21.Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas. 2023. Finding neurons in a haystack: Case studies with sparse probing. arXiv preprint arXiv:2305.01610.
  22. 22.Bofeng Huang. 2023. vigogne-2-13b-instruct. https://huggingface.co/bofenghuang/vigogne-2-13b-instruct. Accessed: 2024-01-22.
  23. 23.Haoyang Huang, Tianyi Tang, Dongdong Zhang, Wayne Xin Zhao, Ting Song, Yan Xia, and Furu Wei. 2023. Not all languages are created equal in llms: Improving multilingual capability by cross-lingual-thought prompting.
  24. 24.Jaavid Aktar Husain, Raj Dabre, Aswanth Kumar, Ratish Puduppully, and Anoop Kunchukuttan. 2024. Romansetu: Efficiently unlocking multilingual capabilities of large language models models via romanization.
  25. 25.Daekeun Kim. 2023. Llama-2-ko-dpo-13b. https://huggingface.co/daekeun-ml/Llama-2-ko-DPO-13B. Accessed: 2024-01-22.
  26. 26.Kenneth Li, Aspen K Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2022. Emergent world representations: Exploring a sequence model trained on a synthetic task. arXiv preprint arXiv:2210.13382.
  27. 27.Jindřich Libovický, Rudolf Rosa, and Alexander Fraser. 2020. On the language neutrality of pre-trained multilingual representations. arXiv preprint arXiv:2004.05160.
  28. 28.Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, and Xian Li. 2022. Few-shot learning with multilingual generative language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics.
  29. 29.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics, 8:726–742.
  30. 30.Samuel Marks and Max Tegmark. 2023. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets. arXiv preprint arXiv:2310.06824.
  31. 31.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in gpt. Advances in Neural Information Processing Systems, 35:17359–17372.
  32. 32.Giovanni Monea, Maxime Peyrard, Martin Josifoski, Vishrav Chaudhary, Jason Eisner, Emre Kıcıman, Hamid Palangi, Barun Patra, and Robert West. 2023. A glitch in the matrix? locating and detecting language model grounding with fakepedia. arXiv preprint arXiv:2312.02073.
  33. 33.Benjamin Muller, Yanai Elazar, Benoît Sagot, and Djamé Seddah. 2021. First align, then predict: Understanding the cross-lingual ability of multilingual bert. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2214–2231.
  34. 34.Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt. 2023. Progress measures for grokking via mechanistic interpretability. arXiv preprint arXiv:2301.05217.
  35. 35.Nostalgebraist. 2020. Interpreting gpt: The logit lens. LessWrong.
  36. 36.Rafael E. Núñez and Eve Sweetser. 2006. With the future behind them: Convergent evidence from aymara language and gesture in the crosslinguistic comparison of spatial construals of time. Cognitive Science, 30(3):401–450.
  37. 37.OpenAI. 2023. Gpt-4 technical report.
  38. 38.Isabel Papadimitriou, Kezia Lopez, and Dan Jurafsky. 2022. Multilingual bert has an accent: Evaluating english influences on fluency in multilingual models.
  39. 39.Rait Piir. 2023. Finland’s chatgpt equivalent begins to think in estonian as well. ERR News.
  40. 40.Björn Plüster. 2023. LeoLM: Ein Impuls für Deutschsprachige LLM-Forschung. https://laion.ai/blog-de/leo-lm/. Accessed: 2024-01-22.
  41. 41.Lucia Quirke, Lovis Heindrich, Wes Gurnee, and Neel Nanda. 2023. Training dynamics of contextual n-grams in language models. arXiv preprint arXiv:2311.00863.
  42. 42.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  43. 43.Nina Rimsky. 2023. Decoding intermediate activations in Llama-2-7b. LessWrong.
  44. 44.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  45. 45.Tal Schuster, Ori Ram, Regina Barzilay, and Amir Globerson. 2019. Cross-lingual alignment of contextual word embeddings, with applications to zero-shot dependency parsing. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1599–1613, Minneapolis, Minnesota. Association for Computational Linguistics.
  46. 46.Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. 2022. Language models are multilingual chain-of-thought reasoners.
  47. 47.Oleh Shliazhko, Alena Fenogenova, Maria Tikhonova, Vladislav Mikhailov, Anastasia Kozlova, and Tatiana Shavrina. 2022. mgpt: Few-shot learners go multilingual.
  48. 48.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  49. 49.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  50. 50.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  51. 51.Xiangpeng Wei, Haoran Wei, Huan Lin, Tianhao Li, Pei Zhang, Xingzhang Ren, Mei Li, Yu Wan, Zhiwei Cao, Binbin Xie, Tianxiang Hu, Shangjie Li, Binyuan Hui, Bowen Yu, Dayiheng Liu, Baosong Yang, Fei Huang, and Jun Xie. 2023. Polylm: An open source polyglot large language model.
  52. 52.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mt5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics.
  53. 53.Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson. 2014. How transferable are features in deep neural networks? Advances in neural information processing systems, 27.
  54. 54.Jason Yosinski, Jeff Clune, Anh Nguyen, Thomas Fuchs, and Hod Lipson. 2015. Understanding neural networks through deep visualization. arXiv preprint arXiv:1506.06579.
  55. 55.Xiang Zhang, Senyu Li, Bradley Hauer, Ning Shi, and Grzegorz Kondrak. 2023. Don’t trust ChatGPT when your question is not in English: A study of multilingual abilities and types of LLMs. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 7915–7927, Singapore. Association for Computational Linguistics.
  56. 56.Wenhao Zhu, Shujian Huang, Fei Yuan, Shuaijie She, Jiajun Chen, and Alexandra Birch. 2024. Question translation training for better multilingual reasoning.
  57. 57.Wenhao Zhu, Yunzhe Lv, Qingxiu Dong, Fei Yuan, Jingjing Xu, Shujian Huang, Lingpeng Kong, Jiajun Chen, and Lei Li. 2023. Extrapolating large language models to non-english by aligning languages.

Citation

MLA
Wendler, C., et al. “Do Llamas Work in English? On the Latent Language of Multilingual Transformers”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 15366–94, https://doi.org/10.18653/v1/2024.acl-long.820.
APA
Wendler, C., Veselovsky, V., Monea, G., & West, R. (2024). Do Llamas Work in English? On the Latent Language of Multilingual Transformers. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 15366–15394. https://doi.org/10.18653/v1/2024.acl-long.820
Chicago
Wendler, C., V. Veselovsky, G. Monea, and R. West. 2024. “Do Llamas Work in English? On the Latent Language of Multilingual Transformers”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 15366–94. https://doi.org/10.18653/v1/2024.acl-long.820.
Harvard
Wendler, C. et al. (2024) “Do Llamas Work in English? On the Latent Language of Multilingual Transformers”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 15366–15394. Available at: https://doi.org/10.18653/v1/2024.acl-long.820.
Vancouver
1. Wendler C, Veselovsky V, Monea G, West R (2024) Do Llamas Work in English? On the Latent Language of Multilingual Transformers. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 15366–15394

BibTeX

@inproceedings{wendler-etal-2024-llamas,
    title = "Do Llamas Work in {E}nglish? On the Latent Language of Multilingual Transformers",
    author = "Wendler, Chris  and
      Veselovsky, Veniamin  and
      Monea, Giovanni  and
      West, Robert",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.820/",
    doi = "10.18653/v1/2024.acl-long.820",
    pages = "15366--15394"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/