Do Llamas Work in English? On the Latent Language of Multilingual Transformers
Chris WendlerVeniamin VeselovskyGiovanni MoneaRobert West
Reveals through logit lens analysis that multilingual transformer models route non-English inputs through an internal, English-aligned concept space in intermediate layers before decoding them into the target language.
Modern multilingual artificial intelligence models achieve strong performance across numerous global languages despite being trained predominantly on English text. This disparity raises a critical question regarding whether these systems internally translate inputs into English to process information before generating output in the requested language. Understanding whether models rely on an internal English bridge is essential for identifying subtle linguistic biases, cultural skew, and performance limitations when deploying artificial intelligence globally.
The main objective of the article is to evaluate empirically whether multilingual transformer models use English as an internal pivot language during text generation and to demonstrate the geometrical structure of internal representations across processing layers.
To evaluate this question, the researchers applied mechanistic interpretability methods across multiple sizes of the open-weight Llama-2 family (7B, 13B, and 70B parameters) and validated key findings on Mistral-7B. They designed controlled text-completion benchmarks covering Chinese, German, French, Russian, and Estonian across translation, word repetition, and fill-in-the-blank cloze tasks. The analysis used the logit lens technique—which prematurely projects hidden numerical states in intermediate layers into human-readable word probabilities—alongside geometric analysis measuring the mathematical alignment and energy between hidden vectors and output token embeddings.
The investigation produced three main findings. First, internal processing moves consistently through three distinct operational phases across model scales: an initial input-processing phase in early layers where representations show high entropy and no language emerges; a middle concept phase where entropy drops sharply and semantically correct English words dominate intermediate predictions; and a final target-language phase in the last layers where output probability abruptly shifts to the requested language. Second, geometric analysis revealed that intermediate latent states do not represent literal English text; instead, they operate in an abstract concept space that carries non-linguistic context but sits geometrically closer to English due to training data dominance. Third, tokenization heavily impacts this trajectory: languages with dedicated single-token vocabulary entries (such as Chinese) bypass or reduce the English detour during simple repetition tasks, whereas languages split into multi-token fragments (such as Estonian) are forced deeper through English-aligned concept spaces.
These findings indicate that multilingual models do not perform explicit step-by-step translation, but rather navigate a semantic concept space that is structurally biased toward English. For organizations deploying multilingual systems, this introduces risks of Anglocentric bias, subtle shifts in emotional framing, and distorted reasoning on non-Western cultural topics. Furthermore, underrepresented languages face increased computational inefficiencies and error rates due to vocabulary fragmentation.
To mitigate these risks, developers and policymakers should prioritize rebalancing pretraining corpora, expanding multilingual vocabularies, and adopting culturally balanced tokenizers. Before deploying systems in sensitive multilingual environments, organizations should conduct targeted audits for Anglocentric cognitive drift. Further research should extend these evaluations to complex, multi-token reasoning tasks, explore training interventions on balanced datasets, and examine closed-source commercial models where internal weights cannot currently be accessed.
- Paper: BERT Rediscovers the Classical NLP Pipeline, Ian Tenney et al. (2019). It establishes the layer-by-layer probing methodology showing how transformer representations evolve from surface syntax to abstract semantics across network depth.
- Paper: How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings, Kawin Ethayarajh (2019). It provides the geometric foundations for analyzing representation spaces, anisotropy, and layerwise contextuality in transformer embeddings.
- Paper: Language Contamination Helps Explains the Cross-lingual Capabilities of English Pretrained Models, Terra Blevins et al. (2022). It documents how incidental multilingual pretraining data in English-centric corpora establishes the latent cross-lingual transfer capabilities examined in the source.
- Paper: How Multilingual is Multilingual BERT?, Telmo Pires et al. (2019). It introduces essential empirical investigations into how multilingual transformers generalize across languages without explicit alignment signals.
- Paper: Analyzing Encoded Concepts in Transformer Language Models, Hassan Sajjad et al. (2022). It details how contextual representations cluster into distinct human-interpretable concepts across the intermediate layers of transformer architectures.
- Paper: What Does BERT Learn about the Structure of Language?, Ganesh Jawahar et al. (2019). It maps the structural hierarchy of transformer representations from surface lexical tokens to abstract semantic representations across successive layers.
- Paper: TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents, Bofei Zhang et al. (2026). It extends the mechanistic study of multilingual routing by isolating specific language versus concept neurons across BLOOM and LLaMA2 architectures.
- Paper: Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages, Zihao Li et al. (2025). It builds directly on the finding of English-centric intermediate geometry to formulate a quantitative metric for ranking multilingual LLM proficiency.
- Paper: The Linear Representation Hypothesis and the Geometry of Large Language Models, Kiho Park et al. (2024). It formalizes the linear representation and geometric concept space properties that underpin intermediate latent steering in large models like LLaMA-2.
- Paper: Unintended Impacts of LLM Alignment on Global Representation, Michael J. Ryan et al. (2024). It investigates the downstream sociolinguistic disparities and alignment impacts that emerge from the Anglocentric biases identified in multilingual foundation models.
- Paper: MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks, Sanchit Ahuja et al. (2024). It broadens the empirical evaluation of cross-lingual performance gaps across dozens of typologically diverse languages and tasks.
