Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models
Tianyi TangWenyang LuoHaoyang HuangDongdong ZhangXiaolei WangXin ZhaoFuru WeiJi-Rong Wen
Reveals that multilingual proficiency in large language models relies on a tiny subset of neurons located in the top and bottom layers, introducing an entropy-based detection method to identify and steer these language-specific units.
Large language models often demonstrate strong multilingual capabilities despite receiving most of their training data in English. Understanding the internal mechanisms that enable these models to process multiple languages remains a critical challenge. This article evaluates how internal network components handle different languages, aiming to locate language-specific regions and determine whether manipulating these components can control output behavior.
To identify these regions, the article introduces Language Activation Probability Entropy, a metric that measures how selectively individual neurons activate across different languages. The analysis evaluated foundation models across various architectures and sizes—primarily LLaMA-2 models ranging from 7 billion to 70 billion parameters and BLOOM with 7.1 billion parameters—tested on multilingual datasets spanning English, Simplified Chinese, French, Spanish, Vietnamese, Indonesian, and Japanese. The researchers measured language modeling performance through perplexity and open-ended text quality using automated evaluation.
The findings show that an LLM's proficiency in a specific language relies predominantly on a very small subset of neurons, representing about 1% or fewer of the total network. Selectively deactivating this tiny fraction severely degrades both comprehension and generation in the targeted language while leaving other languages largely unaffected. Language-specific neurons are concentrated heavily in the bottom and top layers of the model in a U-shaped distribution, while middle layers remain largely language-agnostic. The analysis also revealed that low-resource languages align around dominant training languages, requiring fewer dedicated English neurons in English-heavy models like LLaMA-2. Furthermore, manually activating language-specific neurons successfully directs the model's output language, significantly mitigating errors where the model responds in the wrong language.
These findings indicate that multilingual processing in language models is modular rather than uniformly distributed. This modularity offers significant operational benefits, suggesting that organizations can improve cross-lingual generation and resolve off-target language errors through targeted neuron manipulation without retraining entire models. Practitioners can leverage these insights to optimize cross-lingual transfer from high-resource to low-resource languages and refine continual pre-training workflows.
Future efforts should explore developing efficient strategies for continual pre-training and expanding neuron manipulation techniques to support a broader array of low-resource languages. Decision-makers should note that the identification method depends on relative comparisons across multiple languages and cannot determine language-specific regions in single-language settings. The results demonstrate high empirical consistency across various model sizes, though further testing on models trained with broader multilingual corpora is recommended.
- Paper: Transformer Feed-Forward Layers Are Key-Value Memories, Mor Geva et al. (2020). Establishes the foundational view that feed-forward layers in Transformers act as key-value memory neurons, providing the mechanistic baseline for detecting neuron-level specialization.
- Paper: Language Contamination Helps Explains the Cross-lingual Capabilities of English Pretrained Models, Terra Blevins et al. (2022). Analyzes cross-lingual transfer and multilingual capabilities arising without explicit parallel corpora in foundation models, directly motivating the search for language-specific internal mechanisms.
- Paper: Towards a Common Understanding of Contributing Factors for Cross-Lingual Transfer in Multilingual Language Models: A Review, Fred Philippy et al. (2023). Surveys the underlying factors contributing to emergent cross-lingual transfer in multilingual language models, contextualizing the mechanistic role of language-specific regions.
- Paper: MoEfication: Transformer Feed-forward Layers are Mixtures of Experts, Zhengyan Zhang et al. (2022). Demonstrates sparse neuron activation and functional specialization in Transformer feed-forward networks, forming the basis for isolating specialized language subnetworks.
- Paper: How Multilingual is Multilingual BERT?, Telmo Pires et al. (2019). Pioneered empirical probing into how multilingual models internally process and map distinct languages without explicit translation supervision.
- Paper: What Does BERT Learn about the Structure of Language?, Ganesh Jawahar et al. (2019). Dissects the layer-by-layer hierarchical processing of Transformer representations, motivating the structural discovery of language-specific neurons across top and bottom layers.
- Paper: LLaMA: Open and Efficient Foundation Language Models, Hugo Touvron et al. (2023). Introduces the LLaMA architecture and open foundation models that serve as primary experimental subjects for neuron identification and steering.
- Paper: TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents, Bofei Zhang et al. (2026). Builds directly on neuron-level analysis in BLOOM and LLaMA-2 to investigate language-specific neuron deactivation and internal shared representations across languages.
- Paper: Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages, Zihao Li et al. (2025). Leverages internal layer representations across models like LLaMA-2 and Mistral to quantify and grade cross-lingual performance disparities across high- and low-resource languages.
- Paper: Towards Best Practices of Activation Patching in Language Models: Metrics and Methods, Fred Zhang et al. (2024). Extends the methodology of localized neuron interventions by systematizing metrics and best practices for activation patching and causal steering in language models.
