Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models

Tianyi TangWenyang LuoHaoyang HuangDongdong ZhangXiaolei WangXin ZhaoFuru WeiJi-Rong Wen

article2024ACL109 citations

Reveals that multilingual proficiency in large language models relies on a tiny subset of neurons located in the top and bottom layers, introducing an entropy-based detection method to identify and steer these language-specific units.

Listen

Large language models often demonstrate strong multilingual capabilities despite receiving most of their training data in English. Understanding the internal mechanisms that enable these models to process multiple languages remains a critical challenge. This article evaluates how internal network components handle different languages, aiming to locate language-specific regions and determine whether manipulating these components can control output behavior.

To identify these regions, the article introduces Language Activation Probability Entropy, a metric that measures how selectively individual neurons activate across different languages. The analysis evaluated foundation models across various architectures and sizes—primarily LLaMA-2 models ranging from 7 billion to 70 billion parameters and BLOOM with 7.1 billion parameters—tested on multilingual datasets spanning English, Simplified Chinese, French, Spanish, Vietnamese, Indonesian, and Japanese. The researchers measured language modeling performance through perplexity and open-ended text quality using automated evaluation.

The findings show that an LLM's proficiency in a specific language relies predominantly on a very small subset of neurons, representing about 1% or fewer of the total network. Selectively deactivating this tiny fraction severely degrades both comprehension and generation in the targeted language while leaving other languages largely unaffected. Language-specific neurons are concentrated heavily in the bottom and top layers of the model in a U-shaped distribution, while middle layers remain largely language-agnostic. The analysis also revealed that low-resource languages align around dominant training languages, requiring fewer dedicated English neurons in English-heavy models like LLaMA-2. Furthermore, manually activating language-specific neurons successfully directs the model's output language, significantly mitigating errors where the model responds in the wrong language.

These findings indicate that multilingual processing in language models is modular rather than uniformly distributed. This modularity offers significant operational benefits, suggesting that organizations can improve cross-lingual generation and resolve off-target language errors through targeted neuron manipulation without retraining entire models. Practitioners can leverage these insights to optimize cross-lingual transfer from high-resource to low-resource languages and refine continual pre-training workflows.

Future efforts should explore developing efficient strategies for continual pre-training and expanding neuron manipulation techniques to support a broader array of low-resource languages. Decision-makers should note that the identification method depends on relative comparisons across multiple languages and cannot determine language-specific regions in single-language settings. The results demonstrate high empirical consistency across various model sizes, though further testing on models trained with broader multilingual corpora is recommended.

arXiv: 2402.16438
Cover for Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models

Abstract

Large language models (LLMs) demonstrate remarkable multilingual capabilities without being pre-trained on specially curated multilingual parallel corpora. It remains a challenging problem to explain the underlying mechanisms by which LLMs process multilingual texts. In this paper, we delve into the composition of Transformer architectures in LLMs to pinpoint language-specific regions. Specifically, we propose a novel detection method, language activation probability entropy (LAPE), to identify language-specific neurons within LLMs. Based on LAPE, we conduct comprehensive experiments on several representative LLMs, such as LLaMA-2, BLOOM, and Mistral. Our findings indicate that LLMs' proficiency in processing a particular language is predominantly due to a small subset of neurons, primarily situated in the models' top and bottom layers. Furthermore, we showcase the feasibility to “steer” the output language of LLMs by selectively activating or deactivating language-specific neurons. Our research provides important evidence to the understanding and exploration of the multilingual capabilities of LLMs.

Table of Contents

  • 1 Introduction
  • 2 Identifying Language-Specific Regions
  • 2.1 Background
  • 2.2 Language Activation Probability Entropy
  • 3 Experiments
  • 3.1 Experimental Setup
  • 3.2 Main Perturbation Experiments
  • 3.3 Further Analysis
  • 3.3.1 Distribution and Identification Ratio
  • 3.3.2 Structural Distribution Analysis
  • 3.3.3 Language Dominance Analysis
  • 3.3.4 Case Study
  • 4 Related Work
  • 5 Conclusion
  • Acknowledgement
  • Limitations
  • References
  • A. Appendix

Knowls

  1. Knowl 1 — Language Activation Probability Entropy for Identifying Language-Specific Neurons

    equation

    For a Transformer feed-forward network (FFN), let ll denote the total number of evaluated languages. For the jj-th neuron in the ii-th layer, the activation probability when processing monolingual texts written in language k∈{1,…,l}k \in \{1, \dots, l\} is defined as:

    pi,jk=E[I(act_fn(h~iW1i)j>0)∣language k]p_{i,j}^k = \mathbb{E}\left[\mathbb{I}\left(\text{act\_fn}(\tilde{\mathbf{h}}^i \mathbf{W}_1^i)_j > 0\right) \mid \text{language } k\right]

    where h~i∈Rd\tilde{\mathbf{h}}^i \in \mathbb{R}^d is the hidden state vector output by the multi-head self-attention module in layer ii, W1i∈Rd×4d\mathbf{W}_1^i \in \mathbb{R}^{d \times 4d} is the input projection weight matrix of the FFN, act_fn(⋅)\text{act\_fn}(\cdot) is the non-linear activation function, and I(⋅)\mathbb{I}(\cdot) is the indicator function returning 11 if the activation value strictly exceeds zero and 00 otherwise.

    The vector of activation probabilities pi,j=(pi,j1,pi,j2,…,pi,jl)\mathbf{p}_{i,j} = (p_{i,j}^1, p_{i,j}^2, \dots, p_{i,j}^l) across all evaluated languages is normalized using L1L_1 normalization to yield a probability distribution pi,j′\mathbf{p}'_{i,j}:

    p′i,jk=pi,jk∑m=1lpi,jm{p'}_{i,j}^k = \frac{p_{i,j}^k}{\sum_{m=1}^l p_{i,j}^m}

    The Language Activation Probability Entropy (LAPE\text{LAPE}) for neuron jj in layer ii is calculated as:

    LAPEi,j=−∑k=1lp′i,jklog⁡(p′i,jk)\text{LAPE}_{i,j} = -\sum_{k=1}^l {p'}_{i,j}^k \log({p'}_{i,j}^k)

    A low LAPEi,j\text{LAPE}_{i,j} score signifies that the neuron is selectively activated by one or two specific languages while exhibiting negligible activation for others, designating it as a candidate language-specific neuron.

  2. Knowl 2 — Language-Specific Neuron Detection Procedure

    algorithm

    The detection procedure identifies neurons within an autoregressive Transformer large language model that specialize in processing individual languages by evaluating their Language Activation Probability Entropy (LAPE\text{LAPE}) over monolingual corpora and applying an absolute activation threshold.

    Input: Pre-trained LLM with LL layers and NN neurons per FFN layer, monolingual corpora for ll languages K={1,…,l}K = \{1, \dots, l\} (each comprising 100M tokens sampled from Wikipedia), selection percentile α=0.01\alpha = 0.01 (lowest 1%), activation threshold τ\tau (set to the 95th percentile of all empirical activation probabilities across the model, e.g., τ=0.515\tau = 0.515 for LLaMA-2 70B)
    Output: Language-specific neuron sets SkS_k for each language k∈Kk \in K
    for each layer i∈{1,…,L}i \in \{1, \dots, L\} do
        for each neuron j∈{1,…,N}j \in \{1, \dots, N\} do
            for each language k∈Kk \in K do
                Compute activation probability pi,jk=E[I(act_fn(h~iW1i)j>0)∣language k]p_{i,j}^k = \mathbb{E}[\mathbb{I}(\text{act\_fn}(\tilde{\mathbf{h}}^i \mathbf{W}_1^i)_j > 0) \mid \text{language } k]
            end for
            Normalize distribution p′i,jk=pi,jk/∑m=1lpi,jm{p'}_{i,j}^k = p_{i,j}^k / \sum_{m=1}^l p_{i,j}^m for all k∈Kk \in K
            Compute entropy LAPEi,j=−∑k=1lp′i,jklog⁡(p′i,jk)\text{LAPE}_{i,j} = -\sum_{k=1}^l {p'}_{i,j}^k \log({p'}_{i,j}^k)
        end for
    end for
    Compute threshold θ\theta as the α\alpha-th percentile value of all LAPEi,j\text{LAPE}_{i,j} scores across the entire model
    Initialize Sk←∅S_k \leftarrow \emptyset for every k∈Kk \in K
    for each layer i∈{1,…,L}i \in \{1, \dots, L\} and neuron j∈{1,…,N}j \in \{1, \dots, N\} do
        if LAPEi,j≤θ\text{LAPE}_{i,j} \le \theta then
            for each language k∈Kk \in K do
                if pi,jk≥τp_{i,j}^k \ge \tau then
                    Sk←Sk∪{(i,j)}S_k \leftarrow S_k \cup \{(i, j)\}
                end if
            end for
        end if
    end for
    return {Sk}k∈K\{S_k\}_{k \in K}
  3. Knowl 3 — FFN Neuron Definition and Activation Criterion in Transformer LLMs

    definition

    In an autoregressive Transformer large language model with hidden dimension dd, the multi-head self-attention (MHA) module at layer ii computes intermediate token representation h~i∈Rd\tilde{\mathbf{h}}^i \in \mathbb{R}^d from sequence hidden states Hi−1\mathbf{H}^{i-1}:

    h~i=Attn(hi−1Wqi,Hi−1Wki,Hi−1Wvi)⋅Woi\tilde{\mathbf{h}}^i = \text{Attn}(\mathbf{h}^{i-1}\mathbf{W}_q^i, \mathbf{H}^{i-1}\mathbf{W}_k^i, \mathbf{H}^{i-1}\mathbf{W}_v^i) \cdot \mathbf{W}_o^i

    where Wqi,Wki,Wvi∈Rd×dk\mathbf{W}_q^i, \mathbf{W}_k^i, \mathbf{W}_v^i \in \mathbb{R}^{d \times d_k} and Woi∈Rdk×d\mathbf{W}_o^i \in \mathbb{R}^{d_k \times d} are trainable projection matrices.

    The subsequent Feed-Forward Network (FFN) with intermediate dimension 4d4d maps h~i\tilde{\mathbf{h}}^i to layer output hi∈Rd\mathbf{h}^i \in \mathbb{R}^d. In standard architectures (e.g., BLOOM with GELU activation):

    hi=act_fn(h~iW1i)⋅W2i\mathbf{h}^i = \text{act\_fn}(\tilde{\mathbf{h}}^i \mathbf{W}_1^i) \cdot \mathbf{W}_2^i

    where W1i∈Rd×4d\mathbf{W}_1^i \in \mathbb{R}^{d \times 4d} and W2i∈R4d×d\mathbf{W}_2^i \in \mathbb{R}^{4d \times d}. In architectures using Gated Linear Units (e.g., LLaMA-2 with SwiGLU):

    hi=(act_fn(h~iW1i)⊗(h~iW3i))⋅W2i\mathbf{h}^i = \left(\text{act\_fn}(\tilde{\mathbf{h}}^i \mathbf{W}_1^i) \otimes (\tilde{\mathbf{h}}^i \mathbf{W}_3^i)\right) \cdot \mathbf{W}_2^i

    where W3i∈Rd×4d\mathbf{W}_3^i \in \mathbb{R}^{d \times 4d} and ⊗\otimes denotes element-wise multiplication.

    A single neuron is defined as the linear transformation associated with a single column of W1i\mathbf{W}_1^i followed by the non-linear activation function act_fn(⋅)\text{act\_fn}(\cdot). The jj-th neuron (j∈{1,…,4d}j \in \{1, \dots, 4d\}) in the ii-th FFN layer is considered activated for a given token if and only if its activation value strictly exceeds zero:

    act_fn(h~iW1i)j>0\text{act\_fn}(\tilde{\mathbf{h}}^i \mathbf{W}_1^i)_j > 0

  4. Knowl 4 — Targeted Language Performance Degradation via Language-Specific Neuron Deactivation

    empirical result

    Deactivating the bottom 1% of neurons identified by LAPE for a particular language (setting their activation values to zero during forward inference) induces severe, language-selective capability degradation in both language modeling and open-ended text generation, while leaving unrelated languages largely unaffected.

    On language modeling (measured by perplexity / PPL increase on held-out Wikipedia corpora sampled after September 2022):

    • In LLaMA-2 (7B), deactivating Chinese neurons increases Chinese PPL by +0.58, with cross-lingual PPL changes ≤0.08\le 0.08 on English, French, Spanish, Vietnamese, and Indonesian (Japanese changes by +0.33 due to shared characters). Deactivating French neurons increases French PPL by +0.44 and Spanish by +0.29, with other languages changing by ≤0.07\le 0.07.
    • In LLaMA-2 (70B), deactivating language-specific neurons selectively increases PPL for Chinese (+1.02), French (+0.34), Spanish (+0.44), Vietnamese (+0.13), Indonesian (+0.64), and Japanese (+0.99), with off-diagonal PPL changes remaining below 0.07 (except the Chinese-Japanese pair).
    • Identical selective degradation patterns hold across model architectures and sizes, including BLOOM (7.1B), OPT (6.7B), Mistral (7B), and Phi-2 (2.7B).

    On open-ended generation evaluated by GPT-4 (scored 1 to 10 on the multilingual Vicuna dataset using LLaMA-2 70B), deactivating specific neurons sharply degrades performance in the target language compared to normal and random deactivation baselines:

    Deactivation Setting zh fr es vi id ja
    Normal (no deactivation) 4.30 4.19 3.51 3.70 4.16 2.86
    Random (1% random neurons) 4.18 4.22 3.35 3.53 4.42 2.99
    Chinese (zh) neurons deactivated 2.46 3.56 2.96 3.64 3.56 2.31
    French (fr) neurons deactivated 3.69 2.50 2.29 3.01 3.59 2.76
    Spanish (es) neurons deactivated 3.51 2.57 2.01 3.14 3.34 2.56
    Vietnamese (vi) neurons deactivated 3.93 3.19 2.49 2.74 3.59 2.74
    Indonesian (id) neurons deactivated 3.67 3.10 2.67 3.21 2.84 2.80
    Japanese (ja) neurons deactivated 3.21 3.69 3.07 3.49 3.37 1.84
  5. Knowl 5 — Skewed U-Shaped Layer Distribution of Language-Specific Neurons

    empirical result

    Language-specific neurons identified by LAPE exhibit a skewed "U"-shaped distribution across the depth of autoregressive Transformer LLMs, concentrating overwhelmingly in the bottom-most and top-most layers, while intermediate layers are dominated by language-agnostic representations.

    In LLaMA-2 (70B, containing 80 layers and approximately 2.29M total neurons across layers):

    • Layer 1 contains 994 language-specific neurons across all evaluated languages (238 en, 199 zh, 45 fr, 43 es, 28 vi, 47 id, 195 ja).
    • Layer 2 contains the global bottom peak with 7,444 language-specific neurons (117 en, 886 zh, 1,056 fr, 1,155 es, 1,589 vi, 897 id, 1,184 ja).
    • Middle layers (layers 5 through 47) contain roughly 100 or fewer language-specific neurons per layer in total (e.g., layer 16 has 0 en, 4 zh, 1 fr, 4 es, 3 vi, 1 id, 5 ja).
    • Upper layers exhibit a continuous rise: layer 77 contains 1,362 neurons, layer 78 contains 1,356 neurons, layer 79 contains 1,878 neurons, and layer 80 contains 1,813 neurons.

    This structural U-shaped concentration also holds for LLaMA-2 7B (32 layers), LLaMA-2 13B (40 layers), and BLOOM 7B (30 layers), confirming that bottom layers handle language-specific input projection into shared semantics, while top layers handle output vocabulary mapping.

  6. Knowl 6 — Inverse Relationship Between Neuron Density and Sentence Embedding Similarity Across Layers

    empirical result

    The depth-wise density of language-specific neurons is inversely correlated with the degree of cross-lingual semantic alignment across Transformer layers.

    When evaluating semantically aligned multilingual sentence pairs from the multilingual Vicuna dataset across layers of LLaMA-2 (70B), the mean Sentence Embedding Similarity (SES) between all language pairs demonstrates a inverted trajectory compared to the U-shaped neuron distribution:

    1. Bottom Layers (Layers 1–4): SES begins near zero and increases rapidly, coinciding with high language-specific neuron counts (e.g., over 7,000 neurons in layer 2) required to project surface text of diverse languages into a unified semantic space.
    2. Intermediate Layers (Layers 5–50): SES plateaus at its maximum (approaching 1.0), indicating a language-agnostic, shared conceptual space where language-specific neuron counts drop to their minimum (~100 neurons per layer).
    3. Upper Layers (Layers 51–80): SES steadily decreases from ~0.8 to ~0.4, coinciding with the sharp resurgence in language-specific neurons (exceeding 1,000 per layer in layers 77–80) as representations are projected from the shared semantic space into specific vocabulary tokens for token generation.
  7. Knowl 7 — Pre-training Language Dominance and Cross-Lingual Semantic Centering

    empirical result

    In pre-trained LLMs, the quantity of language-specific neurons allocated to a language reflects its proportion in the pre-training corpus, and multilingual representations center around dominant pre-training languages.

    Language Code LLaMA-2 Pre-train % BLOOM Pre-train % LLaMA-2 (70B) Specific Neurons
    English en 89.70% 33.68% 836
    Chinese zh 0.13% 18.13% 5,153
    French fr 0.16% 14.46% 6,082
    Spanish es 0.13% 12.16% 6,154
    Vietnamese vi 0.08% 3.04% 4,980
    Indonesian id 0.03% 1.39% 6,106
    Japanese ja 0.10% 0.00% 5,216

    To quantify language dominance, sentence embeddings hki\mathbf{h}_k^i of language kk at layer ii are transformed into the space centered on target language cc using language mean vectors vki\mathbf{v}_k^i and vci\mathbf{v}_c^i:

    h^ki=hki−vki+vci\hat{\mathbf{h}}_k^i = \mathbf{h}_k^i - \mathbf{v}_k^i + \mathbf{v}_c^i

    Computing the mean SES between aligned texts of language kk and target language cc reveals:

    • In LLaMA-2 (70B), English achieves a dominance score (mean SES ≈0.82\approx 0.82) significantly higher than all other languages (0.65−0.720.65 - 0.72), indicating that low-resource representations are anchored to English.
    • In BLOOM (170B), which was trained on balanced multilingual data, multiple languages (both English and Chinese) exhibit high dominance scores (mean SES ≈0.8\approx 0.8).
  8. Knowl 8 — Output Language Steering and Off-Target Mitigation via Language Neuron Activation

    empirical result

    Selectively activating or deactivating language-specific neurons allows direct steering of an LLM's output language during generation and mitigates the off-target response problem (where models erroneously reply in English to non-English queries).

    1. Off-Target Mitigation: Manually activating a language's specific neurons by clamping their activation values to the average activation value observed for that language substantially improves language adherence and generation quality on the multilingual Vicuna benchmark for LLaMA-2 (70B):
    Metric Setting zh fr es vi id ja
    Language Accuracy Normal 0.87 0.73 0.81 0.60 0.40 0.79
    Language Accuracy Steered 0.99 0.90 0.93 0.97 0.99 1.00
    Content Quality (1–10) Normal 4.30 4.19 3.51 3.70 4.16 2.86
    Content Quality (1–10) Steered 4.57 4.35 4.02 3.57 4.28 2.91
    1. Cross-Lingual Generation Steering: When an input question is posed in one language (e.g., Spanish) while Spanish-specific neurons are deactivated and Chinese-specific neurons are activated, LLaMA-2 generates a fluent, grammatically coherent response written entirely in Chinese answering the Spanish query.
  9. Knowl 9 — Cross-Lingual Neuron Sharing Between Chinese and Japanese

    empirical result

    In Transformer LLMs such as LLaMA-2 (7B and 70B), language-specific neurons identified for Chinese and Japanese exhibit substantial physical overlap, with approximately 25% of their identified language-specific neurons being shared between the two languages.

    Because of this neuron sharing:

    • Deactivating Chinese-specific neurons causes a cross-lingual perplexity (PPL) increase on Japanese text (+0.33 on LLaMA-2 7B, +0.37 on LLaMA-2 70B), which is substantially higher than the impact on European or Southeast Asian languages (≤0.08\le 0.08).
    • Deactivating Japanese-specific neurons causes a cross-lingual PPL increase on Chinese text (+0.29 on LLaMA-2 7B, +0.45 on LLaMA-2 70B), while impacting other languages by ≤0.06\le 0.06.
    • In open-ended generation with LLaMA-2 (70B), deactivating Chinese neurons degrades Japanese text generation score from 2.86 to 2.31, and deactivating Japanese neurons degrades Chinese text generation score from 4.30 to 3.21. This shared allocation is attributed to the overlap in Han characters (Kanji) and shared lexical components between the two written languages.
  10. Knowl 10 — Methodological Limitations of Entropy-Based Multilingual Neuron Detection

    limitation

    The Language Activation Probability Entropy (LAPE) method and its associated neuron manipulation techniques have several defined constraints:

    1. Relative Evaluation Dependency: LAPE computes entropy relative to a specified set of multiple comparison languages. It cannot establish an absolute threshold for language specificity when only a single language corpus is available.
    2. Undefined Distinction Criteria for Language Resources: The method identifies empirical differences between high- and low-resource languages (e.g., lower neuron counts in dominant languages) but does not provide formal theoretical criteria for defining the transition between high-resource and low-resource regimes inside LLM architectures.
    3. Scalability to Massive Language Sets: The behavior of LAPE has not been verified on massive language sets spanning hundreds of low-resource languages where high-quality monolingual corpora are scarce.
    4. Rudimentary Activation Steering: Neuron manipulation by clamping activation values to empirical means is an early-stage proof of concept; leveraging identified language-specific neurons for continual pre-training, parameter-efficient adaptation, or cross-lingual transfer remains unexplored.

Coverage note — None was omitted; all primary contributions—including the mathematical LAPE formulation, neuron detection algorithm, FFN neuron activation definitions, cross-model perturbation evaluations, U-shaped layer distributions, sentence embedding similarity dynamics, language dominance findings, steering experiments, cross-lingual neuron overlap, and stated limitations—have been fully extracted.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  2. 2.Kabir Ahuja, Shanu Kumar, Sandipan Dandapat, and Monojit Choudhury. 2022. Multi task learning for zero shot performance prediction of multilingual models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5454–5467, Dublin, Ireland. Association for Computational Linguistics.
  3. 3.Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023. Palm 2 technical report. arXiv preprint arXiv:2305.10403.
  4. 4.Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2018. Identifying and controlling important neurons in neural machine translation. arXiv preprint arXiv:1811.01157.
  5. 5.David Bau, Jun-Yan Zhu, Hendrik Strobelt, Agata Lapedriza, Bolei Zhou, and Antonio Torralba. 2020. Understanding the role of individual units in a deep neural network. Proceedings of the National Academy of Sciences, 117(48):30071–30078.
  6. 6.Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders. 2023. Language models can explain neurons in language models. URL https://openaipublic. blob. core. windows. net/neuron-explainer/paper/index. html.(Date accessed: 14.05. 2023).
  7. 7.Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah. 2023. Towards monosemanticity: Decomposing language models with dictionary learning. Transformer Circuits Thread. Https://transformer-circuits.pub/2023/monosemantic-features/index.html.
  8. 8.Steven Cao, Nikita Kitaev, and Dan Klein. 2020. Multilingual alignment of contextual word representations. arXiv preprint arXiv:2002.03518.
  9. 9.Yuheng Chen, Pengfei Cao, Yubo Chen, Kang Liu, and Jun Zhao. 2023a. Journey to the center of the knowledge neurons: Discoveries of language-independent knowledge neurons and degenerate knowledge neurons. arXiv preprint arXiv:2308.13198.
  10. 10.Zhihong Chen, Shuo Yan, Juhao Liang, Feng Jiang, Xiangbo Wu, Fei Yu, Guiming Hardy Chen, Junying Chen, Hongbo Zhang, Li Jianquan, Wan Xiang, and Benyou Wang. 2023b. MultilingualSIFT: Multilingual Supervised Instruction Fine-tuning.
  11. 11.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  12. 12.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020a. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  13. 13.Alexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer, and Veselin Stoyanov. 2020b. Emerging cross-lingual structure in pretrained language models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6022–6034, Online. Association for Computational Linguistics.
  14. 14.Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2022. Knowledge neurons in pretrained transformers. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8493–8502, Dublin, Ireland. Association for Computational Linguistics.
  15. 15.Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, Anthony Bau, and James Glass. 2019. What is one grain of sand in the desert? analyzing individual neurons in deep nlp models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 6309–6317.
  16. 16.Fahim Dalvi, Hassan Sajjad, Nadir Durrani, and Yonatan Belinkov. 2020. Analyzing redundancy in pretrained transformer models. arXiv preprint arXiv:2004.04010.
  17. 17.Ameet Deshpande, Partha Talukdar, and Karthik Narasimhan. 2022. When is BERT multilingual? isolating crucial ingredients for cross-lingual transfer. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3610–3623, Seattle, United States. Association for Computational Linguistics.
  18. 18.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  19. 19.Philipp Dufter and Hinrich Schütze. 2020. Identifying elements essential for BERT’s multilinguality. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4423–4437, Online. Association for Computational Linguistics.
  20. 20.Angela D Friederici. 2011. The brain basis of language processing: from structure to function. Physiological reviews, 91(4):1357–1392.
  21. 21.Jiatao Gu, Yong Wang, Kyunghyun Cho, and Victor O.K. Li. 2019. Improved zero-shot neural machine translation via ignoring spurious correlations. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1258–1268, Florence, Italy. Association for Computational Linguistics.
  22. 22.Wes Gurnee, Theo Horsley, Zifan Carl Guo, Tara Rezaei Kheirkhah, Qinyi Sun, Will Hathaway, Neel Nanda, and Dimitris Bertsimas. 2024. Universal neurons in gpt2 language models. arXiv preprint arXiv:2401.12181.
  23. 23.Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas. 2023. Finding neurons in a haystack: Case studies with sparse probing. arXiv preprint arXiv:2305.01610.
  24. 24.Dan Hendrycks and Kevin Gimpel. 2016. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415.
  25. 25.Mojan Javaheripi, Sébastien Bubeck, Marah Abdin, Jyoti Aneja, Sebastien Bubeck, Caio César Teodoro Mendes, Weizhu Chen, Allie Del Giorno, Ronen Eldan, Sivakanth Gopi, et al. 2023. Phi-2: The surprising power of small language models. Microsoft Research Blog.
  26. 26.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825.
  27. 27.Arjun R Khanna, William Muñoz, Young Joon Kim, Yoav Kfir, Angelique C Paulk, Mohsen Jamali, Jing Cai, Martina L Mustroph, Irene Caprara, Richard Hardstone, et al. 2024. Single-neuronal elements of speech production in humans. Nature, pages 1–8.
  28. 28.Saurabh Kulshreshtha, José Luis Redondo-García, and Ching-Yun Chang. 2020. Cross-lingual alignment methods for multilingual bert: A comparative study. arXiv preprint arXiv:2009.14304.
  29. 29.Anne Lauscher, Vinit Ravishankar, Ivan Vulić, and Goran Glavaš. 2020. From zero to hero: On the limitations of zero-shot language transfer with multilingual Transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4483–4499, Online. Association for Computational Linguistics.
  30. 30.Jesse Mu and Jacob Andreas. 2020. Compositional explanations of neurons. Advances in Neural Information Processing Systems, 33:17153–17163.
  31. 31.Vinod Nair and Geoffrey E. Hinton. 2010. Rectified linear units improve restricted boltzmann machines. In Proceedings of the 27th International Conference on International Conference on Machine Learning, ICML’10, page 807–814, Madison, WI, USA. Omnipress.
  32. 32.Xuan-Phi Nguyen, Wenxuan Zhang, Xin Li, Mahani Aljunied, Qingyu Tan, Liying Cheng, Guanzheng Chen, Yue Deng, Sen Yang, Chaoqun Liu, et al. 2023. Seallms–large language models for southeast asia. arXiv preprint arXiv:2312.00738.
  33. 33.Thomas Parr, Giovanni Pezzulo, and Karl J Friston. 2022. Active inference: the free energy principle in mind, brain, and behavior. MIT Press.
  34. 34.Fred Philippy, Siwen Guo, and Shohreh Haddadan. 2023. Towards a common understanding of contributing factors for cross-lingual transfer in multilingual language models: A review. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5877–5891, Toronto, Canada. Association for Computational Linguistics.
  35. 35.Telmo Pires, Eva Schlinger, and Dan Garrette. 2019. How multilingual is multilingual BERT? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4996–5001, Florence, Italy. Association for Computational Linguistics.
  36. 36.Alec Radford, Rafal Jozefowicz, and Ilya Sutskever. 2017. Learning to generate reviews and discovering sentiment. arXiv preprint arXiv:1704.01444.
  37. 37.Hassan Sajjad, Nadir Durrani, and Fahim Dalvi. 2022. Neuron-level interpretation of deep NLP models: A survey. Transactions of the Association for Computational Linguistics, 10:1285–1303.
  38. 38.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  39. 39.Rico Sennrich, Jannis Vamvas, and Alireza Mohammadshahi. 2023. Mitigating hallucinations and off-target machine translation with source-contrastive and language-contrastive decoding. arXiv preprint arXiv:2309.07098.
  40. 40.Noam Shazeer. 2020. Glu variants improve transformer. arXiv preprint arXiv:2002.05202.
  41. 41.Karolina Stanczak, Edoardo Ponti, Lucas Torroba Hennigen, Ryan Cotterell, and Isabelle Augenstein. 2022. Same neurons, different languages: Probing morphosyntax in multilingual pre-trained models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1589–1598, Seattle, United States. Association for Computational Linguistics.
  42. 42.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  43. 43.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023a. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  44. 44.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023b. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  45. 45.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  46. 46.Elena Voita, Javier Ferrando, and Christoforos Nalmpantis. 2023. Neurons in large language models: Dead, n-gram, positional. arXiv preprint arXiv:2309.04827.
  47. 47.Xiaozhi Wang, Kaiyue Wen, Zhengyan Zhang, Lei Hou, Zhiyuan Liu, and Juanzi Li. 2022. Finding skill neurons in pre-trained transformer-based language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11132–11152, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  48. 48.Zihan Wang, Stephen Mayhew, Dan Roth, et al. 2019. Cross-lingual ability of multilingual bert: An empirical study. arXiv preprint arXiv:1912.07840.
  49. 49.Ji Xin, Jimmy Lin, and Yaoliang Yu. 2019. What part of the neural network does this? understanding lstms by measuring and dissecting neurons. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5823–5830.
  50. 50.Shaoyang Xu, Junzhuo Li, and Deyi Xiong. 2023. Language representation projection: Can we transfer factual knowledge across languages in multilingual language models? In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 3692–3702, Singapore. Association for Computational Linguistics.
  51. 51.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online. Association for Computational Linguistics.
  52. 52.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  53. 53.Jun Zhao, Zhihao Zhang, Yide Ma, Qi Zhang, Tao Gui, Luhui Gao, and Xuanjing Huang. 2023a. Unveiling a core linguistic region in large language models. arXiv preprint arXiv:2310.14928.
  54. 54.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023b. A survey of large language models. arXiv preprint arXiv:2303.18223.
  55. 55.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685.

Citation

MLA
Tang, T., et al. “Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 5701–15, https://doi.org/10.18653/v1/2024.acl-long.309.
APA
Tang, T., Luo, W., Huang, H., Zhang, D., Wang, X., Zhao, X., Wei, F., & Wen, J.-R. (2024). Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5701–5715. https://doi.org/10.18653/v1/2024.acl-long.309
Chicago
Tang, T., W. Luo, H. Huang, et al. 2024. “Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5701–15. https://doi.org/10.18653/v1/2024.acl-long.309.
Harvard
Tang, T. et al. (2024) “Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5701–5715. Available at: https://doi.org/10.18653/v1/2024.acl-long.309.
Vancouver
1. Tang T, Luo W, Huang H, Zhang D, Wang X, Zhao X, Wei F, Wen J-R (2024) Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 5701–5715

BibTeX

@inproceedings{tang-etal-2024-language,
    title = "Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models",
    author = "Tang, Tianyi  and
      Luo, Wenyang  and
      Huang, Haoyang  and
      Zhang, Dongdong  and
      Wang, Xiaolei  and
      Zhao, Xin  and
      Wei, Furu  and
      Wen, Ji-Rong",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.309/",
    doi = "10.18653/v1/2024.acl-long.309",
    pages = "5701--5715"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/