Finding Skill Neurons in Pre-trained Transformer-based Language Models

Xiaozhi WangKaiyue WenZhengyan ZhangLei HouZhiyuan LiuJuanzi Li

article2022EMNLP79 citations

Reveals that pre-trained Transformers contain task-predictive "skill neurons" formed during pre-training, providing a concrete method to identify them via prompt tuning and apply them to model pruning and transferability estimation.

Listen

Transformer-based language models have become the foundation for modern natural language processing, yet how these large systems store and execute specific skills internally remains a major black-box challenge. Understanding where task-specific capabilities reside is essential for controlling model behavior, diagnosing errors, and reducing the computational cost of running large models.

The article demonstrates that individual components within pre-trained Transformers, specifically individual units called skill neurons located in feed-forward networks, directly encode the high-level capabilities required to solve specific tasks.

To identify these components, the researchers used prompt tuning, a lightweight adaptation method that adds small trainable prompt tokens to inputs while keeping the underlying model frozen. By measuring neuron activations on these prompt tokens across seven English classification benchmarks spanning sentiment analysis, natural language inference, and topic classification, the researchers calculated each neuron's ability to predict correct labels. They evaluated these identified neurons through noise perturbation experiments, cross-task correlation analyses, untuned baseline comparisons, and downstream efficiency applications.

The analysis yielded five primary findings. First, skill neurons emerge consistently across all evaluated tasks, and the top-ranking neurons alone achieve task accuracy comparable to full prompt-tuned models (e.g., 91.6% versus 91.8% on standard sentiment benchmarks). Second, perturbing these specific neurons with random noise causes severe drops in model accuracy, whereas perturbing random neurons has little effect. Third, skill neurons are task-specific and concentrated in the higher layers of the network, with closely related tasks sharing similar skill neuron rankings. Fourth, skill neurons originate during initial pre-training rather than fine-tuning, as they remain critical across distinct tuning methods like adapters and BitFit and can even be detected with untuned prompts. Fifth, applying skill neuron identification to network pruning allows models to deactivate 98% of neurons in upper layers, reducing total model parameters to 66.6% and delivering an approximate 1.4x inference speedup on central processing units with minimal accuracy loss.

These findings indicate that large language models naturally develop modular, task-specific capabilities during pre-training. For organizations deploying language models, this opens direct opportunities to lower computing costs and deployment latency by selectively pruning inactive neurons. It also provides a practical mechanism to measure task similarity and improve transfer learning, as incorporating top skill neurons increased transferability prediction correlations from 0.53 to 0.71.

Organizations evaluating language model efficiency should consider piloting skill-neuron-based structured pruning for latency-sensitive applications. Engineering teams should also leverage skill neuron metrics to identify compatible tasks for cross-task adaptation. Further analysis should explore finding skill neurons without prompt tuning to streamline identification workflows.

The conclusions carry high empirical confidence across the evaluated settings, but leaders should note three boundaries: the experiments evaluated a single base architecture (RoBERTa), focused exclusively on English text, and examined only classification tasks. Additional validation is recommended before applying these techniques to generative workflows, multilingual deployments, or alternative model architectures.

Wang et al (2022).pdf
Cover for Finding Skill Neurons in Pre-trained Transformer-based Language Models

Abstract

Transformer-based pre-trained language models have demonstrated superior performance on various natural language processing tasks. However, it remains unclear how the skills required to handle these tasks distribute among model parameters. In this paper, we find that after prompt tuning for specific tasks, the activations of some neurons within pre-trained Transformers¹ are highly predictive of the task labels. We dub these neurons skill neurons and confirm they encode task-specific skills by finding that: (1) Skill neurons are crucial for handling tasks. Performances of pre-trained Transformers on a task significantly drop when corresponding skill neurons are perturbed. (2) Skill neurons are task-specific. Similar tasks tend to have similar distributions of skill neurons. Furthermore, we demonstrate the skill neurons are most likely generated in pre-training rather than fine-tuning by showing that the skill neurons found with prompt tuning are also crucial for other fine-tuning methods freezing neuron weights, such as the adapter-based tuning and BitFit. We also explore the applications of skill neurons, including accelerating Transformers with network pruning and building better transferability indicators. These findings may promote further research on understanding Transformers. The source code can be obtained from https://github.com/THU-KEG/Skill-Neuron.

Table of Contents

  • 1 Introduction
  • 2 Preliminary
  • 2.1 Prompt Tuning
  • 2.2 Neurons in Transformers
  • 2.3 Investigation Setup
  • 3 Finding Skill Neurons
  • 3.1 Binary Classification Task
  • 3.2 Multi-class Classification Task
  • 4 Do Skill Neurons Encode Skills?
  • 4.1 Skill Neurons Generally and Stably Emerge
  • 4.2 Skill Neurons are Crucial for Handling Tasks
  • 4.3 Skill Neurons are Task-specific
  • 4.4 Skill Neurons are not from Word Selectivity
  • 5 Where do Skill Neurons Come from?
  • 6 Application
  • 6.1 Network Pruning
  • 6.2 Transferability Indicator
  • 7 Related Work
  • 8 Conclusion and Future Work
  • Limitations
  • Acknowledgements
  • References
  • Appendices
  • A Details about Investigated Tasks
  • A.1 Sentiment Analysis
  • A.2 Natural Language Inference
  • A.3 Topic Classification
  • B Implementation Details
  • C More Predictivity Distributions
  • D More Neuron Perturbation Results
  • D.1 Performance Dropping Trends for Prompt Tuning
  • D.2 Performance Dropping Trends for Adapter-based Tuning
  • D.3 Performance Dropping Trends for BitFit
  • E Layer-wise Correlations between Neuron Predictivity Orders of Different Tasks
  • F More Word Selectivity Results
  • G Discussions on Neuron-Finding Design Choices
  • H Experiments following Morcos et al. (2018)

Knowls

  1. Knowl 1 — Prompt-based scoring identifies task-predictive FFN neurons

    model/method

    A skill neuron is a feed-forward-network (FFN) neuron in a pre-trained Transformer whose activation on a soft-prompt token predicts labels for a particular classification task. In an FFN, a neuron is one coordinate of the hidden activation vector; its activation depends on the input token representation and its corresponding FFN weights. To score neurons for a binary task, prompt-tune the model on the training data while freezing the model parameters. For each neuron NN and soft-prompt token pp, compute its mean training-set activation as the baseline, and classify an example as label 1 when its activation exceeds that baseline. Score the neuron on the development set using accuracy, but count either direction of association as predictive:

    \bar a(N,p) &= \frac{1}{|D_{\mathrm{train}}|}\sum_{(x,y)\in D_{\mathrm{train}}} a(N,p,x),\\ \operatorname{Acc}(N,p) &= \frac{1}{|D_{\mathrm{dev}}|}\sum_{(x,y)\in D_{\mathrm{dev}}}\mathbf{1}\!\left[\mathbf{1}[a(N,p,x)>\bar a(N,p)]=y\right],\\ \operatorname{Pred}(N,p) &= \max\{\operatorname{Acc}(N,p),1-\operatorname{Acc}(N,p)\},\\ \operatorname{Pred}(N) &= \frac{1}{5}\sum_{r=1}^{5}\max_{p\in P_r}\operatorname{Pred}(N,p). \end{aligned}$$ Here $D_{\mathrm{train}}$ and $D_{\mathrm{dev}}$ are the training and development examples; $x$ is an input, $y\in\{0,1\}$ is its label, $a(N,p,x)$ is the activation of neuron $N$ on prompt token $p$ for input $x$, and $P_r$ is the set of soft prompts from random prompt-tuning trial $r$. Thus, the final score averages across five trials, taking the most predictive prompt token within each trial. Rank neurons by this score and select the highest-ranked ones. For a multiclass task, decompose the labels into binary subtasks, rank neurons for each subtask using the original task's prompts, and combine equal numbers of top-ranked neurons from the subtasks.
  2. Knowl 2 — Highly predictive neurons match prompt-tuning performance across seven tasks

    empirical result

    The authors evaluated skill-neuron prediction on RoBERTaBASE (110 million parameters, 12 Transformer layers) across seven English classification datasets: SST-2, IMDB, and TweetEval sentiment; MNLI and QNLI inference; and AG News and DBpedia topic classification. Prompt tuning used 127 soft prompts, initialized from a normal distribution with standard deviation 0.03, and Adam with learning rate 0.001 and batch size 8; results are means and standard deviations over five random trials. The table reports task accuracy in percent. For binary tasks, the skill-neuron score is the predictivity of the top-ranked neuron; for multiclass tasks, it is the accuracy of a logistic-regression classifier using the top neuron from each decomposed binary subtask. Across the tasks, these neuron-based predictors approach the performance of prompt tuning.

    Task Prompt tuning Skill neuron(s)
    SST-2 91.8±0.591.8\pm0.5 91.6±0.391.6\pm0.3
    IMDB 91.6±0.591.6\pm0.5 92.0±0.392.0\pm0.3
    Tweet 70.0±0.270.0\pm0.2 56.0±3.256.0\pm3.2
    MNLI 76.8±1.876.8\pm1.8 74.7±2.574.7\pm2.5
    QNLI 85.7±0.785.7\pm0.7 86.0±0.486.0\pm0.4
    AG News 98.8±0.198.8\pm0.1 98.9±0.198.9\pm0.1
    DBpedia 99.7±0.199.7\pm0.1 99.8±0.199.8\pm0.1
  3. Knowl 3 — Perturbing task-ranked skill neurons damages task performance more than random perturbation

    empirical result

    To test whether the identified neurons matter for task performance, the authors added Gaussian noise with mean 00 and standard deviation 0.10.1 to selected neuron activations in prompt-tuned RoBERTaBASE models. Neurons were perturbed in descending order of their task predictivity, and the resulting accuracy curves were compared with curves from perturbing neurons in random order. Across the investigated tasks, perturbing the task-ranked neurons reduced accuracy more than perturbing randomly selected neurons; the paper illustrates this contrast for Tweet and reports similar trends for the other tasks. This intervention supports the claim that high-predictivity neurons contribute to handling the corresponding tasks, rather than merely correlating with their labels.

  4. Knowl 4 — Skill-neuron rankings and intervention effects are task-specific

    empirical result

    The authors compared neuron-predictivity rankings across seven tasks using Spearman rank correlations, averaged over RoBERTaBASE's 12 layers. Rankings were more strongly correlated between tasks of the same type—sentiment tasks with one another, inference tasks with one another, and topic tasks with one another—than between dissimilar task types. They also defined the neuronal importance of a source task to an evaluation task as the area between the evaluation-accuracy curve obtained by perturbing neurons in the source task's predictivity order and the curve obtained by random perturbation. After normalizing source-task importances as z-scores within each evaluation task, same-type source tasks generally had greater importance. Layer-wise comparisons further showed that task rankings become more distinct toward higher layers. Together, these results support task-specific rather than purely task-general interpretations of the top-ranked skill neurons.

  5. Knowl 5 — Untuned prompts reveal predictive neurons in pre-trained models

    empirical result

    The authors tested whether identifying predictive neurons required prompt tuning by using randomly generated soft prompts and human-written hard prompts without tuning. The table gives the accuracy (%) of the resulting top skill neurons, compared with random guessing and with results from a randomly initialized model given random prompts. The nontrivial scores from the pre-trained model under both untuned prompt conditions, and their generally large advantage over the randomly initialized model, are evidence the authors use to argue that skill neurons are most likely acquired during pre-training. This is empirical evidence, not a demonstration of when the neurons arose.

    Task Random guess Random model Random prompt Hard prompt
    SST-2 50.0 52.8±0.452.8\pm0.4 78.1±0.478.1\pm0.4 83.3
    IMDB 50.0 58.0±0.758.0\pm0.7 76.7±2.076.7\pm2.0 75.1
    Tweet 33.3 48.3±0.048.3\pm0.0 48.2±1.848.2\pm1.8 48.6
    MNLI 33.3 32.2±0.432.2\pm0.4 39.8±1.139.8\pm1.1 40.5
    QNLI 50.0 54.3±0.854.3\pm0.8 69.5±0.569.5\pm0.5 65.2
    AG News 50.0 62.7±0.362.7\pm0.3 96.0±0.396.0\pm0.3 95.9
    DBpedia 50.0 60.9±0.460.9\pm0.4 98.8±0.198.8\pm0.1 99.2

    The random-model column refers to randomly initialized models evaluated with random prompts. Standard deviations are reported for the random-model and random-prompt results over five trials.

  6. Knowl 6 — Prompt-discovered neurons remain important after adapter tuning and BitFit

    empirical result

    The authors tested whether skill neurons discovered through prompt tuning also mattered in models adapted by methods with different training dynamics. Adapter-based tuning updates added adapter layers, and BitFit updates bias vectors; both methods leave the original FFN neuron weights fixed. In both types of adapted models, perturbing neurons in the prompt-derived predictivity order reduced task performance more than random-order perturbation, and source tasks showed greater neuronal importance for evaluation tasks of the same type. The authors interpret this persistence as further evidence that the relevant task-associated neurons are present in the pre-trained model rather than created by prompt tuning. The paper reports the pattern across the tasks but does not provide a single aggregate numerical effect size.

  7. Knowl 7 — Skill-neuron pruning retains task performance with lower model size and faster CPU inference

    empirical result

    For each task, the pruning experiment kept the top 2% of skill neurons active and fixed the remaining 98% of neuron activations at their baseline values, folding those fixed contributions into bias terms. Applied to the top nine RoBERTaBASE layers, this procedure reduced the model to 66.6% of its original parameter count. The table compares ordinary prompt tuning with prompt tuning on the pruned model; speedup is measured on a single CPU. The pruned model remained broadly comparable in accuracy and achieved speedups of 1.32–1.38, approximately 1.4 overall.

    Task Prompt tuning accuracy (%) Pruned-model accuracy (%) Speedup
    SST-2 91.8±0.591.8\pm0.5 89.3±2.089.3\pm2.0 1.34
    IMDB 91.6±0.591.6\pm0.5 87.6±3.087.6\pm3.0 1.34
    Tweet 70.0±0.270.0\pm0.2 69.0±0.969.0\pm0.9 1.34
    MNLI 76.8±1.876.8\pm1.8 70.0±1.170.0\pm1.1 1.38
    QNLI 85.7±0.785.7\pm0.7 81.0±1.081.0\pm1.0 1.36
    AG News 98.8±0.198.8\pm0.1 99.8±0.199.8\pm0.1 1.32
    DBpedia 99.7±0.199.7\pm0.1 99.0±0.199.0\pm0.1 1.33

    Accuracy entries are means and standard deviations over five trials.

  8. Knowl 8 — Task prediction is not explained by obvious word selectivity or label-word choice

    empirical result

    To assess whether skill neurons simply respond to task-related keywords, the authors inspected the words associated with top neurons through embedding-to-weight cosine similarity and average activation. In the reported examples, these associated words did not provide evident clues to the task labels, leading the authors to argue that the detected neurons are not merely keyword-selective. They also repeated the analysis with different randomly selected label words for prompt tuning. Across the seven tasks, neuron-predictivity rankings obtained using five random label-word choices had an average Spearman correlation of 0.87, indicating that the rankings were largely consistent across label-word choices.

  9. Knowl 9 — Restricting the activated-neuron transfer metric to skill neurons improves its correlation

    empirical result

    The authors modified the overlapping rate of activated neurons (ON), a metric used to estimate prompt transferability between tasks. Instead of calculating overlap over all neurons, they calculated it using only the top 20% of skill neurons for the target tasks, to reduce the influence of neurons they considered redundant or lacking task-specific skills. Across the investigated tasks, this change raised the average Spearman correlation between ON and prompt transferability from 0.53 to 0.71.

  10. Knowl 10 — Evidence is limited to FFN neurons in one English-language model and classification tasks

    limitation

    The experiments were conducted on RoBERTaBASE alone, so the paper does not establish whether the reported skill-neuron phenomena generalize to other pre-trained Transformer models. All evaluated datasets were English classification tasks; non-English data and non-classification tasks were not tested. The analysis also examined FFN neurons rather than other model components such as attention heads. The findings and applications therefore remain preliminary within these boundaries.

Coverage note — No substantial contributed finding was omitted. Fine-grained dataset split counts and implementation details beyond the prompt-tuning configuration were left out because they do not materially change the reported method or conclusions.

References

  1. 1.Pulkit Agrawal, Ross B. Girshick, and Jitendra Malik. 2014. Analyzing the performance of multilayer neural networks for object recognition. In Proceedings of ECCV, pages 329–344.
  2. 2.Omer Antverg and Yonatan Belinkov. 2022. On the pitfalls of analyzing individual neurons in language models. In Proceedings of ICLR.
  3. 3.Sajid Anwar, Kyuyeon Hwang, and Wonyong Sung. 2017. Structured pruning of deep convolutional neural networks. ACM Journal on Emerging Technologies in Computing Systems (JETC), 13(3):1–18.
  4. 4.Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang. 2018. Stronger generalization bounds for deep nets via a compression approach. In Proceedings of ICML, pages 254–263.
  5. 5.Sören Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary Ives. 2007. DBpedia: A nucleus for a web of open data. In Proceedings of ISWC/ASWC, pages 722–735.
  6. 6.Francesco Barbieri, Jose Camacho-Collados, Luis Espinosa Anke, and Leonardo Neves. 2020. TweetEval: Unified benchmark and comparative evaluation for tweet classification. In Findings of EMNLP, pages 1644–1650.
  7. 7.Horace B Barlow. 1972. Single units and sensation: A neuron doctrine for perceptual psychology? Perception, 1(4):371–394.
  8. 8.Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2018. Identifying and controlling important neurons in neural machine translation. In Proceedings of ICLR.
  9. 9.David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba. 2017. Network dissection: Quantifying interpretability of deep visual representations. Proceedings of CVPR, pages 3319–3327.
  10. 10.David Bau, Jun-Yan Zhu, Hendrik Strobelt, Agata Lapedriza, Bolei Zhou, and Antonio Torralba. 2020. Understanding the role of individual units in a deep neural network. Proceedings of the National Academy of Sciences, 117(48):30071–30078.
  11. 11.Elad Ben-Zaken, Shauli Ravfogel, and Yoav Goldberg. 2022. BitFit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. In Proceedings of ACL.
  12. 12.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Proceedings of NeurIPS, pages 1877–1901.
  13. 13.Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019. What does BERT look at? an analysis of BERT’s attention. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 276–286.
  14. 14.Adam Coates, Andrej Karpathy, and A. Ng. 2012. Emergence of object-selective features in unsupervised feature learning. In Proceedings of NeurIPS, pages 2681–2689.
  15. 15.Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, and Furu Wei. 2021. Knowledge neurons in pretrained transformers. arXiv preprint, arXiv:2104.08696.
  16. 16.Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, Anthony Bau, and James Glass. 2019. What is one grain of sand in the desert? analyzing individual neurons in deep nlp models. In Proceedings of AAAI, pages 6309–6317.
  17. 17.Fahim Dalvi, Hassan Sajjad, Nadir Durrani, and Yonatan Belinkov. 2020. Analyzing redundancy in pretrained transformer models. In Proceedings of EMNLP, pages 4908–4926.
  18. 18.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT, pages 4171–4186.
  19. 19.Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas. 2021. Attention is not all you need: pure attention loses rank doubly exponentially with depth. In Proceedings of ICML, pages 2793–2803.
  20. 20.Nadir Durrani, Hassan Sajjad, Fahim Dalvi, and Yonatan Belinkov. 2020. Analyzing individual neurons in pre-trained language models. In Proceedings of EMNLP, pages 4865–4880.
  21. 21.Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021. Transformer feed-forward layers are key-value memories. In Proceedings of EMNLP, pages 5484–5495.
  22. 22.Mitchell A Gordon, Kevin Duh, and Nicholas Andrews. 2020. Compressing bert: Studying the effects of weight pruning on transfer learning. arXiv preprint arXiv:2002.08307.
  23. 23.Song Han, Jeff Pool, John Tran, and William Dally. 2015. Learning both weights and connections for efficient neural network. In Proceedings of NeurIPS, pages 1135–1143.
  24. 24.Xu Han, Zhengyan Zhang, Ning Ding, Yuxian Gu, Xiao Liu, Yuqi Huo, Jiezhong Qiu, Liang Zhang, Wentao Han, Minlie Huang, et al. 2021. Pre-trained models: Past, present and future. AI Open, pages 225–250.
  25. 25.Lucas Torroba Hennigen, Adina Williams, and Ryan Cotterell. 2020. Intrinsic probing through dimension selection. In Proceedings of EMNLP, pages 197–216.
  26. 26.John Hewitt and Christopher D Manning. 2019. A structural probe for finding syntax in word representations. In Proceedings of NACCL-HLT, pages 4129–4138.
  27. 27.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In Proceedings of ICML, pages 2790–2799.
  28. 28.Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. 2020. How can we know what language models know? Transactions of the Association for Computational Linguistics, 8:423–438.
  29. 29.Andrej Karpathy, Justin Johnson, and Li Fei-Fei. 2015. Visualizing and understanding recurrent networks. arXiv preprint arXiv:1506.02078, pages 818–833.
  30. 30.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Proceedings of ICLR.
  31. 31.Quoc V. Le, Marc’Aurelio Ranzato, Rajat Monga, Matthieu Devin, Gregory S. Corrado, Kai Chen, Jeffrey Dean, and A. Ng. 2013. Building high-level features using large scale unsupervised learning. In Proceedings of ICASSP, pages 8595–8598.
  32. 32.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In Proceedings of EMNLP, pages 3045–3059.
  33. 33.Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf. 2021. Datasets: A community library for natural language processing. In Proceedings of EMNLP, pages 175–184.
  34. 34.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of ACL, pages 4582–4597.
  35. 35.Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019a. Linguistic knowledge and transferability of contextual representations. In Proceedings of NAACL-HLT, pages 1073–1094.
  36. 36.Xiao Liu, Kaixuan Ji, Yicheng Fu, Zhengxiao Du, Zhilin Yang, and Jie Tang. 2022. P-Tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks. In Proceedings of ACL.
  37. 37.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b. RoBERTa: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907,11692.
  38. 38.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of ACL-HLT, pages 142–150.
  39. 39.Eran Malach, Gilad Yehudai, Shai Shalev-shwartz, and Ohad Shamir. 2020. Proving the lottery ticket hypothesis: Pruning is all you need. In Proceedings of ICML.
  40. 40.Paul Michel, Omer Levy, and Graham Neubig. 2019. Are sixteen heads really better than one? In Proceedings of NeurIPS, pages 14014–14024.
  41. 41.Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning. 2021. Fast model editing at scale. In Proceedings of ICLR.
  42. 42.Sparsh Mittal, Poonam Rajput, and Sreenivas Subramoney. 2021. A survey of deep learning on CPUs: opportunities and co-optimizations. IEEE Transactions on Neural Networks and Learning Systems, pages 1–21.
  43. 43.Ari S Morcos, David GT Barrett, Neil C Rabinowitz, and Matthew Botvinick. 2018. On the importance of single directions for generalization. In Proceedings of ICLR.
  44. 44.Jesse Mu and Jacob Andreas. 2020. Compositional explanations of neurons. In Proceedings of NeurIPS, pages 17153–17163.
  45. 45.Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In Proceedings of EMNLP-IJCNLP, pages 2463–2473.
  46. 46.Ofir Press, Noah A. Smith, and Omer Levy. 2020. Improving transformer models by reordering their sublayers. In Proceedings of ACL, pages 2996–3005.
  47. 47.Guanghui Qin and Jason Eisner. 2021. Learning how to ask: Querying LMs with mixtures of soft prompts. In Proceedings of NAACL-HLT, pages 5203–5212.
  48. 48.R Quian Quiroga, Leila Reddy, Gabriel Kreiman, Christof Koch, and Itzhak Fried. 2005. Invariant visual representation by single neurons in the human brain. Nature, 435(7045):1102–1107.
  49. 49.Alec Radford, Rafal Józefowicz, and Ilya Sutskever. 2017. Learning to generate reviews and discovering sentiment. arXiv preprint arXiv:1704.01444.
  50. 50.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67.
  51. 51.Sara Rosenthal, Noura Farra, and Preslav Nakov. 2017. SemEval-2017 task 4: Sentiment analysis in Twitter. In Proceedings of SemEval, pages 502–518.
  52. 52.Bernardo Rudy, Gordon Fishell, SooHyun Lee, and Jens Hjerling-Leffler. 2011. Three groups of interneurons account for nearly 100% of neocortical gabaergic neurons. Developmental neurobiology, 71(1):45–61.
  53. 53.Timo Schick and Hinrich Schütze. 2021. Exploiting cloze-questions for few-shot text classification and natural language inference. In Proceedings of EACL, pages 255–269.
  54. 54.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of EMNLP, pages 1631–1642.
  55. 55.Charles Spearman. 1987. The proof and measurement of association between two things. In Proceedings of AJP, 3/4, pages 441–471.
  56. 56.Yusheng Su, Xiaozhi Wang, Yujia Qin, Chi-Min Chan, Yankai Lin, Zhiyuan Liu, Peng Li, Juanzi Li, Lei Hou, Maosong Sun, et al. 2021. On transferability of prompt tuning for natural language understanding. arXiv preprint arXiv:2111.06719.
  57. 57.Xavier Suau, Luca Zappella, and Nicholas Apostoloff. 2020. Finding experts in transformer models. arXiv preprint arXiv:2005.07647.
  58. 58.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of NeurIPS, pages 5998–6008.
  59. 59.Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. In Proceedings of NAACL, pages 5797–5808.
  60. 60.Tu Vu, Brian Lester, Noah Constant, Rami Al-Rfou, and Daniel Cer. 2021. Spot: Better frozen model adaptation through soft prompt transfer. arXiv preprint arxiv:2110.07904.
  61. 61.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of ICLR.
  62. 62.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of NAACL-HLT, pages 1112–1122.
  63. 63.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of EMNLP, pages 38–45.
  64. 64.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019. Xlnet: Generalized autoregressive pretraining for language understanding. In Proceedings of NeurIPS, pages 5754–5764.
  65. 65.Matthew D. Zeiler and Rob Fergus. 2014. Visualizing and understanding convolutional networks. In Proceedings of ECCV.
  66. 66.Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In Proceedings of NeurIPS, pages 649–657.
  67. 67.Zhengyan Zhang, Yankai Lin, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou. 2021. Moefication: Conditional computation of transformer models for efficient inference. arXiv preprint arXiv:2110.01786.
  68. 68.Zexuan Zhong, Dan Friedman, and Danqi Chen. 2021. Factual probing is [MASK]: Learning vs. learning to recall. In Proceedings of NAACL-HLT, pages 5017–5033.
  69. 69.Bolei Zhou, Aditya Khosla, Àgata Lapedriza, Aude Oliva, and Antonio Torralba. 2015. Object detectors emerge in deep scene cnns. In Proceedings of ICLR.

Citation

MLA
Wang, X., et al. “Finding Skill Neurons in Pre-trained Transformer-based Language Models”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 11132–52, https://doi.org/10.18653/v1/2022.emnlp-main.765.
APA
Wang, X., Wen, K., Zhang, Z., Hou, L., Liu, Z., & Li, J. (2022). Finding Skill Neurons in Pre-trained Transformer-based Language Models. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 11132–11152. https://doi.org/10.18653/v1/2022.emnlp-main.765
Chicago
Wang, X., K. Wen, Z. Zhang, L. Hou, Z. Liu, and J. Li. 2022. “Finding Skill Neurons in Pre-trained Transformer-based Language Models”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 11132–52. https://doi.org/10.18653/v1/2022.emnlp-main.765.
Harvard
Wang, X. et al. (2022) “Finding Skill Neurons in Pre-trained Transformer-based Language Models”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 11132–11152. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.765.
Vancouver
1. Wang X, Wen K, Zhang Z, Hou L, Liu Z, Li J (2022) Finding Skill Neurons in Pre-trained Transformer-based Language Models. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 11132–11152

BibTeX

@inproceedings{wang-etal-2022-finding-skill,
    title = "Finding Skill Neurons in Pre-trained Transformer-based Language Models",
    author = "Wang, Xiaozhi  and
      Wen, Kaiyue  and
      Zhang, Zhengyan  and
      Hou, Lei  and
      Liu, Zhiyuan  and
      Li, Juanzi",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.765/",
    doi = "10.18653/v1/2022.emnlp-main.765",
    pages = "11132--11152"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/