Finding Skill Neurons in Pre-trained Transformer-based Language Models
Xiaozhi WangKaiyue WenZhengyan ZhangLei HouZhiyuan LiuJuanzi Li
Reveals that pre-trained Transformers contain task-predictive "skill neurons" formed during pre-training, providing a concrete method to identify them via prompt tuning and apply them to model pruning and transferability estimation.
Transformer-based language models have become the foundation for modern natural language processing, yet how these large systems store and execute specific skills internally remains a major black-box challenge. Understanding where task-specific capabilities reside is essential for controlling model behavior, diagnosing errors, and reducing the computational cost of running large models.
The article demonstrates that individual components within pre-trained Transformers, specifically individual units called skill neurons located in feed-forward networks, directly encode the high-level capabilities required to solve specific tasks.
To identify these components, the researchers used prompt tuning, a lightweight adaptation method that adds small trainable prompt tokens to inputs while keeping the underlying model frozen. By measuring neuron activations on these prompt tokens across seven English classification benchmarks spanning sentiment analysis, natural language inference, and topic classification, the researchers calculated each neuron's ability to predict correct labels. They evaluated these identified neurons through noise perturbation experiments, cross-task correlation analyses, untuned baseline comparisons, and downstream efficiency applications.
The analysis yielded five primary findings. First, skill neurons emerge consistently across all evaluated tasks, and the top-ranking neurons alone achieve task accuracy comparable to full prompt-tuned models (e.g., 91.6% versus 91.8% on standard sentiment benchmarks). Second, perturbing these specific neurons with random noise causes severe drops in model accuracy, whereas perturbing random neurons has little effect. Third, skill neurons are task-specific and concentrated in the higher layers of the network, with closely related tasks sharing similar skill neuron rankings. Fourth, skill neurons originate during initial pre-training rather than fine-tuning, as they remain critical across distinct tuning methods like adapters and BitFit and can even be detected with untuned prompts. Fifth, applying skill neuron identification to network pruning allows models to deactivate 98% of neurons in upper layers, reducing total model parameters to 66.6% and delivering an approximate 1.4x inference speedup on central processing units with minimal accuracy loss.
These findings indicate that large language models naturally develop modular, task-specific capabilities during pre-training. For organizations deploying language models, this opens direct opportunities to lower computing costs and deployment latency by selectively pruning inactive neurons. It also provides a practical mechanism to measure task similarity and improve transfer learning, as incorporating top skill neurons increased transferability prediction correlations from 0.53 to 0.71.
Organizations evaluating language model efficiency should consider piloting skill-neuron-based structured pruning for latency-sensitive applications. Engineering teams should also leverage skill neuron metrics to identify compatible tasks for cross-task adaptation. Further analysis should explore finding skill neurons without prompt tuning to streamline identification workflows.
The conclusions carry high empirical confidence across the evaluated settings, but leaders should note three boundaries: the experiments evaluated a single base architecture (RoBERTa), focused exclusively on English text, and examined only classification tasks. Additional validation is recommended before applying these techniques to generative workflows, multilingual deployments, or alternative model architectures.
- Paper: Transformer Feed-Forward Layers Are Key-Value Memories, Mor Geva et al. (2020). This study establishes the feed-forward-layer memory mechanism that provides essential context for locating task-specific skills in those same layers.
- Paper: MoEfication: Transformer Feed-forward Layers are Mixtures of Experts, Zhengyan Zhang et al. (2022). Its account of sparse, specialized FFN experts prepares readers to understand the source’s claim that individual FFN neurons support distinct skills.
- Paper: P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks, Xiao Liu et al. (2022). Because the source identifies neurons using prompt tuning, this work introduces the prompt-tuning method and its use for adapting frozen language models.
- Paper: BERT Rediscovers the Classical NLP Pipeline, Ian Tenney et al. (2019). Its layer-wise probing of BERT shows how pretrained Transformers can be analyzed for localized task information, a useful precursor to neuron-level analysis.
- Paper: Task-Specific Skill Localization in Fine-tuned Language Models, Abhishek Panigrahi et al. (2023). This follow-up shifts from skills already present in pretrained neurons to tiny parameter subsets that capture capabilities acquired through fine-tuning.
- Paper: Rethinking the Role of Scale for In-Context Learning: An Interpretability-based Case Study at 66 Billion Scale, Hritik Bansal et al. (2023). It extends the source’s efficiency question from task-specific FFN neurons to pruning attention heads and feed-forward networks in a much larger model.
- Paper: A Fast Post-Training Pruning Framework for Transformers, Woosuk Kwon et al. (2022). Building on structured pruning as a route to efficient inference, this work develops a fast post-training framework for pruning Transformer components under resource constraints.
