Are Sixteen Heads Really Better than One?

Paul MichelOmer LevyGraham Neubig

article2019NeurIPS1,472 citations

Reveals that a majority of attention heads in pretrained Transformer models can be pruned at test time without degrading performance, providing simple greedy algorithms that decrease inference latency and memory consumption.

Listen

Modern state-of-the-art natural language processing systems rely heavily on Transformer models that use multi-headed attention. These models distribute computational focus across numerous parallel attention mechanisms, called heads, to process text. However, standard architectures allocate substantial computational power and memory to these mechanisms—often representing roughly one third of all model parameters—without a clear operational understanding of whether all heads are necessary during deployment.

The article aims to evaluate whether multi-headed attention layers genuinely require multiple heads at test time and demonstrates the practical extent to which redundant heads can be removed to improve computational efficiency without sacrificing accuracy.

To evaluate head necessity, the authors conducted ablation experiments on major Transformer architectures across machine translation and natural language inference tasks, including standard benchmarks such as WMT English-to-French translation and BERT on MultiNLI. The researchers systematically masked individual heads and reduced entire layers down to single heads. They then developed an iterative, sensitivity-based pruning algorithm that ranks head importance using gradient metrics, evaluating the compounding impact of pruning heads across the entire network as well as on out-of-domain datasets.

The investigation produced four central findings. First, the vast majority of individual attention heads are redundant at test time; removing an isolated head rarely impacts performance, and in some cases, removal slightly improves accuracy. Second, many entire layers can be reduced to a single attention head without statistically significant performance loss. Third, applying network-wide iterative pruning allows between 20% and 40% of all heads across the models to be removed—reaching up to 50% to 60% on certain classification tasks—with negligible impact on task performance. Fourth, removing 50% of heads directly delivers operational speed improvements, accelerating BERT inference speed by up to 17.5% at higher batch sizes. However, pruning sensitivity varies considerably across components: encoder-decoder attention in translation models is highly vulnerable to head removal compared to self-attention layers, and over-pruning beyond baseline thresholds causes severe performance degradation.

These findings indicate that current language models carry substantial structural over-parameterization during inference, creating unnecessary operational costs in latency and memory consumption. Practitioners deploying large language models can immediately optimize runtime performance and reduce infrastructure costs by pruning low-impact attention heads rather than serving full-scale models. When implementing pruning, engineering teams must tailor reductions to specific architectural components, ensuring that sensitive layers like encoder-decoder attention are preserved to avoid catastrophic failure.

Decision-makers should consider systematic head pruning as an effective optimization step before deploying Transformer models in latency-critical or resource-constrained production environments. While confidence in these test-time reductions is supported by cross-domain evaluations, the findings rely primarily on specific Transformer variants. Future work should pilot dynamic pruning routines during training and assess whether specialized architectures can be designed with fewer heads from the outset.

  • Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). Vaswani et al. introduce the original multi-head self-attention Transformer architecture whose head redundancy and post-training pruning potential are directly investigated in this work.
  • Paper: Effective Approaches to Attention-based Neural Machine Translation, Minh-Thang Luong et al. (2015). Luong et al. formulate the foundational attention mechanisms for sequence-to-sequence translation models that underpin the attention heads analyzed in this paper.
  • Paper: Rethinking the Value of Network Pruning, Zhuang Liu et al. (2019). Liu et al. critically re-evaluate the purpose of structural pruning and weight initialization, providing key conceptual framing for pruning complex neural network modules.
  • Paper: Pruning Convolutional Neural Networks for Resource Efficient Inference, Pavlo Molchanov et al. (2016). Molchanov et al. develop first-order Taylor expansion criteria for structured pruning that directly inform greedy sensitivity-based component removal algorithms.
  • Paper: Learning both Weights and Connections for Efficient Neural Networks, Song Han et al. (2015). Han et al. establish the classic iterative pruning-and-retraining paradigm for eliminating redundant neural network parameters without sacrificing accuracy.
  • Paper: Optimal Brain Damage, Yann LeCun et al. (1989). LeCun et al. introduce principled, saliency-based network pruning derived from loss sensitivity, which serves as the theoretical precursor to evaluating individual attention head importance.
Cover for Are Sixteen Heads Really Better than One?

Abstract

Attention is a powerful and ubiquitous mechanism for allowing neural models to focus on particular salient pieces of information by taking their weighted average when making predictions. In particular, multi-headed attention is a driving force behind many recent state-of-the-art NLP models such as Transformer-based MT models and BERT. These models apply multiple attention mechanisms in parallel, with each attention "head" potentially focusing on different parts of the input, which makes it possible to express sophisticated functions beyond the simple weighted average. In this paper we make the surprising observation that even if models have been trained using multiple heads, in practice, a large percentage of attention heads can be removed at test time without significantly impacting performance. In fact, some layers can even be reduced to a single head. We further examine greedy algorithms for pruning down models, and the potential speed, memory efficiency, and accuracy improvements obtainable therefrom. Finally, we analyze the results with respect to which parts of the model are more reliant on having multiple heads, and provide precursory evidence that training dynamics play a role in the gains provided by multi-head attention.

Table of Contents

  • 1 Introduction
  • 2 Background: Attention, Multi-headed Attention, and Masking
  • 2.1 Single-headed Attention
  • 2.2 Multi-headed Attention
  • 2.3 Masking Attention Heads
  • 3 Are All Attention Heads Important?
  • 3.1 Experimental Setup
  • 3.2 Ablating One Head
  • 3.3 Ablating All Heads but One
  • 3.4 Are Important Heads the Same Across Datasets?
  • 4 Iterative Pruning of Attention Heads
  • 4.1 Head Importance Score for Pruning
  • 4.2 Effect of Pruning on BLEU/Accuracy
  • 4.3 Effect of Pruning on Efficiency
  • 5 When Are More Heads Important? The Case of Machine Translation
  • 6 Dynamics of Head Importance during Training
  • 7 Related work
  • 8 Conclusion
  • References
  • A Ablating All Heads but One: Additional Experiment.
  • B Additional Pruning Experiments

Knowls

  1. Knowl 1 — Taylor Expansion Gradient Sensitivity Score for Attention Head Importance

    equation

    In multi-head attention models, the importance IhI_h of an attention head hh can be quantified by computing the expected magnitude of the first-order derivative of the loss function L(x)\mathcal{L}(x) with respect to an indicator mask variable ξh∈{0,1}\xi_h \in \{0, 1\} that modulates head hh's output:

    Ih=Ex∼X∣∂L(x)∂ξh∣=Ex∼X∣Atth(x)T∂L(x)∂Atth(x)∣I_h = \mathbb{E}_{x \sim X} \left| \frac{\partial \mathcal{L}(x)}{\partial \xi_h} \right| = \mathbb{E}_{x \sim X} \left| \text{Att}_h(x)^T \frac{\partial \mathcal{L}(x)}{\partial \text{Att}_h(x)} \right|

    where:

    • XX is the data distribution (or a representative sample dataset),
    • xx is an input data example,
    • L(x)\mathcal{L}(x) is the scalar task loss on example xx,
    • Atth(x)∈Rd\text{Att}_h(x) \in \mathbb{R}^d is the vector output of attention head hh for input xx,
    • ξh\xi_h is the multiplicative scalar gate applied to Atth(x)\text{Att}_h(x).

    The absolute value prevents positive and negative gradient contributions across individual samples from cancelling each other out. When ranking heads across different layers of a network, the scores IhI_h are normalized layer-wise using the ℓ2\ell_2 norm.

  2. Knowl 2 — Masked Multi-Head Attention Formulation

    equation

    To evaluate head contributions and prune heads at inference time, multi-head attention (MHA) incorporates a binary scalar mask ξh∈{0,1}\xi_h \in \{0, 1\} for each head h∈{1,…,Nh}h \in \{1, \dots, N_h\}:

    MHAtt(x,q)=∑h=1NhξhAttWkh,Wqh,Wvh,Woh(x,q)\text{MHAtt}(x, q) = \sum_{h=1}^{N_h} \xi_h \text{Att}_{W_k^h, W_q^h, W_v^h, W_o^h}(x, q)

    where:

    • x=(x1,…,xn)∈Rn×dx = (x_1, \dots, x_n) \in \mathbb{R}^{n \times d} is the sequence of nn key/value vectors of dimension dd,
    • q∈Rdq \in \mathbb{R}^d is the query vector,
    • NhN_h is the number of attention heads in the layer,
    • ξh∈{0,1}\xi_h \in \{0, 1\} is a binary mask variable (where ξh=1\xi_h = 1 retains the head and ξh=0\xi_h = 0 masks it out by zeroing its output),
    • Wkh,Wqh,Wvh∈Rdh×dW_k^h, W_q^h, W_v^h \in \mathbb{R}^{d_h \times d} are the head-specific key, query, and value projection matrices with dh=d/Nhd_h = d / N_h,
    • Woh∈Rd×dhW_o^h \in \mathbb{R}^{d \times d_h} is the head-specific output projection matrix,
    • AttWkh,Wqh,Wvh,Woh(x,q)=Woh∑i=1nαihWvhxi\text{Att}_{W_k^h, W_q^h, W_v^h, W_o^h}(x, q) = W_o^h \sum_{i=1}^n \alpha_i^h W_v^h x_i with attention weights αih=softmax(qT(Wqh)TWkhxidh)\alpha_i^h = \text{softmax}\left( \frac{q^T (W_q^h)^T W_k^h x_i}{\sqrt{d_h}} \right).
  3. Knowl 3 — Iterative Head Importance Pruning Algorithm

    algorithm

    An iterative greedy procedure prunes attention heads across all layers of a trained Transformer model using layer-normalized first-order gradient sensitivity scores.

    Input: Trained Transformer model with parameters Θ\Theta and HH total heads across all layers, dataset DD, target pruning fraction P∈(0,1)P \in (0, 1), step size fraction Δp∈(0,1)\Delta p \in (0, 1)
    Output: Pruned Transformer model with (1−P)H(1 - P)H active heads
    Initialize head mask ξh←1\xi_h \leftarrow 1 for all heads $h \in \{1, \dots, H\}
    Calculate current pruning percentage p←0p \leftarrow 0
    while p<Pp < P do
        for each active head hh do
            Compute raw sensitivity: I^h←1∣D∣∑x∈D∣Atth(x)T∂L(x)∂Atth(x)∣\hat{I}_h \leftarrow \frac{1}{|D|} \sum_{x \in D} |\text{Att}_h(x)^T \frac{\partial \mathcal{L}(x)}{\partial \text{Att}_h(x)}|
        end for
        
        for each layer ll in the model do
            Let HlH_l be the set of active heads in layer ll
            norml←∑h∈HlI^h2\text{norm}_l \leftarrow \sqrt{\sum_{h \in H_l} \hat{I}_h^2}
            for each head h∈Hlh \in H_l do
                Ih←I^hnormlI_h \leftarrow \frac{\hat{I}_h}{\text{norm}_l}
            end for
        end for
        
        Determine number of heads to prune: k←⌈Δp⋅H⌉k \leftarrow \lceil \Delta p \cdot H \rceil
        Identify the kk active heads across the whole model with the smallest normalized scores IhI_h
        
        for each identified head h∗h^* do
            ξh∗←0\xi_{h^*} \leftarrow 0
        end for
        
        p←p+Δpp \leftarrow p + \Delta p
    end while
    Physically remove weights corresponding to heads where ξh=0\xi_h = 0
    return Pruned Transformer model

    Computing IhI_h requires only a single forward-backward pass over dataset DD, making it computationally efficient while avoiding combinatorial search over subsets of heads.

  4. Knowl 4 — Redundancy and Cross-Domain Consistency of Individual Attention Heads

    empirical result

    Ablating individual attention heads at test time in trained Transformer models reveals widespread redundancy across tasks and domains:

    1. WMT14 English-to-French Translation: In a 6-layer, 16-head Transformer-large model (base score 36.05 BLEU on newstest2013), only 8 out of 96 encoder self-attention heads produce a statistically significant change in BLEU (p<0.01p < 0.01, paired bootstrap resampling) when removed individually. For 4 of those 8 heads, removing the head actually increases the BLEU score.

    2. BERT-base on Multi-Genre NLI (MultiNLI): In a 12-layer, 12-head model, individual ablation of most heads causes negligible changes in matched validation accuracy, with several individual removals slightly increasing accuracy.

    3. Cross-Domain Correlation: Individual head impact generalizes across disparate test distributions. Ablation BLEU deltas on the in-domain WMT newstest2013 dataset correlate positively with deltas on the noisy out-of-domain MTNT test set (Pearson r=0.56r = 0.56, p<0.001p < 0.001). Similarly, BERT head ablation accuracies on the MultiNLI matched validation set correlate with those on the mismatched validation set (Pearson r=0.68r = 0.68, p<0.001p < 0.001), indicating that head importance is consistent across evaluation domains.

  5. Knowl 5 — Layer Reduction to a Single Attention Head at Test Time

    empirical result

    Many individual Transformer layers can be reduced to a single attention head at test time without statistically significant performance degradation:

    1. BERT-base on MultiNLI: When every attention head in a layer except the single best head is removed, zero layers experience a statistically significant reduction in validation accuracy (p<0.01p < 0.01). Even when selecting the best head per layer on a separate 5,000-example training subset, accuracy drops by under 1.0 percentage point across all 12 layers (layer 1: −0.01%-0.01\%, layer 2: −0.02%-0.02\%, layer 3: −0.26%-0.26\%, layer 4: −0.53%-0.53\%, layer 5: −0.29%-0.29\%, layer 6: −0.52%-0.52\%, layer 7: +0.05%+0.05\%, layer 8: −0.72%-0.72\%, layer 9: −0.96%-0.96\%, layer 10: +0.07%+0.07\%, layer 11: −0.19%-0.19\%, layer 12: −0.15%-0.15\%).

    2. WMT14 English-to-French Transformer-Large: When selecting the best head per layer on newstest2013 and evaluating on newstest2014, 50% of the layers show no statistically significant drop in BLEU (p<0.01p < 0.01). Decoder self-attention layers 1 through 6 change by +0.03+0.03, −0.13-0.13, 0.000.00, −0.31-0.31, −0.66-0.66, and −0.03-0.03 BLEU, respectively, demonstrating that self-attention layers can operate with 1/161/16 of their multi-head parameters.

  6. Knowl 6 — Disproportionate Reliance of Encoder-Decoder Attention on Multi-Headedness

    empirical result

    In sequence-to-sequence Transformer models, multi-headedness is far more critical for encoder-decoder (Enc-Dec) cross-attention than for encoder self-attention (Enc-Enc) or decoder self-attention (Dec-Dec):

    1. Single-Layer Reduction Degradation: In the WMT14 English-to-French model, reducing the final encoder-decoder cross-attention layer (layer 6) to its single best head degrades performance by 13.56 BLEU points on newstest2013 and 18.89 BLEU points on newstest2014 (p<0.01p < 0.01). In contrast, reducing layer 6 of the encoder self-attention or decoder self-attention to a single head results in drops of only 0.36 and 0.03 BLEU points, respectively.

    2. Component Pruning Trajectories: When incrementally pruning heads within specific attention types using gradient sensitivity scores, removing more than 60% of encoder-decoder attention heads leads to catastrophic failure (BLEU falls near 0 at 80% pruning). By contrast, encoder self-attention and decoder self-attention layers retain viable translation scores (around 30 BLEU) even when 80% of their heads are removed.

  7. Knowl 7 — Global Attention Head Pruning Capacity Across NLP Tasks

    empirical result

    Iterative gradient-sensitivity pruning across all layers of trained models without retraining achieves high sparsity before incurring steep accuracy drops:

    • WMT14 En-Fr (Transformer-Large): Up to 20% of all attention heads across the entire network can be removed without noticeable impact on BLEU score.
    • BERT-base on MultiNLI: Up to 40% of all attention heads can be removed without noticeable loss in classification accuracy.
    • BERT-base on Additional GLUE Tasks: Up to 60% of heads can be pruned on SST-2 (Stanford Sentiment Treebank), and up to 50% of heads on CoLA (Corpus of Linguistic Acceptability) and MRPC (Microsoft Research Paraphrase Corpus), without significant degradation in task metrics.
    • IWSLT 2014 De-En (Transformer-Small, 6 layers, 8 heads/layer): Up to 40% of heads can be pruned while retaining ≥30\ge 30 BLEU (unpruned baseline is 34.9 BLEU).

    Beyond these thresholds (20–40% in translation, 40–60% in classification), performance drops precipitously toward zero, showing that while models are overparameterized, a core set of multiple heads across layers remains essential.

  8. Knowl 8 — Inference Throughput Gains from Structured Head Pruning

    data/table

    Physically pruning 50% of all attention heads in BERT-base yields measurable inference speedups, especially at higher batch sizes where matrix multiplications are throughput-bound.

    The table below reports the average inference speed of BERT-base on the MultiNLI matched validation set in examples per second (±\pm standard deviation) across 6 runs on GeForce GTX 1080Ti GPUs:

    Model Batch size 1 Batch size 4 Batch size 16 Batch size 64
    Original BERT 17.0±0.317.0 \pm 0.3 67.3±1.367.3 \pm 1.3 114.0±3.6114.0 \pm 3.6 124.7±2.9124.7 \pm 2.9
    Pruned (50% heads) 17.3±0.617.3 \pm 0.6 69.1±1.369.1 \pm 1.3 134.0±3.6134.0 \pm 3.6 146.6±3.4146.6 \pm 3.4
    Speedup +1.9% +2.7% +17.5% +17.5%

    While head pruning provides limited speedup at batch size 1 (+1.9%), throughput increases by +17.5% at batch sizes 16 and 64, driven by the reduction in multi-head attention parameters, which constitute roughly one-third of the total parameter count in the model.

  9. Knowl 9 — Two-Phase Dynamics of Head Importance During Training

    empirical result

    Evaluating attention head pruning across training epochs of an IWSLT 2014 German-to-English Transformer (6 layers, 8 heads per layer) reveals two distinct learning regimes:

    1. Early Uniform Contribution Phase (Epochs 1–2): During the earliest epochs (unpruned BLEU of 3.0 at epoch 1 and 6.2 at epoch 2), relative BLEU degradation scales linearly with the percentage of pruned heads. Pruning heads ranked by IhI_h yields a performance drop identical to arbitrary head pruning, indicating that heads are equally important early in training.

    2. Specialization and Concentration Phase (Epoch 10 Onward): From epoch 10 through convergence at epoch 40 (unpruned BLEU 34.9), a sharp distinction emerges between critical and redundant heads: up to 40% of heads can be pruned while maintaining 85% to 90% of the unpruned BLEU score.

    This indicates that head differentiation and redundancy develop early during training (by epoch 10 out of 40) rather than existing at initialization or emerging only at convergence.

Coverage note — None was omitted; all primary methodological contributions, experimental findings across architectures and tasks, efficiency benchmarks, and training dynamics analyses from the paper are represented.

References

  1. 1.Karim Ahmed, Nitish Shirish Keskar, and Richard Socher. Weighted transformer network for machine translation. arXiv preprint arXiv:1711.02132, 2017.
  2. 2.Sajid Anwar, Kyuyeon Hwang, and Wonyong Sung. Structured pruning of deep convolutional neural networks. J. Emerg. Technol. Comput. Syst., pages 32:1–32:18, 2017.
  3. 3.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. In Proceedings of the International Conference on Learning Representations (ICLR), 2015.
  4. 4.Alexander Binder, Grégoire Montavon, Sebastian Lapuschkin, Klaus-Robert Müller, and Wojciech Samek. Layer-wise relevance propagation for neural networks with local renormalization layers. In International Conference on Artificial Neural Networks, pages 63–71, 2016.
  5. 5.Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, and Marcello Federico. Report on the 11 th iwslt evaluation campaign , iwslt 2014. In Proceedings of the 2014 International Workshop on Spoken Language Translation (IWSLT), 2015.
  6. 6.Jianpeng Cheng, Li Dong, and Mirella Lapata. Long short-term memory-networks for machine reading. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 551–561, 2016.
  7. 7.Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations using rnn encoder–decoder for statistical machine translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1724–1734, 2014.
  8. 8.Zihang Dai, Zhilin Yang, Yiming Yang, William W Cohen, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. Transformer-xl: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860, 2019.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), 2018.
  10. 10.William B. Dolan and Chris Brockett. Automatically constructing a corpus of sentential paraphrases. In Proceedings of the The 3rd International Workshop on Paraphrasing (IWP), 2005. URL http://aclweb.org/anthology/I05-5002.
  11. 11.Song Han, Jeff Pool, John Tran, and William Dally. Learning both weights and connections for efficient neural network. In Proceedings of the 29th Annual Conference on Neural Information Processing Systems (NIPS), pages 1135–1143, 2015.
  12. 12.Babak Hassibi and David G. Stork. Second order derivatives for network pruning: Optimal brain surgeon. In Proceedings of the 5th Annual Conference on Neural Information Processing Systems (NIPS), pages 164–171, 1993.
  13. 13.Yoon Kim and Alexander M. Rush. Sequence-level knowledge distillation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1317–1327, 2016.
  14. 14.Philipp Koehn. Statistical significance tests for machine translation evaluation. In Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 388–395, 2004.
  15. 15.Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondˇrej Bojar, Alexandra Constantin, and Evan Herbst. Moses: Open source toolkit for statistical machine translation. In Proceedings of the 45th Annual Meeting of the Association for Computational Linguistics (ACL), pages 177–180, 2007.
  16. 16.Yann LeCun, John S. Denker, and Sara A. Solla. Optimal brain damage. In Proceedings of the 2nd Annual Conference on Neural Information Processing Systems (NIPS), pages 598–605, 1990.
  17. 17.Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf. Pruning filters for efficient convnets. arXiv preprint arXiv:1608.08710, 2016.
  18. 18.Thang Luong, Hieu Pham, and Christopher D. Manning. Effective approaches to attention-based neural machine translation. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1412–1421, 2015.
  19. 19.Paul Michel and Graham Neubig. MTNT: A testbed for machine translation of noisy text. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 543–553, 2018.
  20. 20.Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz. Pruning convolutional neural networks for resource efficient inference. In Proceedings of the International Conference on Learning Representations (ICLR), 2017.
  21. 21.Kenton Murray and David Chiang. Auto-sizing neural networks: With applications to n-gram language models. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 908–916, 2015.
  22. 22.Graham Neubig, Zi-Yi Dou, Junjie Hu, Paul Michel, Danish Pruthi, and Xinyi Wang. compare-mt: A tool for holistic comparison of language generation systems. In Meeting of the North American Chapter of the Association for Computational Linguistics (NAACL) Demo Track, Minneapolis, USA, June 2019. URL http://arxiv.org/abs/1903.07926.
  23. 23.Myle Ott, Sergey Edunov, David Grangier, and Michael Auli. Scaling neural machine translation. In Proceedings of the 3rd Conference on Machine Translation (WMT), pages 1–9, 2018.
  24. 24.Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit. A decomposable attention model for natural language inference. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2249–2255, 2016.
  25. 25.Romain Paulus, Caiming Xiong, and Richard Socher. A deep reinforced model for abstractive summarization. In Proceedings of the International Conference on Learning Representations (ICLR), 2017.
  26. 26.Alec Radford, Karthik Narasimhan, Time Salimans, and Ilya Sutskever. Improving language understanding with unsupervised learning. Technical report, Technical report, OpenAI, 2018.
  27. 27.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. OpenAI Blog, 1:8, 2019.
  28. 28.Alessandro Raganato and Jörg Tiedemann. An analysis of encoder representations in transformer-based machine translation. In Proceedings of the Workshop on BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 287–297, 2018.
  29. 29.Abigail See, Minh-Thang Luong, and Christopher D. Manning. Compression of neural machine translation models via pruning. In Proceedings of the Computational Natural Language Learning (CoNLL), pages 291–301, 2016.
  30. 30.Tao Shen, Tianyi Zhou, Guodong Long, Jing Jiang, Shirui Pan, and Chengqi Zhang. Disan: Directional self-attention network for rnn/cnn-free language understanding. In Proceedings of the 32nd Meeting of the Association for Advancement of Artificial Intelligence (AAAI), 2018.
  31. 31.Ravid Shwartz-Ziv and Naftali Tishby. Opening the black box of deep neural networks via information. arXiv preprint arXiv:1703.00810, 2017.
  32. 32.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1631–1642, 2013.
  33. 33.Emma Strubell, Patrick Verga, Daniel Andor, David Weiss, and Andrew McCallum. Linguistically-informed self-attention for semantic role labeling. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5027–5038, 2018.
  34. 34.Gongbo Tang, Mathias Müller, Annette Rios, and Rico Sennrich. Why self-attention? a targeted evaluation of neural machine translation architectures. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4263–4272, 2018.
  35. 35.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Proceedings of the 30th Annual Conference on Neural Information Processing Systems (NIPS), pages 5998–6008, 2017.
  36. 36.Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Titov Ivan. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), page to appear, 2019.
  37. 37.Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. Neural network acceptability judgments. arXiv preprint arXiv:1805.12471, 2018.
  38. 38.Adina Williams, Nikita Nangia, and Samuel Bowman. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), pages 1112–1122, 2018.

Citation

MLA
Michel, P., et al. “Are Sixteen Heads Really Better Than One?”. arXiv, 2019, http://arxiv.org/abs/1905.10650v3.
APA
Michel, P., Levy, O., & Neubig, G. (2019). Are Sixteen Heads Really Better than One?. arXiv. http://arxiv.org/abs/1905.10650v3
Chicago
Michel, P., O. Levy, and G. Neubig. 2019. “Are Sixteen Heads Really Better Than One?”. arXiv. http://arxiv.org/abs/1905.10650v3.
Harvard
Michel, P., Levy, O. and Neubig, G. (2019) “Are Sixteen Heads Really Better than One?”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1905.10650v3.
Vancouver
1. Michel P, Levy O, Neubig G (2019) Are Sixteen Heads Really Better than One?. arXiv

BibTeX

@article{michel2019are,
  title = {Are Sixteen Heads Really Better than One?},
  author = {Michel, Paul and Levy, Omer and Neubig, Graham},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1905.10650v3},
  eprint = {1905.10650}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors