A Fast Post-Training Pruning Framework for Transformers

Woosuk KwonSehoon KimMichael W. MahoneyJoseph HassounKurt KeutzerAmir Gholami

article2022NeurIPS246 citations

Proposes a retraining-free structured pruning framework for Transformers that achieves up to a 2.0x reduction in FLOPs with under 1% accuracy loss in less than three minutes on a single GPU.

Listen

Modern natural language processing relies heavily on Transformer neural networks, but their large computational size and high operational latency make practical deployment challenging and costly. Conventional structured pruning approaches reduce inference overhead by removing redundant network components, but they typically require extensive model retraining on full datasets. This retraining process can extend development timelines by hours or days, demands complex engineering and hyperparameter tuning, and often fails to adapt directly to specific hardware and operational constraints.

To overcome these hurdles, the article introduces a fast, post-training structured pruning framework for Transformers that eliminates the need for retraining. The framework takes a fine-tuned Transformer model, a small sample calibration dataset, and a target resource budget (computational operations or target hardware latency), automatically generating a compact, deployment-ready architecture within minutes.

The framework accomplishes this through a three-stage mathematical pipeline that treats the base model weights as fixed and optimizes only lightweight scaling masks. First, a lightweight mask search evaluates the importance of attention heads and feed-forward filters using Fisher information, identifying an optimal initial subset to prune under the specified computational constraint. Second, a mask rearrangement stage captures interactions between components within the same layer to refine pruning selections. Third, a mask tuning step optimizes the remaining nonzero components via linear least squares to reconstruct the original layer-by-layer output signals. The approach was evaluated on popular Transformer architectures (BERT-Base and DistilBERT) across standard language understanding and question-answering benchmarks (GLUE and SQuAD) using an NVIDIA V100 GPU.

The experimental findings show that the proposed framework delivers substantial efficiency gains without degrading task performance. Across all evaluated tasks, the method achieves a 30% to 50% reduction in floating-point operations while keeping model accuracy loss below 1%. On physical hardware, this computational reduction translates to real inference speedups of up to 1.56 times (averaging 1.47 times across tasks at batch size 256). Furthermore, the entire end-to-end pruning process finishes in roughly 39 to 135 seconds on a single GPU—two to three orders of magnitude faster than conventional methods that require 5 to 33 hours of retraining—while matching or exceeding their accuracy.

These results demonstrate that extensive retraining is unnecessary for moderate compression levels in Transformers. By cutting pruning workflows from days to minutes and removing manual hyperparameter tuning, this framework significantly reduces development costs, cloud compute expenses, and time-to-market for enterprise artificial intelligence deployments. It effectively bridges the gap between theoretical model compression and standard automated software deployment pipelines.

Organizations seeking to lower operational latency and inference costs should consider adopting post-training structured pruning for moderate compression needs (up to 50% compute reduction). Because the framework supports target-hardware latency tables, engineering teams should calibrate pruning against their specific production hardware to maximize utilization and avoid hardware-underutilization thresholds. If extreme compression beyond 50% is required, teams should combine the initial pruning framework with light fine-tuning. Future work should explore extending these retraining-free principles to generative language models and vision-based Transformer architectures.

arXiv: 2204.09656
Cover for A Fast Post-Training Pruning Framework for Transformers

Abstract

Pruning is an effective way to reduce the huge inference cost of Transformer models. However, prior work on pruning Transformers requires retraining the models. This can add high training cost and high complexity to model deployment, making it difficult to use in many practical situations. To address this, we propose a fast post-training pruning framework for Transformers that does not require any retraining. Given a resource constraint and a sample dataset, our framework automatically prunes the Transformer model using structured sparsity methods. To retain high accuracy without retraining, we introduce three novel techniques: (i) a lightweight mask search algorithm that finds which heads and filters to prune based on the Fisher information; (ii) mask rearrangement that complements the search algorithm; and (iii) mask tuning that reconstructs the output activations for each layer. We apply our method to BERT_BASE and DistilBERT, and we evaluate its effectiveness on GLUE and SQuAD benchmarks. Our framework achieves up to 2.0× reduction in FLOPs and 1.56× speedup in inference latency, while maintaining < 1% loss in accuracy. Importantly, our framework prunes Transformers in less than 3 minutes on a single GPU, which is over two orders of magnitude faster than existing pruning approaches that retrain the models.1

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Overview
  • 3.1 Background
  • 3.2 Framework Overview
  • 4 Methodology
  • 4.1 Fisher-based Mask Search
  • 4.2 Fisher-based Mask Rearrangement
  • 4.3 Mask Tuning
  • 5 Evaluation
  • 5.1 Experimental Setup
  • 5.2 Performance Evaluation
  • 5.3 Comparison with the Prior Methods
  • 5.4 Discussion
  • 6 Conclusion
  • Acknowledgments and Disclosure of Funding
  • References

Knowls

  1. Knowl 1 — Retraining-free constrained pruning framework

    model/method

    The paper introduces a post-training pruning framework that takes a fine-tuned Transformer, a small task-specific sample dataset, and either a FLOPs constraint or a target-hardware latency constraint as input. It outputs a smaller dense Transformer without retraining or user-guided hyperparameter search. The sample dataset normally contains 1–2K training examples, and latency constraints require a lookup table measured on the target hardware.

    The framework has three stages: Fisher-based binary mask search, Fisher-based mask rearrangement, and real-valued mask tuning. It prunes attention heads in multi-head attention and filters in feed-forward networks, while leaving embeddings and the final classifier unchanged because their contribution to inference latency is negligible. The resulting architecture can therefore run on standard hardware without specialized sparse-kernel support. The workflow diagram on page 4 depicts the progression from all-one masks to binary searched masks, rearranged masks, and finally tuned nonzero mask values.

  2. Knowl 2 — Structured head and filter mask parameterization

    model/method

    Consider an encoder Transformer with LL layers, HH attention heads per multi-head attention layer, and NN feed-forward filters per layer. For layer ll, the framework associates a head mask mlMHA∈RHm_l^{\mathrm{MHA}}\in\mathbb{R}^{H} and a filter mask mlFFN∈RNm_l^{\mathrm{FFN}}\in\mathbb{R}^{N} with the layer outputs. The masked operations are

    MHA⁡(x;mlMHA)=∑i=1Hml,iMHAAtt⁡i(x),\operatorname{MHA}(x;m_l^{\mathrm{MHA}})=\sum_{i=1}^{H}m_{l,i}^{\mathrm{MHA}}\operatorname{Att}_i(x), FFN⁡(x;mlFFN)=∑i=1Nml,iFFNW:,i(2) ϕ ⁣(Wi,:(1)x+bi(1))+b(2),\operatorname{FFN}(x;m_l^{\mathrm{FFN}})=\sum_{i=1}^{N}m_{l,i}^{\mathrm{FFN}}W^{(2)}_{:,i}\,\phi\!\left(W^{(1)}_{i,:}x+b_i^{(1)}\right)+b^{(2)},

    where xx is the layer input, Att⁡i\operatorname{Att}_i is attention head ii, W(1),W(2),b(1),b(2)W^{(1)},W^{(2)},b^{(1)},b^{(2)} are feed-forward parameters, and ϕ\phi is the activation function, typically GELU. All masks are initialized to 11, so the initial masked model is identical to the original model. Setting a mask entry to 00 removes the associated head or filter. The masks across all layers can be flattened into a vector m∈RL(H+N)m\in\mathbb{R}^{L(H+N)}; the first two stages restrict entries to 00 or 11, while the tuning stage permits nonzero entries to take real values.

  3. Knowl 3 — Fisher-based approximation to constrained pruning

    equation

    Let m∈Rqm\in\mathbb{R}^{q} be the flattened head-and-filter mask, where q=L(H+N)q=L(H+N), let L(m)\mathcal{L}(m) be the task loss of the fixed Transformer with mask mm, let Cost⁡(m)\operatorname{Cost}(m) be its FLOPs or latency, and let CC be the user-specified cost limit. The pruning problem is

    m⋆∈arg⁡min⁡mL(m)subject toCost⁡(m)≤C.m^\star\in\arg\min_m \mathcal{L}(m)\quad\text{subject to}\quad \operatorname{Cost}(m)\le C.

    The framework approximates the loss around the all-one mask 1\mathbf{1} by

    L(m)≈L(1)+12(1−m)TI(1−m),\mathcal{L}(m)\approx \mathcal{L}(\mathbf{1})+\frac{1}{2}(\mathbf{1}-m)^\mathsf{T}\mathcal{I}(\mathbf{1}-m),

    where the model is assumed to be near a local loss minimum, so the first-order gradient term is treated as negligible, and I\mathcal{I} is the empirical Fisher information matrix computed on a sample dataset D\mathcal{D}:

    I=1∣D∣∑(x,y)∈D(∂ℓ(x,y;1)∂m)(∂ℓ(x,y;1)∂m)T,\mathcal{I}=\frac{1}{|\mathcal{D}|}\sum_{(x,y)\in\mathcal{D}}\left(\frac{\partial\ell(x,y;\mathbf{1})}{\partial m}\right)\left(\frac{\partial\ell(x,y;\mathbf{1})}{\partial m}\right)^\mathsf{T},

    with (x,y)(x,y) an input-label pair and ℓ\ell its per-example loss. Under the diagonal approximation I≈diag⁡(I11,…,Iqq)\mathcal{I}\approx\operatorname{diag}(\mathcal{I}_{11},\ldots,\mathcal{I}_{qq}) and binary masks mi∈{0,1}m_i\in\{0,1\}, selecting a mask reduces to minimizing the sum of diagonal Fisher scores of the pruned units:

    min⁡mi∈{0,1}∑i:mi=0Iii.\min_{m_i\in\{0,1\}}\sum_{i:m_i=0}\mathcal{I}_{ii}.

    Thus, Iii\mathcal{I}_{ii} is used as the importance score of the head or filter associated with mask entry ii.

  4. Knowl 4 — Optimal polynomial-time search under a FLOPs constraint

    algorithm

    For FLOPs-constrained pruning, suppose every attention head costs the same FheadF_{\mathrm{head}} FLOPs and every feed-forward filter costs the same FfilterF_{\mathrm{filter}} FLOPs. Given a diagonal Fisher matrix I\mathcal{I} and a total FLOPs limit CC, the search enumerates the number nn of retained heads, computes the largest feasible number of retained filters, and chooses the lowest-importance units to remove.

    Input: FLOPs limit CC, diagonal Fisher scores Iii\mathcal{I}_{ii}, layer counts L,H,NL,H,N, and per-unit costs Fhead,FfilterF_{\mathrm{head}},F_{\mathrm{filter}}
    Output: Binary head masks mMHAm^{\mathrm{MHA}} and filter masks mFFNm^{\mathrm{FFN}}
    For each feasible number nn of retained heads from 00 to LHLH:
        Set kH=LH−nk_H=LH-n and select the kHk_H heads with the smallest Fisher scores
        Set f=min⁡(LN,⌊(C−nFhead)/Ffilter⌋)f=\min\left(LN,\left\lfloor(C-nF_{\mathrm{head}})/F_{\mathrm{filter}}\right\rfloor\right)
        Set kF=LN−fk_F=LN-f and select the kFk_F filters with the smallest Fisher scores
        Compute the candidate score as the sum of Fisher scores of the selected heads and filters
    Choose the candidate with the smallest score
    Initialize every head and filter mask entry to 11
    Set the selected head and filter entries to 00
    Return the resulting binary masks

    The procedure is optimal for the diagonal-Fisher objective under the stated equal-cost assumptions: for any feasible binary mask mm, its selected mask m⋆m^\star satisfies

    ∑i:mi⋆=0Iii≤∑i:mi=0Iii.\sum_{i:m_i^\star=0}\mathcal{I}_{ii}\leq\sum_{i:m_i=0}\mathcal{I}_{ii}.

    The result follows from the nonnegativity of Fisher diagonal entries and the fact that, for a fixed number of heads or filters to prune, choosing the lowest-scoring units is optimal. The enumeration has LH+1LH+1 head-count cases and avoids a general knapsack search.

  5. Knowl 5 — Piecewise-linear latency-aware cost model

    model/method

    Actual hardware latency is not generally linear in the number of retained heads or filters because parallel hardware can be underutilized after sufficient pruning. For a layer with rr retained heads or filters, the framework approximates lookup-table latency with

    LAT⁡(r)={0,r=0,c,0<r≤T,a(r−T)+c,r>T,\operatorname{LAT}(r)= \begin{cases} 0, & r=0,\\ c, & 0<r\le T,\\ a(r-T)+c, & r>T, \end{cases}

    where c≥0c\ge 0 is the fixed layer overhead, TT is the retention threshold below which the overhead dominates, and a≥0a\ge 0 is the slope above the threshold. The parameters aa, cc, and TT are fitted to measured target-hardware latency by minimizing mean squared error.

    For a Transformer with layer-level retained-unit counts rlMHAr_l^{\mathrm{MHA}} and rlFFNr_l^{\mathrm{FFN}}, the latency constraint is approximated by

    ∑l=1LLAT⁡(rlMHA)+∑l=1LLAT⁡(rlFFN)≤C.\sum_{l=1}^{L}\operatorname{LAT}(r_l^{\mathrm{MHA}})+\sum_{l=1}^{L}\operatorname{LAT}(r_l^{\mathrm{FFN}})\le C.

    The latency search separates the constant-overhead part from the linear part and then applies the FLOPs-style Fisher search to the linear component. The page-7 latency schematic illustrates why pruning below the threshold may reduce the model size without producing proportional measured speedup.

  6. Knowl 6 — Fisher-based mask rearrangement captures within-layer interactions

    model/method

    The diagonal Fisher approximation treats mask variables independently, although two heads or filters in the same layer can interact: pruning one may be harmless while pruning both substantially damages the output. The framework therefore uses a block-diagonal Fisher approximation in which each multi-head attention layer and each feed-forward layer forms one block Il\mathcal{I}_l; interactions across different layers are ignored.

    Let ml⋆m_l^\star be the binary mask obtained by the initial search. The rearrangement stage fixes the number of retained heads or filters in every layer, so ∥ml∥0=∥ml⋆∥0\|m_l\|_0=\|m_l^\star\|_0, and approximately solves the layer-wise problem

    m^l∈arg⁡min⁡ml(1−ml)TIl(1−ml).\widehat{m}_l\in\arg\min_{m_l}(\mathbf{1}-m_l)^\mathsf{T}\mathcal{I}_l(\mathbf{1}-m_l).

    It warm-starts from ml⋆m_l^\star. During greedy exchanges, a currently pruned head or filter with high Fisher importance is considered for swapping with an unpruned unit in the same layer; the exchange is accepted only if it lowers the block objective. Each pruned unit is processed once. Because the number of retained units in every layer is unchanged, rearrangement preserves the FLOPs or latency of the searched architecture while improving the placement of the zeros by accounting for intra-layer interactions.

  7. Knowl 7 — Layer-wise mask tuning by linear least squares

    algorithm

    After binary search and rearrangement, the framework tunes only the nonzero mask coefficients to compensate for the lost activation signal. Layers are processed from the first to the last. For a layer operation layer⁡\operatorname{layer}, let xlx_l and xl′x_l' be the inputs to that layer in the pruned and original models, respectively, and let mlm_l contain the coefficients of the retained heads or filters. The tuning target compares activations after the residual connection:

    min⁡ml∥xl+layer⁡(xl;ml)−(xl′+layer⁡(xl′;1))∥22.\min_{m_l}\left\|x_l+\operatorname{layer}(x_l;m_l)-\left(x_l'+\operatorname{layer}(x_l';\mathbf{1})\right)\right\|_2^2.

    For fixed inputs, this becomes a linear least-squares problem min⁡ml∥Alml−bl∥22\min_{m_l}\|A_lm_l-b_l\|_2^2, where the columns of AlA_l are the retained head-wise or filter-wise output activations and blb_l is the difference between the original and pruned residual outputs. With TT sample tokens and hidden size DD, AlA_l has TDTD rows and one column per retained head or filter.

    The implementation solves the large least-squares system with CuPy's LSMR solver using regularization parameter damp⁡=1\operatorname{damp}=1. It reparameterizes the coefficients as ml=1+rlm_l=\mathbf{1}+r_l, constrains every tuned coefficient to [−10,10][-10,10], and discards the candidate mask for a layer and stops tuning if the solver produces a coefficient outside that range. Since only existing nonzero entries are tuned, this stage changes neither the number of retained units nor the FLOPs or latency.

  8. Knowl 8 — Experimental protocol for evaluating the framework

    experimental setup

    The framework is implemented with PyTorch and HuggingFace Transformers and is evaluated on fine-tuned BERTBASE_{\mathrm{BASE}} and DistilBERT models. The downstream benchmarks are the GLUE tasks MNLI, QQP, QNLI, SST-2, STS-B, and MRPC, together with SQuAD 1.1 and SQuAD 2.0. Pruning uses 2,000 examples sampled from each task's training set, while accuracy is measured on the corresponding development set.

    Each reported result is averaged over 10 random seeds. Hardware latency is measured with PyTorch on a single NVIDIA V100 GPU in an AWS p3.2xlarge instance. The mask-tuning hyperparameters, LSMR damping value 11 and coefficient range [−10,10][-10,10], are fixed for all models and tasks rather than tuned separately.

  9. Knowl 9 — Accuracy, FLOPs, and latency results

    empirical result

    On GLUE and SQuAD, BERTBASE_{\mathrm{BASE}} retains 60–70% of its original FLOPs while staying within a 1% accuracy drop from the unpruned baseline across all evaluated tasks. DistilBERT achieves up to 50% FLOPs reduction within the same accuracy-drop target, including on the already compressed architecture.

    The reported page-8 V100 latency measurements give the following speedups for BERTBASE_{\mathrm{BASE}}, with the largest speedup satisfying at most a 1% accuracy degradation:

    • Batch size 32: MNLI 1.27×1.27\times, QQP 1.42×1.42\times, QNLI 1.42×1.42\times, SST-2 1.23×1.23\times, STS-B 1.34×1.34\times, MRPC 1.36×1.36\times, SQuAD 1.1 1.33×1.33\times, SQuAD 2.0 1.37×1.37\times, geometric mean 1.34×1.34\times.
    • Batch size 256: MNLI 1.34×1.34\times, QQP 1.54×1.54\times, QNLI 1.53×1.53\times, SST-2 1.56×1.56\times, STS-B 1.54×1.54\times, MRPC 1.55×1.55\times, SQuAD 1.1 1.34×1.34\times, SQuAD 2.0 1.40×1.40\times, geometric mean 1.47×1.47\times.

    Thus the measured maximum speedup is 1.56×1.56\times, while the method reaches roughly 2.0×2.0\times FLOPs reduction at the reported accuracy tolerance.

  10. Knowl 10 — Comparison with retraining-based structured pruning

    empirical result

    Compared with prior structured Transformer pruning methods including Flop, SLIP, layer dropping by Sajjad et al., DynaBERT, EBERT, Block Movement Pruning, and CoFi, the proposed method has a comparable or better FLOPs–accuracy trade-off on BERTBASE_{\mathrm{BASE}} GLUE tasks at moderate sparsity, despite using no model retraining. The comparison measures accuracy loss relative to each method's own baseline and excludes knowledge distillation and data augmentation. At high sparsity, the authors report that adding retraining to their framework gives comparable or better results at the same pruning cost.

    The page-10 pruning-cost data compare epochs and end-to-end pruning time on MNLI:

    • DynaBERT: 4 epochs and 12 hours.
    • EBERT: 6 epochs and 5 hours.
    • Block Movement Pruning: 20 epochs and 17 hours.
    • CoFi: 40 epochs and 33 hours.
    • Proposed method: 0 epochs and 0.01 hours.

    The proposed pipeline therefore finishes in under a minute for this comparison and is reported to be 2–3 orders of magnitude faster than the retraining-based alternatives, while also using only the two fixed mask-tuning hyperparameters.

  11. Knowl 11 — Ablation verifies the roles of rearrangement and tuning

    data/table

    The page-10 ablation evaluates BERTBASE_{\mathrm{BASE}} pruned to 60% of its original FLOPs. Scores are reported for eight tasks, with each successive row adding one component of the proposed pipeline:

    • Baseline: MNLI 84.53, QQP 91.00, QNLI 91.41, SST-2 93.57, STS-B 88.90, MRPC 86.27, SQuAD 1.1 88.48, SQuAD 2.0 76.82.
    • Fisher mask search: MNLI 81.21, QQP 89.99, QNLI 88.38, SST-2 92.13, STS-B 87.10, MRPC 83.14, SQuAD 1.1 82.66, SQuAD 2.0 71.12.
    • Search plus Fisher mask rearrangement: MNLI 81.81, QQP 90.08, QNLI 88.77, SST-2 92.09, STS-B 87.68, MRPC 83.23, SQuAD 1.1 84.47, SQuAD 2.0 72.38, with a reported average difference of +0.60+0.60 over mask search.
    • Search, rearrangement, and mask tuning: MNLI 82.51, QQP 90.35, QNLI 90.06, SST-2 92.49, STS-B 88.00, MRPC 85.27, SQuAD 1.1 86.72, SQuAD 2.0 75.26, with a reported average difference of +1.27+1.27 over the preceding configuration.

    The results show that rearrangement improves the binary mask and tuning recovers additional accuracy, with tuning particularly important. A separate comparison against weight-magnitude and gradient-based masks finds that those alternatives substantially degrade retraining-free accuracy at low sparsity, and their accuracy is not fully recovered by mask tuning; this supports the need for the Fisher search and rearrangement stages.

Coverage note — Detailed proof derivations, appendix-only implementation details, larger-sample analyses, and supplementary high-sparsity retraining experiments were omitted because they support the reported results but are not separate load-bearing contributions.

References

  1. 1.Yonathan Aflalo, Asaf Noy, Ming Lin, Itamar Friedman, and Lihi Zelnik. Knapsack pruning with inner distillation. arXiv preprint arXiv:2002.08258, 2020.
  2. 2.Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli. wav2vec 2.0: A framework for self-supervised learning of speech representations. arXiv preprint arXiv:2006.11477, 2020.
  3. 3.Ron Banner, Yury Nahshan, Elad Hoffer, and Daniel Soudry. Post-training 4-bit quantization of convolution networks for rapid-deployment. arXiv preprint arXiv:1810.05723, 2018.
  4. 4.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
  5. 5.Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia. Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation. arXiv preprint arXiv:1708.00055, 2017.
  6. 6.Daoyuan Chen, Yaliang Li, Minghui Qiu, Zhen Wang, Bofang Li, Bolin Ding, Hongbo Deng, Jun Huang, Wei Lin, and Jingren Zhou. Adabert: Task-adaptive bert compression with differentiable neural architecture search. arXiv preprint arXiv:2001.04246, 2020.
  7. 7.Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al. Wavlm: Large-scale self-supervised pre-training for full stack speech processing. arXiv preprint arXiv:2110.13900, 2021.
  8. 8.Tianlong Chen, Yu Cheng, Zhe Gan, Lu Yuan, Lei Zhang, and Zhangyang Wang. Chasing sparsity in vision transformers: An end-to-end exploration. Advances in Neural Information Processing Systems, 34:19974–19988, 2021.
  9. 9.Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Zhangyang Wang, and Michael Carbin. The lottery ticket hypothesis for pre-trained BERT networks. arXiv preprint arXiv:2007.12223, 2020.
  10. 10.Xiaohan Chen, Yu Cheng, Shuohang Wang, Zhe Gan, Zhangyang Wang, and Jingjing Liu. Earlybert: Efficient bert training via early-bird lottery tickets. arXiv preprint arXiv:2101.00063, 2020.
  11. 11.Ido Dagan, Oren Glickman, and Bernardo Magnini. The pascal recognising textual entailment challenge. In Machine Learning Challenges Workshop, pages 177–190. Springer, 2005.
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  13. 13.William B Dolan and Chris Brockett. Automatically constructing a corpus of sentential paraphrases. In Proceedings of the Third International Workshop on Paraphrasing (IWP2005), 2005.
  14. 14.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  15. 15.Angela Fan, Edouard Grave, and Armand Joulin. Reducing transformer depth on demand with structured dropout. arXiv preprint arXiv:1909.11556, 2019.
  16. 16.Jonathan Frankle and Michael Carbin. The lottery ticket hypothesis: Finding sparse, trainable neural networks. arXiv preprint arXiv:1803.03635, 2018.
  17. 17.Elias Frantar and Dan Alistarh. Spdy: Accurate pruning with speedup guarantees. arXiv preprint arXiv:2201.13096, 2022.
  18. 18.Trevor Gale, Erich Elsen, and Sara Hooker. The state of sparsity in deep neural networks. arXiv preprint arXiv:1902.09574, 2019.
  19. 19.Google. Tensorflow Lite: https://www.tensorflow.org/lite, 2017.
  20. 20.Cong Guo, Bo Yang Hsueh, Jingwen Leng, Yuxian Qiu, Yue Guan, Zehuan Wang, Xiaoying Jia, Xipeng Li, Minyi Guo, and Yuhao Zhu. Accelerating sparse dnn models without hardware-support via tile-wise sparsity. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–15. IEEE, 2020.
  21. 21.Tae Jun Ham, Sung Jun Jung, Seonghak Kim, Young H Oh, Yeonhong Park, Yoonho Song, Jung-Hun Park, Sanghee Lee, Kyoung Park, Jae W Lee, et al. Aˆ 3: Accelerating attention mechanisms in neural networks with approximation. In 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA), pages 328–341. IEEE, 2020.
  22. 22.Tae Jun Ham, Yejin Lee, Seong Hoon Seo, Soosung Kim, Hyunji Choi, Sung Jun Jung, and Jae W Lee. Elsa: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks. In 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA), pages 692–705. IEEE, 2021.
  23. 23.Yihui He, Xiangyu Zhang, and Jian Sun. Channel pruning for accelerating very deep neural networks. In Proceedings of the IEEE international conference on computer vision, pages 1389–1397, 2017.
  24. 24.Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016.
  25. 25.Lu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang, Xiao Chen, and Qun Liu. Dynabert: Dynamic bert with adaptive width and depth. arXiv preprint arXiv:2004.04037, 2020.
  26. 26.Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. arXiv preprint arXiv:2106.07447, 2021.
  27. 27.Itay Hubara, Yury Nahshan, Yair Hanani, Ron Banner, and Daniel Soudry. Improving post training neural quantization: Layer-wise calibration and integer programming. arXiv preprint arXiv:2006.10518, 2020.
  28. 28.Forrest N Iandola, Albert E Shaw, Ravi Krishna, and Kurt W Keutzer. Squeezebert: What can computer vision teach nlp about efficient neural networks? arXiv preprint arXiv:2006.11316, 2020.
  29. 29.Intel. OpenVINO: https://docs.openvino.ai/latest/index.html, 2021.
  30. 30.Shankar Iyer, Nikhil Dandekar, and Kornl Csernai. First quora dataset release: Question pairs.(2017). URL https://data. quora. com/First-Quora-Dataset-Release-Question-Pairs, 2017.
  31. 31.Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. Tinybert: Distilling bert for natural language understanding. arXiv preprint arXiv:1909.10351, 2019.
  32. 32.Ashish Khetan and Zohar Karnin. schubert: Optimizing elements of bert. arXiv preprint arXiv:2005.06628, 2020.
  33. 33.Sehoon Kim, Amir Gholami, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer. I-bert: Integer-only bert quantization. arXiv preprint arXiv:2101.01321, 2021.
  34. 34.Woojeong Kim, Suhyun Kim, Mincheol Park, and Geonseok Jeon. Neuron merging: Compensating for pruned neurons. arXiv preprint arXiv:2010.13160, 2020.
  35. 35.Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. Reformer: The efficient transformer. In International Conference on Learning Representations, 2019.
  36. 36.Eldar Kurtic, Daniel Campos, Tuan Nguyen, Elias Frantar, Mark Kurtz, Benjamin Fineran, Michael Goin, and Dan Alistarh. The optimal bert surgeon: Scalable and accurate second-order pruning for large language models. arXiv preprint arXiv:2203.07259, 2022.
  37. 37.Woosuk Kwon, Gyeong-In Yu, Eunji Jeong, and Byung-Gon Chun. Nimble: Lightweight and parallel gpu task scheduling for deep learning. arXiv preprint arXiv:2012.02732, 2020.
  38. 38.François Lagunas, Ella Charlaix, Victor Sanh, and Alexander M Rush. Block pruning for faster transformers. arXiv preprint arXiv:2109.04838, 2021.
  39. 39.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. Albert: A lite bert for self-supervised learning of language representations. arXiv preprint arXiv:1909.11942, 2019.
  40. 40.Ivan Lazarevich, Alexander Kozlov, and Nikita Malinin. Post-training deep neural network pruning via layer-wise calibration. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 798–805, 2021.
  41. 41.Yann LeCun, John S Denker, and Sara A Solla. Optimal brain damage. In Advances in neural information processing systems, pages 598–605, 1990.
  42. 42.Hector Levesque, Ernest Davis, and Leora Morgenstern. The winograd schema challenge. In Thirteenth International Conference on the Principles of Knowledge Representation and Reasoning. Citeseer, 2012.
  43. 43.Bingbing Li, Zhenglun Kong, Tianyun Zhang, Ji Li, Zhengang Li, Hang Liu, and Caiwen Ding. Efficient transformer-based large scale language representations using hardware-friendly block structured pruning. arXiv preprint arXiv:2009.08065, 2020.
  44. 44.Zi Lin, Jeremiah Zhe Liu, Zi Yang, Nan Hua, and Dan Roth. Pruning redundant mappings in transformer models via spectral-normalized identity prior. arXiv preprint arXiv:2010.01791, 2020.
  45. 45.Liyang Liu, Shilong Zhang, Zhanghui Kuang, Aojun Zhou, Jing-Hao Xue, Xinjiang Wang, Yimin Chen, Wenming Yang, Qingmin Liao, and Wayne Zhang. Group fisher pruning for practical network compression. In International Conference on Machine Learning, pages 7021–7032. PMLR, 2021.
  46. 46.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  47. 47.Yuanxin Liu, Zheng Lin, and Fengcheng Yuan. Rosita: Refined bert compression with integrated techniques. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 8715–8722, 2021.
  48. 48.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030, 2021.
  49. 49.Zejian Liu, Fanrong Li, Gang Li, and Jian Cheng. Ebert: Efficient bert inference with dynamic structured pruning. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4814–4823, 2021.
  50. 50.Lingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue, Youshan Miao, Wei Cui, Wenxiang Hu, Fan Yang, Lintao Zhang, and Lidong Zhou. Rammer: Enabling holistic deep learning compiler optimizations with rtasks. In 14th {USENIX} Symposium on Operating Systems Design and Implementation ({OSDI} 20), pages 881–897, 2020.
  51. 51.Paul Michel, Omer Levy, and Graham Neubig. Are sixteen heads really better than one? arXiv preprint arXiv:1905.10650, 2019.
  52. 52.Pavlo Molchanov, Arun Mallya, Stephen Tyree, Iuri Frosio, and Jan Kautz. Importance estimation for neural network pruning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11264–11272, 2019.
  53. 53.Ben Mussay, Margarita Osadchy, Vladimir Braverman, Samson Zhou, and Dan Feldman. Data-independent neural pruning via coresets. arXiv preprint arXiv:1907.04018, 2019.
  54. 54.Markus Nagel, Rana Ali Amjad, Mart Van Baalen, Christos Louizos, and Tijmen Blankevoort. Up or down? adaptive rounding for post-training quantization. In International Conference on Machine Learning, pages 7197–7206. PMLR, 2020.
  55. 55.ROYUD Nishino and Shohei Hido Crissman Loomis. Cupy: A numpy-compatible library for nvidia gpu calculations. 31st confernce on neural information processing systems, 151, 2017.
  56. 56.NVIDIA. TensorRT: https://developer.nvidia.com/tensorrt, 2018.
  57. 57.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32:8026–8037, 2019.
  58. 58.Sai Prasanna, Anna Rogers, and Anna Rumshisky. When BERT plays the lottery, all tickets are winning. arXiv preprint arXiv:2005.00561, 2020.
  59. 59.Valentin Radu, Kuba Kaszyk, Yuan Wen, Jack Turner, José Cano, Elliot J Crowley, Björn Franke, Amos Storkey, and Michael O’Boyle. Performance aware convolutional neural network channel pruning for embedded gpus. In 2019 IEEE International Symposium on Workload Characterization (IISWC), pages 24–34. IEEE, 2019.
  60. 60.Pranav Rajpurkar, Robin Jia, and Percy Liang. Know what you don’t know: Unanswerable questions for squad. arXiv preprint arXiv:1806.03822, 2018.
  61. 61.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. SQuAD: 100,000+ questions for machine comprehension of text. arXiv preprint arXiv:1606.05250, 2016.
  62. 62.Hassan Sajjad, Fahim Dalvi, Nadir Durrani, and Preslav Nakov. On the effect of dropping layers of pre-trained transformer models. arXiv preprint arXiv:2004.03844, 2020.
  63. 63.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019.
  64. 64.Victor Sanh, Thomas Wolf, and Alexander M Rush. Movement pruning: Adaptive sparsity by fine-tuning. arXiv preprint arXiv:2005.07683, 2020.
  65. 65.Maying Shen, Hongxu Yin, Pavlo Molchanov, Lei Mao, Jianna Liu, and Jose M Alvarez. Halp: Hardware-aware latency pruning. arXiv preprint arXiv:2110.10811, 2021.
  66. 66.Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer. Q-bert: Hessian based ultra low precision quantization of bert. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 8815–8821, 2020.
  67. 67.David So, Quoc Le, and Chen Liang. The evolved transformer. In International Conference on Machine Learning, pages 5877–5886. PMLR, 2019.
  68. 68.David R So, Wojciech Manke, Hanxiao Liu, Zihang Dai, Noam Shazeer, and Quoc V Le. Primer: Searching for efficient transformers for language modeling. arXiv preprint arXiv:2109.08668, 2021.
  69. 69.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing, pages 1631–1642, 2013.
  70. 70.Suraj Srinivas and R Venkatesh Babu. Data-free parameter pruning for deep neural networks. arXiv preprint arXiv:1507.06149, 2015.
  71. 71.Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. Patient knowledge distillation for bert model compression. arXiv preprint arXiv:1908.09355, 2019.
  72. 72.Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. Mobilebert: a compact task-agnostic bert for resource-limited devices. arXiv preprint arXiv:2004.02984, 2020.
  73. 73.Thierry Tambe, Coleman Hooper, Lillian Pentecost, Tianyu Jia, En-Yu Yang, Marco Donato, Victor Sanh, Paul Whatmough, Alexander M Rush, David Brooks, et al. Edgebert: Sentence-level energy optimizations for latency-aware multi-task nlp inference. In MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture, pages 830–844, 2021.
  74. 74.Lucas Theis, Iryna Korshunova, Alykhan Tejani, and Ferenc Huszár. Faster gaze prediction with dense networks and fisher pruning. arXiv preprint arXiv:1801.05787, 2018.
  75. 75.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning, pages 10347–10357. PMLR, 2021.
  76. 76.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017.
  77. 77.Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. arXiv preprint arXiv:1905.09418, 2019.
  78. 78.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. GLUE: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461, 2018.
  79. 79.Hanrui Wang, Zhanghao Wu, Zhijian Liu, Han Cai, Ligeng Zhu, Chuang Gan, and Song Han. Hat: Hardware-aware transformers for efficient natural language processing. arXiv preprint arXiv:2005.14187, 2020.
  80. 80.Hanrui Wang, Zhekai Zhang, and Song Han. Spatten: Efficient sparse attention architecture with cascade token and head pruning. In 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA), pages 97–110. IEEE, 2021.
  81. 81.Sinong Wang, Belinda Li, Madian Khabsa, Han Fang, and Hao Ma. Linformer: Self-attention with linear complexity. arXiv preprint arXiv:2006.04768, 2020.
  82. 82.Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers. arXiv preprint arXiv:2002.10957, 2020.
  83. 83.Ziheng Wang, Jeremy Wohlwend, and Tao Lei. Structured pruning of large language models. arXiv preprint arXiv:1910.04732, 2019.
  84. 84.Alex Warstadt, Amanpreet Singh, and Samuel R Bowman. Neural network acceptability judgments. Transactions of the Association for Computational Linguistics, 7:625–641, 2019.
  85. 85.Adina Williams, Nikita Nangia, and Samuel R Bowman. A broad-coverage challenge corpus for sentence understanding through inference. arXiv preprint arXiv:1704.05426, 2017.
  86. 86.Thomas Wolf, Julien Chaumond, Lysandre Debut, Victor Sanh, Clement Delangue, Anthony Moi, Pierric Cistac, Morgan Funtowicz, Joe Davison, Sam Shleifer, et al. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, 2020.
  87. 87.Zhanghao Wu, Zhijian Liu, Ji Lin, Yujun Lin, and Song Han. Lite transformer with long-short range attention. arXiv preprint arXiv:2004.11886, 2020.
  88. 88.Mengzhou Xia, Zexuan Zhong, and Danqi Chen. Structured pruning learns compact and accurate models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1513–1528, 2022.
  89. 89.Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin. Deebert: Dynamic early exiting for accelerating bert inference. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2246–2251, 2020.
  90. 90.Jin Xu, Xu Tan, Renqian Luo, Kaitao Song, Jian Li, Tao Qin, and Tie-Yan Liu. Nas-bert: Task-agnostic and adaptive-size bert compression with neural architecture search. arXiv preprint arXiv:2105.14444, 2021.
  91. 91.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. Xlnet: Generalized autoregressive pretraining for language understanding. Advances in neural information processing systems, 32, 2019.
  92. 92.Zhewei Yao, Linjian Ma, Sheng Shen, Kurt Keutzer, and Michael W Mahoney. Mlpruning: A multilevel structured pruning framework for transformer-based models. arXiv preprint arXiv:2105.14636, 2021.
  93. 93.Yichun Yin, Cheng Chen, Lifeng Shang, Xin Jiang, Xiao Chen, and Qun Liu. Autotinybert: Automatic hyper-parameter optimization for efficient pre-trained language models. arXiv preprint arXiv:2107.13686, 2021.
  94. 94.Edouard Yvinec, Arnaud Dapogny, Matthieu Cord, and Kevin Bailly. Red: Looking for redundancies for data-freestructured compression of deep neural networks. Advances in Neural Information Processing Systems, 34:20863–20873, 2021.
  95. 95.Ali Hadi Zadeh, Isak Edo, Omar Mohamed Awad, and Andreas Moshovos. Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference. In 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), pages 811–824. IEEE, 2020.
  96. 96.Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. Q8bert: Quantized 8bit bert. arXiv preprint arXiv:1910.06188, 2019.
  97. 97.Ritchie Zhao, Yuwei Hu, Jordan Dotzel, Christopher De Sa, and Zhiru Zhang. Improving neural network quantization without retraining using outlier channel splitting. Proceedings of Machine Learning Research, 2019.
  98. 98.Wangchunshu Zhou, Canwen Xu, Tao Ge, Julian McAuley, Ke Xu, and Furu Wei. Bert loses patience: Fast and robust inference with early exit. Advances in Neural Information Processing Systems, 33:18330–18341, 2020.

Citation

MLA
Kwon, W., et al. “A Fast Post-Training Pruning Framework for Transformers”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 24101–16, https://proceedings.neurips.cc/paper_files/paper/2022/file/987bed997ab668f91c822a09bce3ea12-Paper-Conference.pdf.
APA
Kwon, W., Kim, S., Mahoney, M., Hassoun, J., Keutzer, K., & Gholami, A. (2022). A Fast Post-Training Pruning Framework for Transformers. Advances in Neural Information Processing Systems, 35, 24101–24116. https://proceedings.neurips.cc/paper_files/paper/2022/file/987bed997ab668f91c822a09bce3ea12-Paper-Conference.pdf
Chicago
Kwon, W., S. Kim, M. Mahoney, J. Hassoun, K. Keutzer, and A. Gholami. 2022. “A Fast Post-Training Pruning Framework for Transformers”. Advances in Neural Information Processing Systems 35: 24101–16. https://proceedings.neurips.cc/paper_files/paper/2022/file/987bed997ab668f91c822a09bce3ea12-Paper-Conference.pdf.
Harvard
Kwon, W. et al. (2022) “A Fast Post-Training Pruning Framework for Transformers”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 24101–24116. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/987bed997ab668f91c822a09bce3ea12-Paper-Conference.pdf.
Vancouver
1. Kwon W, Kim S, Mahoney M, Hassoun J, Keutzer K, Gholami A (2022) A Fast Post-Training Pruning Framework for Transformers. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 24101–24116

BibTeX

@inproceedings{kwon2022fast,
  title = {A Fast Post-Training Pruning Framework for Transformers},
  author = {Kwon, Woosuk and Kim, Sehoon and Mahoney, Michael and Hassoun, Joseph and Keutzer, Kurt and Gholami, Amir},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {24101-24116},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/987bed997ab668f91c822a09bce3ea12-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors