Structured Pruning Learns Compact and Accurate Models

Mengzhou XiaZexuan ZhongDanqi Chen

article2022ACL259 citationsOutstanding Paper Award

Proposes CoFi, a structured pruning method that jointly eliminates coarse- and fine-grained Transformer components to achieve over tenfold inference speedups while matching the accuracy of expensive distillation baselines without requiring unlabeled pre-training data.

Listen

Modern natural language processing relies heavily on large pre-trained language models that demand substantial memory, storage, and computing power. To deploy these models effectively in production, organizations typically turn to model compression techniques: knowledge distillation, which trains smaller models to mimic larger ones, or model pruning, which removes redundant parameters from existing models. While distillation achieves high inference speedups, it requires computationally expensive pre-training over billions of unlabeled tokens. Conversely, existing structured pruning techniques create flexible sub-networks but struggle to achieve speedups beyond twofold to threefold acceleration.

The article introduces and evaluates CoFi (Coarse- and Fine-grained Pruning), a task-specific structured pruning framework designed to deliver highly parallelizable, compact models. The objective is to demonstrate that structured pruning can match or exceed the accuracy and inference latency of state-of-the-art distillation methods without requiring massive unlabeled datasets or prolonged pre-training cycles.

The approach operates by jointly learning pruning decisions across multiple granularities simultaneously, including entire attention and feed-forward layers, individual attention heads, and hidden dimensions. These components are regularized using sparsity constraints that allow the optimization process to dynamically determine the most efficient network shape. To preserve model quality during pruning, the framework introduces a dynamic layerwise distillation technique that automatically aligns and transfers intermediate representations from the full teacher model to the evolving pruned student model. The method was evaluated across eight GLUE benchmark tasks and the SQuAD question-answering dataset using standard base Transformer architectures.

The evaluation yielded several key findings. First, CoFi achieves over 10-fold inference speedups on GPUs with a 95% parameter sparsity rate while preserving over 90% of the baseline model accuracy. Second, it matches or outperforms leading distillation baselines like TinyBERT while slashing training time from roughly 350 GPU hours down to under 20 GPU hours on a single GPU. Third, joint coarse-and-fine pruning proved crucial: omitting whole-layer pruning severely degraded speedups from 12.1-fold to 7.0–8.3-fold at high sparsity, while omitting hidden-dimension pruning degraded model accuracy. Finally, analysis revealed that feed-forward layers contain substantially more redundancy than attention layers, showing a 71% reduction in intermediate dimensions compared to a 39% reduction in attention heads at 60% overall sparsity.

These results demonstrate that task-specific structured pruning provides an efficient, low-cost path to production-ready language models. Organizations can bypass the heavy computational overhead, engineering complexity, and data management risks associated with large-scale unlabeled data distillation. By adapting network depth and width directly to specific downstream tasks, teams can deploy models that run over ten times faster on standard hardware without significant predictive degradation.

Decision-makers and engineering teams should consider adopting multi-granularity structured pruning as a preferred compression pipeline for task-specific deployments, particularly where compute budgets or training turnarounds are constrained. For existing models, pruning fine-tuned weights directly with dynamic intermediate distillation offers the best balance of speed and retention. Future work should pilot this approach on broader generative architectures and investigate upstream task-agnostic pruning to establish general-purpose compact base models.

The primary limitation of the study is its focus on task-specific compression of encoder architectures, meaning the pruned models cannot be universally reused across distinct tasks without retraining. High confidence in these findings is supported by consistent empirical improvements across multiple standard benchmarks and clear ablation studies.

Cover for Structured Pruning Learns Compact and Accurate Models

Abstract

The growing size of neural language models has led to increased attention in model compression. The two predominant approaches are pruning, which gradually removes weights from a pre-trained model, and distillation, which trains a smaller compact model to match a larger one. Pruning methods can significantly reduce the model size but hardly achieve large speedups as distillation. However, distillation methods require large amounts of unlabeled data and are expensive to train. In this work, we propose a task-specific structured pruning method CoFi¹ (Coarse- and Fine-grained Pruning), which delivers highly parallelizable subnetworks and matches the distillation methods in both accuracy and latency, without resorting to any unlabeled data. Our key insight is to jointly prune coarse-grained (e.g., layers) and fine-grained (e.g., heads and hidden units) modules, which controls the pruning decision of each parameter with masks of different granularity. We also devise a layerwise distillation strategy to transfer knowledge from unpruned to pruned models during optimization. Our experiments on GLUE and SQuAD datasets show that CoFi yields models with over 10× speedups with a small accuracy drop, showing its effectiveness and efficiency compared to previous pruning and distillation approaches.²

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Transformers
  • 2.2 Distillation
  • 2.3 Pruning
  • 3 Method
  • 3.1 Coarse- and Fine-Grained Pruning
  • 3.2 Distillation to Pruned Models
  • 4 Experiments
  • 4.1 Setup
  • 4.2 Main Results
  • 4.3 Ablation Study
  • 4.4 Structures of Pruned Models
  • 5 Related Work
  • 6 Conclusion
  • Acknowledgements
  • References
  • A Reproducibility & Hyperparameters
  • B Optimization Details
  • C Details of Baseline Methods
  • D Data Statistics
  • E TinyBERT 4 w/ Data Augmentation
  • F Additional Comparisons
  • F.1 Comparison to Movement Pruning
  • F.2 Comparison to Block Pruning
  • F.3 More Baselines
  • G More Analyses on Layer Distillation
  • G.1 Layer Alignment
  • G.2 Ablation on Distillation Objectives
  • H FFN/MHA Layers in Pruned Models
  • I RoBERTa Pruning
  • J Training Time Measurement

Knowls

  1. Knowl 1 — CoFi Multi-Granularity Masking Framework for Transformer Pruning

    model/method

    Coarse- and Fine-grained Pruning (CoFi) is a structured pruning method for Transformer models that jointly controls parameter removal across multiple structural granularities: whole layers, sub-layer modules, and hidden vector dimensions.

    For a Transformer with LL layers, hidden dimension dd, NhN_h attention heads per layer (each with projection dimension dh=d/Nhd_h = d / N_h), and feed-forward intermediate dimension dfd_f (typically 4d4d), CoFi introduces five sets of binary mask variables:

    1. Multi-Head Attention (MHA) layer mask: z_{MHA}^{(l)} ackslashin \{0, 1\} for each layer l∈{1,…,L}l \in \{1, \dots, L\}.
    2. Attention head mask: zhead(l,i)∈{0,1}z_{head}^{(l, i)} \in \{0, 1\} for each head i∈{1,…,Nh}i \in \{1, \dots, N_h\} in layer ll.
    3. Feed-Forward Network (FFN) layer mask: zFFN(l)∈{0,1}z_{FFN}^{(l)} \in \{0, 1\} for each layer l∈{1,…,L}l \in \{1, \dots, L\}.
    4. FFN intermediate dimension mask: zint(l,j)∈{0,1}z_{int}^{(l, j)} \in \{0, 1\} for each intermediate unit j∈{1,…,df}j \in \{1, \dots, d_f\} in layer ll.
    5. Hidden dimension mask: zhidden(k)∈{0,1}z_{hidden}^{(k)} \in \{0, 1\} for each output/hidden dimension k∈{1,…,d}k \in \{1, \dots, d\}, shared across all layers due to residual connections.

    Under these masks, the forward pass of multi-head attention and feed-forward layers with input matrix XX is defined as:

    MHA(X)=zMHA(l)⋅∑i=1Nh(zhead(l,i)⋅Att(WQ(l,i),WK(l,i),WV(l,i),WO(l,i),X))\text{MHA}(X) = z_{MHA}^{(l)} \cdot \sum_{i=1}^{N_h} \left( z_{head}^{(l, i)} \cdot \text{Att}\left(W_Q^{(l, i)}, W_K^{(l, i)}, W_V^{(l, i)}, W_O^{(l, i)}, X\right) \right)

    FFN(X)=zFFN(l)⋅GELU(XWU(l))⋅diag(zint(l))⋅WD(l)\text{FFN}(X) = z_{FFN}^{(l)} \cdot \text{GELU}\left(X W_U^{(l)}\right) \cdot \text{diag}\left(z_{int}^{(l)}\right) \cdot W_D^{(l)}

    where WQ(l,i),WK(l,i),WV(l,i)∈Rd×dhW_Q^{(l, i)}, W_K^{(l, i)}, W_V^{(l, i)} \in \mathbb{R}^{d \times d_h}, WO(l,i)∈Rdh×dW_O^{(l, i)} \in \mathbb{R}^{d_h \times d}, WU(l)∈Rd×dfW_U^{(l)} \in \mathbb{R}^{d \times d_f}, and WD(l)∈Rdf×dW_D^{(l)} \in \mathbb{R}^{d_f \times d}.

    The hidden dimension mask is applied across all weight matrices via diag(zhidden)\text{diag}(z_{hidden}). Consequently, an individual weight parameter in the network is dropped if its enclosing layer mask, its sub-layer unit mask, or its corresponding hidden dimension mask is zero.

  2. Knowl 2 — Sparsity Formulation and Hard Concrete Lagrangian Optimization

    model/method

    In CoFi pruning, the expected fraction of remaining parameters s^\hat{s} relative to the full parameter count MM (excluding embedding layers) is computed analytically from the mask variables as:

    s^=1M⋅4⋅dh⋅∑l=1L∑i=1Nh∑k=1dzMHA(l)⋅zhead(l,i)⋅zhidden(k)+1M⋅2⋅∑l=1L∑j=1df∑k=1dzFFN(l)⋅zint(l,j)⋅zhidden(k)\hat{s} = \frac{1}{M} \cdot 4 \cdot d_h \cdot \sum_{l=1}^{L} \sum_{i=1}^{N_h} \sum_{k=1}^{d} z_{MHA}^{(l)} \cdot z_{head}^{(l, i)} \cdot z_{hidden}^{(k)} + \frac{1}{M} \cdot 2 \cdot \sum_{l=1}^{L} \sum_{j=1}^{d_f} \sum_{k=1}^{d} z_{FFN}^{(l)} \cdot z_{int}^{(l, j)} \cdot z_{hidden}^{(k)}

    where LL is the number of layers, NhN_h is the number of attention heads per layer, dhd_h is head dimension, dfd_f is feed-forward intermediate dimension, and dd is hidden dimension.

    To make discrete masks differentiable during training, each binary mask variable zz is parameterized using a Hard Concrete continuous relaxation. For a uniform random variable u∼U(0,1)u \sim \mathcal{U}(0, 1) and learnable parameter log⁡α\log \alpha:

    s=sigmoid(1β(log⁡u−log⁡(1−u))+log⁡α)s = \text{sigmoid}\left(\frac{1}{\beta} \left(\log u - \log(1 - u)\right) + \log \alpha\right)

    s~=s⋅(r−l)+l\tilde{s} = s \cdot (r - l) + l

    z=min⁡(1,max⁡(0,s~))z = \min(1, \max(0, \tilde{s}))

    where β\beta controls the distribution steepness, and l<0l < 0 and r>0r > 0 are stretching constants that push probability mass onto exact 0 and 1.

    To strictly meet a target sparsity t∈(0,1)t \in (0, 1), CoFi replaces standard L0L_0 regularization penalties with a Lagrangian multiplier equality constraint loss LcL_c:

    Lc=λ1⋅(s^−t)+λ2⋅(s^−t)2L_c = \lambda_1 \cdot (\hat{s} - t) + \lambda_2 \cdot (\hat{s} - t)^2

    where λ1\lambda_1 and λ2\lambda_2 are Lagrangian multipliers adjusted during training to penalize deviations from the target sparsity tt.

  3. Knowl 3 — Dynamic Layerwise Distillation for Evolving Pruned Subnetworks

    model/method

    Because structured pruning dynamically modifies network depth by eliminating entire layers during training, static layer-to-layer knowledge distillation maps fail. CoFi introduces dynamic layerwise distillation that computes an adaptive alignment from teacher layers to the closest active student layers.

    Let T\mathcal{T} denote the set of teacher layer indices selected for distillation (e.g., layers 3, 6, 9, 12 in a 12-layer teacher). The dynamic layer mapping function m(i)m(i) assigns each teacher layer i∈Ti \in \mathcal{T} to the student layer jj that minimizes mean squared error among all currently active student FFN layers (i.e., where zFFN(j)>0z_{FFN}^{(j)} > 0):

    m(i)=arg⁡min⁡j:zFFN(j)>0MSE(WlayerHs(j),Ht(i))m(i) = \arg\min_{j : z_{FFN}^{(j)} > 0} \text{MSE}\left(W_{layer} H_s^{(j)}, H_t^{(i)}\right)

    where Ht(i)∈RdH_t^{(i)} \in \mathbb{R}^{d} and Hs(j)∈RdH_s^{(j)} \in \mathbb{R}^{d} are the hidden representation vectors from the ii-th teacher and jj-th student FFN layers respectively, and Wlayer∈Rd×dW_{layer} \in \mathbb{R}^{d \times d} is a learnable linear transformation initialized as an identity matrix. For small datasets, a monotonic order constraint m(i)≥m(i−1)m(i) \ge m(i-1) is enforced to prevent layer inversion.

    The intermediate layer distillation loss is:

    Llayer=∑i∈TMSE(WlayerHsm(i),Ht(i))L_{layer} = \sum_{i \in \mathcal{T}} \text{MSE}\left(W_{layer} H_s^{m(i)}, H_t^{(i)}\right)

    The total distillation objective LdistilL_{distil} linearly combines intermediate layer distillation with output prediction (logit) distillation:

    Ldistil=λLpred+(1−λ)LlayerL_{distil} = \lambda L_{pred} + (1 - \lambda) L_{layer}

    where Lpred=DKL(ps∥pt)L_{pred} = D_{KL}(p_s \parallel p_t) is the Kullback-Leibler divergence between student output probability distribution psp_s and teacher output distribution ptp_t, and λ∈[0,1]\lambda \in [0, 1] balances the two loss terms.

  4. Knowl 4 — CoFi Structured Pruning and Distillation Pipeline

    algorithm

    The complete training and optimization procedure for CoFi structured pruning combines distillation warm-up, constrained mask searching, and post-pruning subnetwork fine-tuning:

    Input: Pre-trained teacher model T, fine-tuned dense base student model S, task training dataset D, target sparsity t, teacher distillation layer set \mathcal{T}, balancing factor \lambda
    Output: Compact pruned student model S*
    Initialize mask distribution parameters \log \alpha for z_{MHA}, z_{FFN}, z_{head}, z_{int}, z_{hidden}
    Initialize linear projection matrix W_{layer} \leftarrow I_{d \times d}
    // Phase 1: Distillation Warmup
    for epoch = 1 to N_{warmup} do
        for batch in D do
            Compute L_{pred} = D_{KL}(p_s \parallel p_t)
            Compute dynamic layer map m(i) for each i \in \mathcal{T}
            Compute L_{layer} using m(i) and W_{layer}
            L_{distil} = \lambda L_{pred} + (1 - \lambda) L_{layer}
            Update student weights and W_{layer} using \nabla L_{distil}
        end for
    end for
    // Phase 2: Joint Pruning and Distillation Search
    for epoch = 1 to N_{prune} do
        Update current target sparsity t_{curr} via linear schedule up to t
        for batch in D do
            Sample masks z via Hard Concrete distributions
            Compute expected sparsity \hat{s} from mask parameters
            Compute L_{distil} with active student layers
            Compute Lagrangian sparsity loss L_c = \lambda_1 (\hat{s} - t_{curr}) + \lambda_2 (\hat{s} - t_{curr})^2
            L_{total} = L_{distil} + L_c
            Update student weights, \log \alpha, W_{layer}, \lambda_1, \lambda_2 using \nabla L_{total}
        end for
    end for
    // Phase 3: Final Subnetwork Extraction and Fine-tuning
    Extract deterministic binary subnetwork S* by thresholding mask parameters \log \alpha
    for epoch = 1 to N_{finetune} do
        for batch in D do
            Compute task loss and L_{distil} on S*
            Update remaining active parameters of S*
        end for
    end for
    return S*
  5. Knowl 5 — Performance and Efficiency of CoFi vs. Distillation and Pruning Baselines

    data/table

    The table below compares the unpruned BERTbase\text{BERT}_{\text{base}} teacher model against distillation and structured pruning baselines at ∼10×\sim 10\times inference acceleration on an NVIDIA V100 GPU across the GLUE benchmark and SQuAD v1.1. Models operate at ∼95%\sim 95\% parameter sparsity (sim4.4M−5.0M\\sim 4.4\text{M}-5.0\text{M} non-embedding parameters). Training time is reported in GPU hours.

    Model SST-2 QNLI MNLI QQP CoLA RTE STS-B MRPC SQuAD Train Time
    BERTbase\text{BERT}_{\text{base}} (teacher) 93.1 91.5 84.8 91.2 61.2 70.0 88.7 85.0 88.4 -
    TinyBERT4\text{TinyBERT}_4 w/o GD 87.7 81.8 78.7 89.5 16.6 47.3 17.8 68.9 - ≤10\le 10 h
    TinyBERT4\text{TinyBERT}_4 89.7 86.7 78.8 90.0 32.5 63.2 85.0 81.4 82.1 ∼350\sim 350 h
    Speedup 11.4×11.4\times 11.4×11.4\times 11.4×11.4\times 11.4×11.4\times 11.4×11.4\times 11.4×11.4\times 11.4×11.4\times 11.4×11.4\times 8.7×8.7\times -
    CoFi Pruning (ours) 90.6 86.1 80.6 90.1 35.6 64.7 83.1 82.6 82.6 ≤20\le 20 h
    Speedup 12.0×12.0\times 12.1×12.1\times 12.1×12.1\times 11.0×11.0\times 11.5×11.5\times 11.9×11.9\times 12.9×12.9\times 11.9×11.9\times 8.7×8.7\times -

    These results demonstrate that CoFi matches or exceeds the accuracy of TinyBERT4\text{TinyBERT}_4 across all tasks while delivering comparable or higher empirical inference speedup (11.0×−12.9×11.0\times - 12.9\times on GLUE, 8.7×8.7\times on SQuAD). Furthermore, CoFi eliminates the need for expensive general distillation (GD) on large unlabeled corpora, requiring at most 20 GPU hours on a single GPU compared to ∼350\sim 350 GPU hours for TinyBERT4\text{TinyBERT}_4.

  6. Knowl 6 — Ablation of Granularity Masks on Pruned Model Latency and Accuracy

    empirical result

    Ablation of CoFi pruning components on QNLI, MNLI, and SQuAD demonstrates the distinct computational and performance roles played by layer masks (zMHA,zFFNz_{MHA}, z_{FFN}) and hidden dimension masks (zhiddenz_{hidden}):

    1. Impact of Layer Masks (zMHA,zFFNz_{MHA}, z_{FFN}): Removing layer-level masks (−layer-layer) while keeping head and intermediate dimension pruning maintains accuracy at 95%95\% sparsity (5M parameters) but severely reduces GPU inference speedup: on QNLI, speedup drops from 12.1×12.1\times to 8.3×8.3\times; on MNLI, from 12.1×12.1\times to 8.4×8.4\times; and on SQuAD, from 8.7×8.7\times to 7.9×7.9\times. At 60%60\% sparsity (34M parameters), removing layer masks does not significantly degrade speedup (2.1×2.1\times vs 2.1×2.1\times). This establishes that dropping entire layers is the primary driver of high inference acceleration in extreme compression regimes.

    2. Impact of Hidden Dimension Masks (zhiddenz_{hidden}): Removing hidden dimension masks (−hidden-hidden) slightly increases inference speedup because the optimizer removes more entire layers to meet the sparsity budget, but it causes consistent accuracy degradation across all tasks: on QNLI (95%95\% sparsity), accuracy falls from 86.186.1 to 85.685.6; on MNLI (95%95\% sparsity), from 80.680.6 to 79.879.8; and on SQuAD (95%95\% sparsity), F1 drops from 82.682.6 to 80.880.8.

    3. Removing both layer and hidden dimension masks (−layer  &  hidden-layer \;\&\; hidden) causes severe degradation in both speedup (7.2×7.2\times on QNLI, 7.0×7.0\times on MNLI, 6.4×6.4\times on SQuAD at 95%95\% sparsity) and task performance (accuracy drops to 84.684.6 on QNLI, 78.478.4 on MNLI, and 74.174.1 F1 on SQuAD).

  7. Knowl 7 — Ablation of Distillation Objectives Across Pruning Sparsities

    empirical result

    Ablation studies across target sparsities (60%60\% to 95%95\%) and GLUE/SQuAD tasks demonstrate the critical contribution of dynamic layerwise distillation in structured pruning:

    1. Distillation Necessity: Completely removing knowledge distillation (−Lpred,−Llayer-L_{pred}, -L_{layer}) from CoFi at 95%95\% sparsity causes a major drop in performance across all tasks: SST-2 falls from 90.690.6 to 86.686.6 (−4.0-4.0), QNLI falls from 86.186.1 to 84.284.2 (−1.9-1.9), MNLI falls from 80.680.6 to 78.278.2 (−2.4-2.4), and SQuAD falls from 82.682.6 to 75.875.8 (−6.8-6.8).

    2. Dynamic vs. Fixed Distillation: Replacing dynamic layer matching with static 1-to-1 layer distillation ("Fixed Hidden Distillation", matching student layer jj directly to teacher layer jj when active) degrades accuracy at 95%95\% sparsity across SST-2 (90.6→90.090.6 \to 90.0), QNLI (86.1→85.886.1 \to 85.8), MNLI (80.6→80.580.6 \to 80.5), and SQuAD (82.6→80.982.6 \to 80.9).

    3. Dynamic Layer Loss Across Sparsity Levels: Adding dynamic layerwise distillation (+Llayer+L_{layer}) onto prediction logit distillation (LpredL_{pred}) improves accuracy at every sparsity tier from 60%60\% to 95%95\%. The gain is most pronounced in moderate-to-high sparsity regimes: on QNLI, +Llayer+L_{layer} provides +1.18+1.18 at 60%60\%, +2.35+2.35 at 75%75\%, +2.85+2.85 at 85%85\%, +3.19+3.19 at 90%90\%, and +1.06+1.06 at 95%95\% sparsity; on MNLI, gains reach up to +1.52+1.52 (90%90\% sparsity).

  8. Knowl 8 — Structural Properties and Layer Asymmetry in Pruned Transformer Subnetworks

    empirical result

    Analysis of subnetworks discovered by CoFi reveals distinct structural patterns in how Transformer parameters are allocated after structured pruning:

    1. FFN vs. MHA Redundancy: Feed-forward (FFN) layers are pruned much more aggressively than Multi-Head Attention (MHA) layers. At 60%60\% overall model sparsity, the average number of FFN intermediate dimensions is reduced by 71%71\% (3072→8843072 \to 884), whereas the average number of attention heads in MHA layers is reduced by only 39%39\% (12→7.312 \to 7.3). Across all sparsity targets (60%−95%60\%-95\%), pruned models systematically preserve more MHA layers than FFN layers.

    2. Depth and Layer Asymmetry: CoFi prunes submodules more heavily from upper layers than from lower layers; upper MHA layers retain fewer attention heads on average than lower MHA layers. Additionally, the first two MHA layers are preserved across almost all runs and datasets, whereas intermediate layers are frequently dropped entirely.

    3. Task-Specific Non-Interleaved Topologies: While standard Transformers strictly alternate MHA and FFN layers, CoFi discovers specialized, non-interleaved architectures with adjacent MHA or adjacent FFN layers that differ across tasks (e.g., preserving the first MHA layer in SST-2 and QNLI, while pruning it in QQP and SQuAD at 95%95\% sparsity).

  9. Knowl 9 — Effect of Task-Specific Data Augmentation on Highly Pruned Models

    empirical result

    When trained with task-specific data augmentation at 95%95\% sparsity (sim5M\\sim 5\text{M} non-embedding parameters, 11×−12×11\times - 12\times GPU inference acceleration), CoFi consistently achieves superior or competitive performance compared to TinyBERT4\text{TinyBERT}_4 under the identical data-augmented setting:

    • SST-2: CoFi improves from 90.690.6 (without augmentation) to 92.492.4 (with augmentation), compared to TinyBERT4\text{TinyBERT}_4 which improves from 89.789.7 to 91.691.6.
    • QNLI: CoFi improves from 86.186.1 to 86.886.8, compared to TinyBERT4\text{TinyBERT}_4 which improves from 86.786.7 to 87.687.6.
    • RTE: CoFi improves from 64.764.7 to 67.567.5, compared to TinyBERT4\text{TinyBERT}_4 which decreases from 63.263.2 to 62.562.5.
    • MRPC: CoFi improves from 82.682.6 to 84.684.6, compared to TinyBERT4\text{TinyBERT}_4 which improves from 81.481.4 to 83.683.6.

Coverage note — Omitted non-essential material: preliminary RoBERTa pruning curves (Appendix I) which follow similar trends to BERT, and secondary comparisons to non-directly comparable architectures (AutoTinyBERT from ELECTRA, MobileBERT with custom designs in Appendix F.3).

References

  1. 1.Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017. SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation. In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), pages 1–14, Vancouver, Canada. Association for Computational Linguistics.
  2. 2.Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Zhangyang Wang, and Michael Carbin. 2020a. The lottery ticket hypothesis for pre-trained BERT networks. In Advances in Neural Information Processing Systems (NeurIPS), pages 15834–15846.
  3. 3.Xiaohan Chen, Yu Cheng, Shuohang Wang, Zhe Gan, Zhangyang Wang, and Jingjing Liu. 2020b. Earlybert: Efficient bert training via early-bird lottery tickets. arXiv preprint arXiv:2101.00063.
  4. 4.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In North American Chapter of the Association for Computational Linguistics (NAACL), pages 4171–4186.
  5. 5.William B Dolan and Chris Brockett. 2005. Automatically constructing a corpus of sentential paraphrases. In Proceedings of the Third International Workshop on Paraphrasing (IWP2005).
  6. 6.Angela Fan, Edouard Grave, and Armand Joulin. 2020. Reducing Transformer depth on demand with structured dropout. In International Conference on Learning Representations (ICLR).
  7. 7.Angela Fan, Pierre Stock, Benjamin Graham, Edouard Grave, Rémi Gribonval, Herve Jegou, and Armand Joulin. 2021. Training with quantization noise for extreme model compression. In International Conference on Learning Representations (ICLR).
  8. 8.Jonathan Frankle and Michael Carbin. 2019. The lottery ticket hypothesis: Finding sparse, trainable neural networks. In International Conference on Learning Representations (ICLR).
  9. 9.Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin. 2020. Linear mode connectivity and the lottery ticket hypothesis. In International Conference on Machine Learning (ICML), pages 3259–3269.
  10. 10.Prakhar Ganesh, Yao Chen, Xin Lou, Mohammad Ali Khan, Yin Yang, Hassan Sajjad, Preslav Nakov, Deming Chen, and Marianne Winslett. 2021. Compressing large-scale Transformer-based models: A case study on BERT. Transactions of the Association of Computational Linguistics (TACL), 9:1061–1080.
  11. 11.Shaopeng Guo, Yujie Wang, Quanquan Li, and Junjie Yan. 2020. Dmcp: Differentiable markov channel pruning for neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  12. 12.Yihui He, Xiangyu Zhang, and Jian Sun. 2017. Channel pruning for accelerating very deep neural networks. In International Conference on Computer Vision (ICCV), pages 1389–1397.
  13. 13.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531.
  14. 14.Lu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang, Xiao Chen, and Qun Liu. 2020. DynaBERT: Dynamic bert with adaptive width and depth. In Advances in Neural Information Processing Systems (NeurIPS), volume 33.
  15. 15.Shaoyi Huang, Dongkuan Xu, Ian EH Yen, Sung-en Chang, Bingbing Li, Shiyang Chen, Mimi Xie, Hang Liu, and Caiwen Ding. 2021. Sparse progressive distillation: Resolving overfitting under pretrain-and-finetune paradigm. arXiv e-prints, pages arXiv–2110.
  16. 16.Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020. TinyBERT: Distilling BERT for natural language understanding. In Findings of Empirical Methods in Natural Language Processing (EMNLP), pages 4163–4174.
  17. 17.François Lagunas, Ella Charlaix, Victor Sanh, and Alexander M Rush. 2021. Block pruning for faster transformers. arXiv preprint arXiv:2109.04838.
  18. 18.Jiaoda Li, Ryan Cotterell, and Mrinmaya Sachan. 2021. Differentiable subset pruning of Transformer heads. Transactions of the Association of Computational Linguistics (TACL).
  19. 19.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019a. RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692.
  20. 20.Zechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo, Xin Yang, Kwang-Ting Cheng, and Jian Sun. 2019b. Metapruning: Meta learning for automatic neural network channel pruning. In International Conference on Computer Vision (ICCV), pages 3296–3305.
  21. 21.Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang. 2017. Learning efficient convolutional networks through network slimming. In International Conference on Computer Vision (ICCV), pages 2736–2744.
  22. 22.Zhuang Liu, Mingjie Sun, Tinghui Zhou, Gao Huang, and Trevor Darrell. 2019c. Rethinking the value of network pruning. In International Conference on Learning Representations (ICLR).
  23. 23.C Louizos, M Welling, and DP Kingma. 2018. Learning sparse neural networks through l0 regularization. In International Conference on Learning Representations (ICLR).
  24. 24.Jian-Hao Luo, Jianxin Wu, and Weiyao Lin. 2017. Thinet: A filter level pruning method for deep neural network compression. In International Conference on Computer Vision (ICCV), pages 5058–5066.
  25. 25.JS McCarley, Rishav Chakravarti, and Avirup Sil. 2019. Structured pruning of a BERT-based question answering model. arXiv preprint arXiv:1910.06360.
  26. 26.Paul Michel, Omer Levy, and Graham Neubig. 2019. Are sixteen heads really better than one? In Advances in Neural Information Processing Systems (NeurIPS).
  27. 27.Pavlo Molchanov, Arun Mallya, Stephen Tyree, Iuri Frosio, and Jan Kautz. 2019. Importance estimation for neural network pruning. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11264–11272.
  28. 28.Matan Ben Noach and Yoav Goldberg. 2020. Compressing pre-trained language models by matrix decomposition. In Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, pages 884–889.
  29. 29.Sai Prasanna, Anna Rogers, and Anna Rumshisky. 2020. When BERT plays the lottery, all tickets are winning. In Empirical Methods in Natural Language Processing (EMNLP), pages 3208–3229.
  30. 30.Ofir Press, Noah A Smith, and Omer Levy. 2020. Improving transformer models by reordering their sublayers. In Association for Computational Linguistics (ACL), pages 2996–3005.
  31. 31.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text Transformer. The Journal of Machine Learning Research (JMLR), 21(140).
  32. 32.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Empirical Methods in Natural Language Processing (EMNLP), pages 2383–2392.
  33. 33.Alex Renda, Jonathan Frankle, and Michael Carbin. 2020. Comparing rewinding and fine-tuning in neural network pruning. In International Conference on Learning Representations (ICLR).
  34. 34.Hassan Sajjad, Fahim Dalvi, Nadir Durrani, and Preslav Nakov. 2020. Poor man’s BERT: Smaller and faster transformer models. arXiv preprint arXiv:2004.03844.
  35. 35.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. DistilBERT, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108.
  36. 36.Victor Sanh, Thomas Wolf, and Alexander Rush. 2020. Movement pruning: Adaptive sparsity by fine-tuning. Advances in Neural Information Processing Systems (NeurIPS), 33.
  37. 37.Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer. 2020. Q-BERT: Hessian based ultra low precision quantization of BERT. In Conference on Artificial Intelligence (AAAI), pages 8815–8821.
  38. 38.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Empirical Methods in Natural Language Processing (EMNLP).
  39. 39.Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019. Patient knowledge distillation for BERT model compression. In Empirical Methods in Natural Language Processing (EMNLP), pages 4314–4323.
  40. 40.Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020. MobileBERT: a compact task-agnostic bert for resource-limited devices. In Association for Computational Linguistics (ACL), pages 2158–2170.
  41. 41.Philippe Tillet, Hsiang-Tsung Kung, and David Cox. 2019. Triton: an intermediate language and compiler for tiled neural network computations. In Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages, pages 10–19.
  42. 42.Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Well-read students learn better: On the importance of pre-training compact models. arXiv preprint arXiv:1908.08962.
  43. 43.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems (NIPS), 30:5998–6008.
  44. 44.Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. In Association for Computational Linguistics (ACL), pages 5797–5808.
  45. 45.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2019. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In International Conference on Learning Representations (ICLR).
  46. 46.Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020a. MiniLM: Deep self-attention distillation for task-agnostic compression of pre-trained Transformers. In Advances in Neural Information Processing Systems (NeurIPS).
  47. 47.Ziheng Wang, Jeremy Wohlwend, and Tao Lei. 2020b. Structured pruning of large language models. In Empirical Methods in Natural Language Processing (EMNLP), pages 6151–6162.
  48. 48.Alex Warstadt, Amanpreet Singh, and Samuel R Bowman. 2019. Neural network acceptability judgments. In Transactions of the Association of Computational Linguistics (TACL), volume 7, pages 625–641.
  49. 49.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT).
  50. 50.Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin. 2020. DeeBERT: Dynamic early exiting for accelerating BERT inference. In Association for Computational Linguistics (ACL), pages 2246–2251.
  51. 51.Dongkuan Xu, Ian En-Hsu Yen, Jinxi Zhao, and Zhibin Xiao. 2021. Rethinking network pruning–under the pre-train and fine-tune paradigm. In North American Chapter of the Association for Computational Linguistics (NAACL), pages 2376–2382.
  52. 52.Zhewei Yao, Linjian Ma, Sheng Shen, Kurt Keutzer, and Michael W Mahoney. 2021. MLPruning: A multilevel structured pruning framework for transformer-based models. arXiv preprint arXiv:2105.14636.
  53. 53.Yichun Yin, Cheng Chen, Lifeng Shang, Xin Jiang, Xiao Chen, and Qun Liu. 2021. AutoTinyBERT: Automatic hyper-parameter optimization for efficient pre-trained language models. In Association for Computational Linguistics (ACL), pages 5146–5157.
  54. 54.Ofir Zafrir, Ariel Larey, Guy Boudoukh, Haihao Shen, and Moshe Wasserblat. 2021. Prune once for all: Sparse pre-trained language models. arXiv preprint arXiv:2111.05754.
  55. 55.Hattie Zhou, Janice Lan, Rosanne Liu, and Jason Yosinski. 2019. Deconstructing lottery tickets: Zeros, signs, and the supermask. In Advances in Neural Information Processing Systems (NeurIPS), volume 32. Curran Associates, Inc.

Citation

MLA
Xia, M., et al. “Structured Pruning Learns Compact and Accurate Models”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 1513–28, https://doi.org/10.18653/v1/2022.acl-long.107.
APA
Xia, M., Zhong, Z., & Chen, D. (2022). Structured Pruning Learns Compact and Accurate Models. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1513–1528. https://doi.org/10.18653/v1/2022.acl-long.107
Chicago
Xia, M., Z. Zhong, and D. Chen. 2022. “Structured Pruning Learns Compact and Accurate Models”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1513–28. https://doi.org/10.18653/v1/2022.acl-long.107.
Harvard
Xia, M., Zhong, Z. and Chen, D. (2022) “Structured Pruning Learns Compact and Accurate Models”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 1513–1528. Available at: https://doi.org/10.18653/v1/2022.acl-long.107.
Vancouver
1. Xia M, Zhong Z, Chen D (2022) Structured Pruning Learns Compact and Accurate Models. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 1513–1528

BibTeX

@inproceedings{xia-etal-2022-structured,
    title = "Structured Pruning Learns Compact and Accurate Models",
    author = "Xia, Mengzhou  and
      Zhong, Zexuan  and
      Chen, Danqi",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.107/",
    doi = "10.18653/v1/2022.acl-long.107",
    pages = "1513--1528"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/