Towards Efficient NLP: A Standard Evaluation and A Strong Baseline

Xiangyang LiuTianxiang SunJunliang HeJiawen WuLingling WuXinyu ZhangHao JiangZhao CaoXuanjing HuangXipeng Qiu

article2022NAACL61 citations

Establishes ELUE, a standardized evaluation benchmark with a public leaderboard to measure Pareto improvements across accuracy and computational cost, while introducing ElasticBERT as a strong baseline capable of both static and dynamic early exiting across any layer.

Listen

Modern natural language processing relies heavily on large pre-trained language models that achieve high accuracy at the cost of immense computational expense and slow execution. Consequently, the research focus has increasingly shifted toward model efficiency and practical usability. However, existing evaluation benchmarks are primarily designed to reward raw accuracy, lacking standardized measurements for efficiency metrics such as computational operations and model size. Furthermore, common baseline models used to validate efficient techniques remain weak or introduce significant discrepancies between how a model is pre-trained and fine-tuned.

The article introduces two core solutions to establish a standardized, multi-dimensional assessment of language models: the Efficient Language Understanding Evaluation (ELUE) benchmark and ElasticBERT, a versatile baseline model designed for flexible deployment.

The authors constructed the ELUE benchmark across six standard language understanding datasets spanning sentiment analysis, natural language inference, and semantic similarity. To measure efficiency consistently without reliance on variable hardware, the platform evaluates submitted models based on total parameter counts and theoretical floating-point operations (FLOPs). Alongside the benchmark, the authors developed ElasticBERT, a transformer model trained on roughly 160 gigabytes of text using multi-exit training across all intermediate layers. They incorporated a gradient equilibrium strategy to balance layer learning and a grouped training approach to optimize computational efficiency during pre-training, evaluating both static layer-pruning configurations and dynamic early-exiting inference against several existing baselines.

The evaluations yielded several key findings regarding model efficiency and performance trade-offs. ElasticBERT demonstrates superior resilience when reduced in depth; for example, a six-layer ElasticBERT achieves an average score of 89.4% on ELUE tasks, outperforming comparable six-layer variants of conventional models like BERT (86.5%) and compressed alternatives like DistilBERT (86.9%). When deployed dynamically to exit early on simpler inputs, ElasticBERT achieves the best efficiency-to-performance trade-off across the benchmark. Additionally, methodological ablations revealed that the grouped training strategy reduced pre-training compute time by approximately 43% without degrading internal layer accuracy.

These findings indicate that models designed with built-in elasticity can significantly lower the operational costs and latency of language processing tasks without sacrificing accuracy. For industry leaders and engineering teams, this provides a clear pathway to reduce infrastructure expenses, lower energy footprints, and deploy responsive models on resource-constrained environments. Rather than relying on separate, fragmented compression methods, teams can train a single multi-exit backbone that dynamically adapts to various performance requirements.

Organizations developing or deploying language technologies should adopt standardized multi-dimensional metrics, such as FLOPs and parameter limits, when evaluating model deployments rather than focusing solely on peak accuracy. Practitioners seeking efficient architectures should consider multi-exit structures like ElasticBERT as strong starting baselines. As next steps, the evaluation framework can be expanded to encompass broader deep learning toolkits, specialized hardware constraints, and generative language tasks.

The conclusions are supported by evaluations across six standardized classification and regression datasets. However, certain limitations remain: physical runtimes may still vary across specific hardware implementations, and the benchmark currently focuses primarily on sentence classification and semantic matching rather than text generation or extremely long context inputs. Nonetheless, the evidence strongly supports the validity of multi-dimensional Pareto benchmarking and multi-exit model architectures.

Cover for Towards Efficient NLP: A Standard Evaluation and A Strong Baseline

Abstract

Supersized pre-trained language models have pushed the accuracy of various natural language processing (NLP) tasks to a new state-of-the-art (SOTA). Rather than pursuing the reachless SOTA accuracy, more and more researchers start paying attention to model efficiency and usability. Different from accuracy, the metric for efficiency varies across different studies, making them hard to be fairly compared. To that end, this work presents ELUE (Efficient Language Understanding Evaluation), a standard evaluation, and a public leaderboard for efficient NLP models. ELUE is dedicated to depicting the Pareto Frontier for various language understanding tasks, such that it can tell whether and how much a method achieves Pareto improvement. Along with the benchmark, we also release a strong baseline, ElasticBERT, which allows BERT to exit at any layer in both static and dynamic ways. We demonstrate the ElasticBERT, despite its simplicity, outperforms or performs on par with SOTA compressed and early exiting models. With ElasticBERT, the proposed ELUE has a strong Pareto Frontier and makes a better evaluation for efficient NLP models.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 ELUE: A Standard Benchmark for Efficient NLP Models
  • 3.1 Design Considerations
  • 3.2 Task and Dataset Selection
  • 3.3 Leaderboard
  • 4 ElasticBERT: A Strong Baseline for Efficient Inference
  • 5 Experiments
  • 5.1 Experimental Setup
  • 5.2 Evaluating ElasticBERT on ELUE
  • 6 Conclusion and Future Work
  • Acknowledgment
  • Ethical Considerations
  • References
  • Appendix
  • A Details of Evaluation
  • A.1 Submission and Evaluation
  • B Experimental Details and Additional Results
  • B.1 Details of Training ElasticBERT
  • B.2 Evaluating ElasticBERT on GLUE
  • B.3 Ablation Study
  • B.4 Overall Comparison
  • C ELUE Website

Knowls

  1. Knowl 1 — ELUE Benchmark for Efficient NLP

    model/method

    The Efficient Language Understanding Evaluation (ELUE) benchmark is a standardized evaluation platform designed to assess both the task accuracy and computational efficiency (FLOPs and parameter count) of Natural Language Processing (NLP) models. ELUE depicts Pareto frontiers across six language understanding datasets spanning three core task categories:

    1. Sentiment Analysis: SST-2 (Stanford Sentiment Treebank binary classification using complete sentences rather than sub-phrases; 8,544 train, 1,101 dev, 2,208 test) and IMDb (movie review polarity classification; 20,000 train, 5,000 dev, 25,000 test).
    2. Natural Language Inference: SNLI (3-way entailment classification using spell-checked splits with unagreed examples omitted; 549,367 train, 9,842 dev, 9,824 test) and SciTail (binary entailment classification derived from science exams; 23,596 train, 1,304 dev, 2,126 test).
    3. Similarity and Paraphrase: MRPC (binary paraphrase identification; 3,668 train, 408 dev, 1,725 test) and STS-B (semantic textual similarity regression scored from 0 to 5; 5,749 train, 1,500 dev, 1,379 test).

    Evaluation metrics are accuracy for SST-2, IMDb, SNLI, and SciTail; the unweighted average of accuracy and F1 score for MRPC; and the average of Pearson and Spearman correlation coefficients for STS-B. To provide fair comparisons across model sizes, the benchmark maintains dedicated parameter-budget tracks for models below 40M, 55M, 70M, and 110M parameters.

  2. Knowl 2 — ELUE Score Formulation

    equation

    To measure the degree to which a model submission (such as dynamic early-exiting models that generate multiple efficiency-performance trade-off operating points) advances beyond the baseline Pareto frontier, ELUE defines the ELUE Score relative to an ElasticBERT baseline curve.

    For a specific dataset, let (piEB,fiEB)i=112(p_i^{\text{EB}}, f_i^{\text{EB}})_{i=1}^{12} be the performance values piEBp_i^{\text{EB}} and FLOPs fiEBf_i^{\text{EB}} evaluated at each of the 12 exit layers of a 12-layer ElasticBERT baseline. Linear interpolation over these points yields a continuous baseline performance function pEB(f)p^{\text{EB}}(f) parameterized by FLOPs ff. For a submitted system evaluated at nn operating points (pi,fi)i=1n(p_i, f_i)_{i=1}^n on that dataset, the task-level ELUE score is:

    ELUEScore=1n∑i=1n[pi−pEB(fi)]\text{ELUEScore} = \frac{1}{n} \sum_{i=1}^n \left[ p_i - p^{\text{EB}}(f_i) \right]

    where pip_i is the model's performance at operating point ii, fif_i is the computational cost in FLOPs at point ii, and pEB(fi)p^{\text{EB}}(f_i) is the interpolated baseline performance of ElasticBERT at FLOPs fif_i. The overall ELUE benchmark score is the unweighted arithmetic mean of the task scores across all six ELUE datasets. A positive score indicates a Pareto improvement over ElasticBERT.

  3. Knowl 3 — Standardized FLOPs and Activation-Trace Evaluation Framework in ELUE

    model/method

    To avoid hardware-dependent and software-dependent variability in physical execution time measurements, ELUE standardizes efficiency evaluation through deterministic static parameter counts and exact per-sample floating-point operations (FLOPs).

    Submissions provide two components:

    1. Prediction and Activation Trace Files: Prediction files where each entry records the sample index, predicted label, and an ordered modules string listing every neural module invoked for that sample along with its input tensor dimensions (e.g., (10),emb; (10,768),layer_1; (768),exit_1). This format natively supports static models, sample-level early-exiting classifiers, and token-level length-reduction methods.
    2. Model Definition File: A Python script defining the PyTorch module architecture.

    The benchmark server executes a unified profiling script against the submitted module dictionary and activation traces to compute average per-sample FLOPs and total parameter counts.

  4. Knowl 4 — ElasticBERT Architecture and Multi-Exit Training Objective

    model/method

    ElasticBERT is a multi-exit Transformer model designed to serve as an elastic static baseline (which can be truncated to any intermediate depth ll without fine-tuning performance collapse) and as a backbone for dynamic early exiting.

    Rather than attaching pre-training heads solely to the final Transformer layer, ElasticBERT places auxiliary prediction heads at every intermediate layer l∈{1,2,…,L}l \in \{1, 2, \dots, L\}. Pre-training optimizes the joint loss across all exits:

    L=∑l=1L(LlMLM+LlSOP)\mathcal{L} = \sum_{l=1}^L \left( \mathcal{L}_l^{\text{MLM}} + \mathcal{L}_l^{\text{SOP}} \right)

    where LlMLM\mathcal{L}_l^{\text{MLM}} is the nn-gram masked language modeling loss computed on the output representations of layer ll, and LlSOP\mathcal{L}_l^{\text{SOP}} is the sentence order prediction loss computed on the pooled representation at layer ll.

    When pruned to ll layers for static deployment, ElasticBERT acts as a complete standalone model. When used dynamically, intermediate classifiers can be fine-tuned without introducing a mismatch between pre-training and downstream inference.

  5. Knowl 5 — Grouped Multi-Exit Pre-Training Algorithm

    algorithm

    Summing losses across all intermediate exits during every backward step increases GPU memory consumption and slows pre-training. ElasticBERT addresses this by partitioning the LL layer exits into GG groups, cycling through one group per training batch while including the top exit layer LL in every group to retain overall model capacity.

    Input: Total layers LL, number of groups GG, exit partitions {G1,G2,…,GGG_1, G_2, \dots, G_G}, training batches {B1,B2,…B_1, B_2, \dots}
    Output: Pre-trained ElasticBERT parameters Θ\Theta
    Initialize parameters Θ\Theta with standard BERT checkpoint weights
    t←0t \leftarrow 0
    while training step t<MaxStepst < \text{MaxSteps} do
        Sample training batch BtB_t
        Select active exit group index i←(t mod G)+1i \leftarrow (t \bmod G) + 1
        Compute total group loss on batch BtB_t:
            L=∑l∈Gi(LlMLM(Bt;Θ)+LlSOP(Bt;Θ))\mathcal{L} = \sum_{l \in G_i} \left( \mathcal{L}_l^{\text{MLM}}(B_t; \Theta) + \mathcal{L}_l^{\text{SOP}}(B_t; \Theta) \right)
        Apply Gradient Equilibrium rescaling to layer gradients
        Update parameters Θ\Theta using Adam optimizer on ∇ΘL\nabla_\Theta \mathcal{L}
        t←t+1t \leftarrow t + 1
    end while
    return Θ\Theta

    For 12-layer ElasticBERT_BASE (L=12,G=2L=12, G=2):

    • G1={1,3,5,7,9,11,12}G_1 = \{1, 3, 5, 7, 9, 11, 12\}
    • G2={2,4,6,8,10,12}G_2 = \{2, 4, 6, 8, 10, 12\}

    For 24-layer ElasticBERT_LARGE (L=24,G=3L=24, G=3):

    • G1={1,4,7,10,13,16,19,22,24}G_1 = \{1, 4, 7, 10, 13, 16, 19, 22, 24\}
    • G2={2,5,8,11,14,17,20,23,24}G_2 = \{2, 5, 8, 11, 14, 17, 20, 23, 24\}
    • G3={3,6,9,12,15,18,21,24}G_3 = \{3, 6, 9, 12, 15, 18, 21, 24\}

    Grouped training reduces pre-training duration by approximately 43% (from 186.0 hours to 106.0 hours on identical hardware) without degrading intermediate exit accuracy.

  6. Knowl 6 — Gradient Equilibrium for Multi-Exit Training

    model/method

    In multi-exit Transformer training, straightforward loss summation across all exits causes lower layers to receive overlapping backward gradient signals from all upper exits. This leads to large gradient variance and unstable optimization.

    ElasticBERT incorporates the Gradient Equilibrium (GE) strategy during pre-training and downstream multi-exit fine-tuning. For any exit loss Lj\mathcal{L}_j attached to layer jj, the gradient ∇θiLj\nabla_{\theta_i} \mathcal{L}_j with respect to the parameters θi\theta_i of an earlier layer i<ji < j is rescaled. This stabilizes parameter updates and balances the gradient magnitudes across lower and higher Transformer layers.

  7. Knowl 7 — ElasticBERT Pre-Training and Fine-Tuning Experimental Setup

    experimental setup

    ElasticBERT_BASE (12 layers, hidden size 768, 12 attention heads, intermediate size 3072, 109M parameters) and ElasticBERT_LARGE (24 layers, hidden size 1024, 16 attention heads, intermediate size 4096, 335M parameters) are initialized with BERT weights.

    Pre-training Setup:

    • Corpus: ~160GB uncompressed English text comprising English Wikipedia (12GB), BookCorpus (4GB), OpenWebText (38GB), and 110GB of the C4 corpus.
    • Optimizer: Adam (β1=0.9,β2=0.999\beta_1 = 0.9, \beta_2 = 0.999, weight decay =0.01= 0.01, peak learning rate =2×10−4= 2 \times 10^{-4} with linear warmup over the first 4% of steps).
    • Schedule & Hardware: Batch size 4,096 sequences, trained for 125,000 steps on 64 ×\times 32GB NVIDIA Tesla V100 GPUs using Megatron-LM.

    Downstream Fine-Tuning Setup:

    • Static models: Fine-tuned for 10 epochs with early stopping using AdamW, batch size 32, learning rates in {1×10−5,2×10−5,3×10−5}\{1 \times 10^{-5}, 2 \times 10^{-5}, 3 \times 10^{-5}\}, with a 6% linear warmup.
    • Dynamic early exit models: Evaluated with entropy-based early exiting (thresholding output entropy) and patience-based early exiting (halting upon consecutive identical layer predictions). Models using two-stage training are trained for 3 epochs per stage; single-stage dynamic models are trained for 5 epochs using AdamW.
  8. Knowl 8 — Static Model Performance on the ELUE Benchmark

    data/table

    Performance of ElasticBERT compared to full and truncated pre-trained models (BERT, ALBERT, RoBERTa, LayerDrop) and dedicated compressed models (DistilBERT, TinyBERT, BERT-of-Theseus, HeadPrune) on the ELUE task test sets. MRPC reports the mean of Accuracy and F1; STS-B reports Pearson and Spearman correlation; other tasks report Accuracy. FLOPs are averages across all datasets.

    Model #Params #FLOPs SST-2 IMDb MRPC STS-B SNLI SciTail Average
    BASE Models
    BERTBASE\text{BERT}_{\text{BASE}} 109M 13399M 85.1 93.0 83.1 84.2 90.4 93.2 88.2
    ALBERTBASE\text{ALBERT}_{\text{BASE}} 12M 13927M 86.6 92.9 87.8 88.3 90.1 93.4 89.9
    RoBERTaBASE\text{RoBERTa}_{\text{BASE}} 125M 13103M 88.3 94.9 88.0 89.6 91.3 92.8 90.8
    LayerDropBASE\text{LayerDrop}_{\text{BASE}} 125M 13103M 88.5 94.2 88.2 87.1 90.7 92.8 90.3
    ElasticBERTBASE\text{ElasticBERT}_{\text{BASE}} 109M 13399M 88.6 93.9 87.9 87.6 91.3 93.8 90.5
    BERTBASE-6L\text{BERT}_{\text{BASE}}\text{-6L} 67M 6700M 83.3 91.0 82.6 82.5 88.9 90.7 86.5
    ALBERTBASE-6L\text{ALBERT}_{\text{BASE}}\text{-6L} 12M 6972M 84.7 92.0 85.3 83.5 89.3 92.3 87.9
    RoBERTaBASE-6L\text{RoBERTa}_{\text{BASE}}\text{-6L} 82M 6552M 86.8 92.6 86.7 84.5 90.2 91.3 88.7
    LayerDropBASE-6L\text{LayerDrop}_{\text{BASE}}\text{-6L} 82M 6552M 86.3 92.9 86.3 86.1 89.5 90.3 88.6
    HeadPrune-BERTBASE\text{HeadPrune-BERT}_{\text{BASE}} 86M 9249M 84.8 84.7 77.8 74.8 87.8 88.3 83.0
    DistilBERT 67M 6700M 84.8 92.0 83.8 81.7 89.2 89.7 86.9
    TinyBERT-6L 67M 6700M 85.3 89.0 86.2 85.7 89.3 90.0 87.6
    BERT-of-Theseus 67M 6700M 84.4 90.7 82.4 85.0 89.4 92.1 87.3
    ElasticBERTBASE-6L\text{ElasticBERT}_{\text{BASE}}\text{-6L} 67M 6700M 87.0 92.7 87.3 86.9 90.1 92.5 89.4
    LARGE Models
    BERTLARGE\text{BERT}_{\text{LARGE}} 335M 47214M 87.9 94.0 85.9 86.7 90.8 93.9 89.9
    ALBERTLARGE\text{ALBERT}_{\text{LARGE}} 18M 48876M 87.7 93.8 88.1 89.3 90.2 93.6 90.5
    RoBERTaLARGE\text{RoBERTa}_{\text{LARGE}} 355M 46042M 90.5 95.7 89.9 90.5 91.6 95.8 92.3
    LayerDropLARGE\text{LayerDrop}_{\text{LARGE}} 355M 46042M 90.4 95.3 89.5 91.0 91.4 95.2 92.1
    ElasticBERTLARGE\text{ElasticBERT}_{\text{LARGE}} 335M 47214M 89.8 95.0 89.8 90.9 91.4 95.7 92.1
    BERTLARGE-6L\text{BERT}_{\text{LARGE}}\text{-6L} 108M 11922M 80.4 89.6 74.3 70.5 87.4 84.4 81.1
    ALBERTLARGE-6L\text{ALBERT}_{\text{LARGE}}\text{-6L} 18M 12397M 84.5 92.0 84.7 85.1 89.4 90.8 87.8
    RoBERTaLARGE-6L\text{RoBERTa}_{\text{LARGE}}\text{-6L} 129M 11664M 83.5 91.7 77.9 72.7 88.6 84.7 83.2
    LayerDropLARGE-6L\text{LayerDrop}_{\text{LARGE}}\text{-6L} 129M 11664M 85.4 92.5 77.3 75.9 88.8 84.1 84.0
    ElasticBERTLARGE-6L\text{ElasticBERT}_{\text{LARGE}}\text{-6L} 108M 11922M 86.8 92.9 86.2 86.3 89.8 92.4 89.1

    When compressed to 6 layers, ElasticBERT_BASE-6L achieves an average of 89.4, outperforming dedicated distillation and replacement methods (DistilBERT at 86.9, TinyBERT-6L at 87.6, and BERT-of-Theseus at 87.3). In 6-layer LARGE configurations, ElasticBERT_LARGE-6L achieves 89.1, outperforming RoBERTa_LARGE-6L (83.2) and LayerDrop_LARGE-6L (84.0).

  9. Knowl 9 — ELUE Leaderboard Scores Across Baselines

    data/table

    ELUE benchmark scores across static baselines and dynamic early-exiting methods, evaluating the average performance differential relative to the 12-layer ElasticBERT baseline curve. ElasticBERT has a reference score of 0.00.

    Model SST-2 IMDb MRPC STS-B SNLI SciTail Average
    ElasticBERTBASE\text{ElasticBERT}_{\text{BASE}} 0.00 0.00 0.00 0.00 0.00 0.00 0.00
    Static Models
    BERTBASE\text{BERT}_{\text{BASE}} -4.55 -2.15 -5.88 -4.75 -1.50 -3.35 -3.70
    ALBERTBASE\text{ALBERT}_{\text{BASE}} -2.41 -1.08 -2.34 -2.81 -1.55 -1.50 -1.95
    RoBERTaBASE\text{RoBERTa}_{\text{BASE}} -0.89 -0.11 -2.95 -5.38 -0.66 -3.32 -2.22
    LayerDropBASE\text{LayerDrop}_{\text{BASE}} -1.17 -0.13 -2.17 -2.98 -1.36 -4.14 -1.99
    HeadPrune-BERTBASE\text{HeadPrune-BERT}_{\text{BASE}} -3.81 -8.61 -9.73 -11.9 -2.89 -4.18 -6.85
    DistilBERT -2.20 -0.70 -3.50 -5.20 -0.90 -2.80 -2.55
    TinyBERT-6L -1.70 -3.70 -2.60 -1.90 -0.80 -2.50 -2.20
    BERT-of-Theseus -4.21 -2.61 -5.13 -1.67 -1.29 -0.38 -2.55
    Dynamic Models
    PABEE -1.33 -0.23 -2.93 -2.13 -0.85 -0.43 -1.50
    DeeBERT -12.1 -14.0 -4.88 - -8.35 -6.19 -
    FastBERT -1.51 0.16 -3.70 - -0.22 -1.23 -
    CascadeBERT -2.13 -0.12 -4.05 - -0.23 0.14 -
    ElasticBERTpatience\text{ElasticBERT}_{\text{patience}} 0.40 0.20 -1.00 -0.44 0.03 0.36 -0.08
    ElasticBERTentropy\text{ElasticBERT}_{\text{entropy}} 0.97 1.02 -0.14 - 0.02 0.64 -

    Dynamic models utilizing ElasticBERT as a backbone (ElasticBERT_patience and ElasticBERT_entropy) outperform earlier early-exit methods (PABEE, DeeBERT, FastBERT, CascadeBERT). ElasticBERT_patience obtains an overall score of -0.08 and achieves positive Pareto gains over ElasticBERT's static curve on SST-2 (+0.40), IMDb (+0.20), SNLI (+0.03), and SciTail (+0.36).

  10. Knowl 10 — Downstream Evaluation on GLUE Benchmark

    data/table

    Performance of ElasticBERT and static baselines on the GLUE benchmark. MRPC reports the mean of Accuracy and F1; STS-B reports Pearson and Spearman correlation; CoLA reports Matthews correlation; other tasks report Accuracy.

    Model #Params #FLOPs CoLA MNLI-m/mm MRPC QNLI QQP RTE SST-2 STS-B Average
    BASE Models (Development Set)
    BERTBASE\text{BERT}_{\text{BASE}} 109M 6615M 56.5 84.6/84.9 87.6 91.2 89.6 69.0 92.9 89.4 82.9
    ALBERTBASE\text{ALBERT}_{\text{BASE}} 12M 6861M 56.8 84.9/85.6 90.5 91.4 89.2 78.3 92.8 90.7 84.5
    RoBERTaBASE\text{RoBERTa}_{\text{BASE}} 125M 6727M 63.6 87.5/87.2 90.8 92.7 90.3 77.5 94.8 90.9 86.1
    LayerDropBASE\text{LayerDrop}_{\text{BASE}} 125M 6727M 64.5 86.4/86.5 91.6 92.2 89.9 71.1 93.7 88.6 84.9
    ElasticBERTBASE\text{ElasticBERT}_{\text{BASE}} 109M 6615M 64.3 85.3/85.9 91.0 92.0 90.2 76.5 94.3 90.7 85.6
    BERTBASE-6L\text{BERT}_{\text{BASE}}\text{-6L} 67M 3308M 44.6 81.4/81.4 84.9 87.4 88.7 65.7 90.9 88.1 79.2
    ALBERTBASE-6L\text{ALBERT}_{\text{BASE}}\text{-6L} 12M 3435M 52.4 82.6/82.2 89.0 89.8 88.7 70.4 90.8 89.6 81.7
    RoBERTaBASE-6L\text{RoBERTa}_{\text{BASE}}\text{-6L} 82M 3364M 44.4 84.2/84.6 87.9 90.5 89.8 60.6 92.1 89.0 80.3
    LayerDropBASE-6L\text{LayerDrop}_{\text{BASE}}\text{-6L} 82M 3364M 53.7 83.8/83.8 87.6 89.8 89.4 64.3 91.3 88.1 81.3
    HeadPrune-BERTBASE\text{HeadPrune-BERT}_{\text{BASE}} 87M 4744M 48.7 71.0/79.7 80.2 86.1 84.7 62.5 89.4 85.2 76.4
    DistilBERT 67M 3308M 55.6 82.1/82.0 86.5 89.2 88.8 63.9 91.3 86.7 80.7
    TinyBERT-6L 67M 3308M 46.3 83.6/83.8 88.7 90.6 89.1 73.6 92.0 89.4 81.9
    BERT-of-Theseus 67M 3308M 45.1 81.4/81.9 88.1 88.1 88.9 70.1 91.4 88.8 80.4
    ElasticBERTBASE-6L\text{ElasticBERT}_{\text{BASE}}\text{-6L} 67M 3308M 53.7 84.3/84.2 89.7 90.8 89.7 74.0 92.7 90.2 83.3
    BASE Models (Test Set Results)
    TinyBERT-6L 67M 3308M 42.5 83.2/82.4 86.2 89.6 79.6 73.0 91.8 85.7 79.3
    ElasticBERTBASE-6L\text{ElasticBERT}_{\text{BASE}}\text{-6L} 67M 3308M 49.1 83.7/83.4 87.3 90.4 79.7 68.7 92.9 86.9 80.3
    LARGE Models (Development Set)
    BERTLARGE\text{BERT}_{\text{LARGE}} 335M 23446M 61.6 86.2/86.0 90.1 92.2 90.1 72.9 93.5 90.4 84.8
    ALBERTLARGE\text{ALBERT}_{\text{LARGE}} 18M 24296M 60.1 86.0/86.1 90.4 91.6 89.6 83.0 95.2 91.4 85.9
    RoBERTaLARGE\text{RoBERTa}_{\text{LARGE}} 355M 23840M 66.4 89.0/89.6 91.6 94.2 90.7 86.6 95.4 92.3 88.4
    LayerDropLARGE\text{LayerDrop}_{\text{LARGE}} 355M 23840M 66.6 89.7/89.6 91.2 93.9 88.5 86.6 95.5 92.6 88.2
    ElasticBERTLARGE\text{ElasticBERT}_{\text{LARGE}} 335M 23446M 66.3 88.0/88.5 92.0 93.6 90.9 83.1 95.3 91.7 87.7
    BERTLARGE-6L\text{BERT}_{\text{LARGE}}\text{-6L} 108M 5863M 20.2 76.5/76.5 76.4 84.3 87.3 58.5 89.7 77.3 78.5
    ALBERTLARGE-6L\text{ALBERT}_{\text{LARGE}}\text{-6L} 18M 6083M 51.7 82.2/82.9 86.5 89.4 88.6 66.4 92.2 89.4 81.0
    RoBERTaLARGE-6L\text{RoBERTa}_{\text{LARGE}}\text{-6L} 129M 5962M 43.3 80.4/80.9 80.0 86.1 88.9 54.9 90.1 80.5 76.1
    LayerDropLARGE-6L\text{LayerDrop}_{\text{LARGE}}\text{-6L} 129M 5962M 44.3 81.4/81.0 79.8 87.1 88.5 53.1 91.4 83.0 76.6
    ElasticBERTLARGE-6L\text{ElasticBERT}_{\text{LARGE}}\text{-6L} 108M 5863M 53.9 83.5/84.3 89.6 90.8 90.1 71.1 91.9 90.1 82.8
    LARGE Models (Test Set Results)
    ALBERTLARGE-6L\text{ALBERT}_{\text{LARGE}}\text{-6L} 18M 6083M 46.5 81.9/82.2 84.7 88.5 78.9 62.3 91.3 85.1 77.9
    ElasticBERTLARGE-6L\text{ElasticBERT}_{\text{LARGE}}\text{-6L} 108M 5863M 47.2 83.2/82.6 86.2 90.4 80.2 67.0 92.5 86.3 79.5

    On the GLUE development set, ElasticBERT_BASE-6L obtains an average score of 83.3, outperforming TinyBERT-6L (81.9) and DistilBERT (80.7). On the official GLUE test sets, ElasticBERT_BASE-6L reaches 80.3 vs. 79.3 for TinyBERT-6L, and ElasticBERT_LARGE-6L reaches 79.5 vs. 77.9 for ALBERT_LARGE-6L.

  11. Knowl 11 — Ablations on Multi-Exit Training Strategies and Exit Grouping

    empirical result

    Ablation experiments evaluate the components of ElasticBERT's training procedure:

    1. Training Strategy Comparison: Comparing Gradient Equilibrium (GE) against two-stage training (freezing the backbone to train internal classifiers) and depth-weighted training (loss weighted by layer depth ll via L=∑l=1Ll⋅Ll/∑l=1Ll\mathcal{L} = \sum_{l=1}^L l \cdot \mathcal{L}_l / \sum_{l=1}^L l) shows that GE consistently achieves the highest classification performance across intermediate exit layers on both SST-2 and MRPC.
    2. Contribution of Gradient Equilibrium: In pre-training ElasticBERT_BASE on Wikipedia and BookCorpus, incorporating GE improves the 12-layer average score across MNLI, MRPC, QNLI, and QQP from 88.3 to 88.4 (MNLI: 85.4/85.7 vs. 85.3/85.1) and improves the 6-layer average score from 86.6 to 87.0.
    3. Exit Grouping Topology: Evaluating average MultiNLI exit accuracy under different exit partitioning schemes for a 12-exit network shows that the alternating odd/even partition (G1={1,3,5,7,9,11,12},G2={2,4,6,8,10,12}G_1=\{1, 3, 5, 7, 9, 11, 12\}, G_2=\{2, 4, 6, 8, 10, 12\}) matches full ungrouped pre-training accuracy at 76.7%. Alternative groupings perform lower: 3-group alternating (G1={1,4,7,10,12},G2={2,5,8,11,12},G3={3,6,9,12}G_1=\{1,4,7,10,12\}, G_2=\{2,5,8,11,12\}, G_3=\{3,6,9,12\}) achieves 75.7%, 3-group block partitioning achieves 75.5%, and 2-group block partitioning (G1={1..6,12},G2={7..12}G_1=\{1..6,12\}, G_2=\{7..12\}) achieves 75.9%.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek Gordon Murray, Benoit Steiner, Paul A. Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016. Tensorflow: A system for largescale machine learning. In 12th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2016, Savannah, GA, USA, November 2-4, 2016, pages 265–283. USENIX Association.
  2. 2.Yoshua Bengio, Nicholas Léonard, and Aaron C. Courville. 2013. Estimating or propagating gradients through stochastic neurons for conditional computation. CoRR, abs/1308.3432.
  3. 3.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, EMNLP 2015, Lisbon, Portugal, September 17-21, 2015, pages 632–642. The Association for Computational Linguistics.
  4. 4.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  5. 5.Daniel M. Cer, Mona T. Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017. Semeval2017 task 1: Semantic textual similarity - multilingual and cross-lingual focused evaluation. CoRR, abs/1708.00055.
  6. 6.Alexis Conneau and Douwe Kiela. 2018. Senteval: An evaluation toolkit for universal sentence representations. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation, LREC 2018, Miyazaki, Japan, May 7-12, 2018. European Language Resources Association (ELRA).
  7. 7.Andrew S. Davis and Itamar Arel. 2014. Low-rank approximations for conditional feedforward computation in deep neural networks. In 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Workshop Track Proceedings.
  8. 8.Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. 2019. Universal transformers. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 4171–4186. Association for Computational Linguistics.
  10. 10.William B. Dolan and Chris Brockett. 2005. Automatically constructing a corpus of sentential paraphrases. In Proceedings of the Third International Workshop on Paraphrasing, IWP@IJCNLP 2005, Jeju Island, Korea, October 2005, 2005. Asian Federation of Natural Language Processing.
  11. 11.Maha Elbayad, Jiatao Gu, Edouard Grave, and Michael Auli. 2020. Depth-adaptive transformer. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  12. 12.Angela Fan, Edouard Grave, and Armand Joulin. 2020. Reducing transformer depth on demand with structured dropout. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  13. 13.Sebastian Gehrmann, Tosin P. Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Aremu Anuoluwapo, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh D. Dhole, Wanyu Du, Esin Durmus, Ondrej Dusek, Chris Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Rubungo Andre Niyongabo, Salomey Osei, Ankur P. Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou. 2021. The GEM benchmark: Natural language generation, its evaluation and metrics. CoRR, abs/2102.01672.
  14. 14.Aaron Gokaslan and Vanya Cohen. 2019. Openwebtext corpus.
  15. 15.Mitchell A. Gordon, Kevin Duh, and Nicholas Andrews. 2020. Compressing BERT: studying the effects of weight pruning on transfer learning. In Proceedings of the 5th Workshop on Representation Learning for NLP, RepL4NLP@ACL 2020, Online, July 9, 2020, pages 143–155. Association for Computational Linguistics.
  16. 16.Alex Graves. 2016. Adaptive computation time for recurrent neural networks. CoRR, abs/1603.08983.
  17. 17.Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020. Tinybert: Distilling BERT for natural language understanding. In Proceedings of EMNLP 2020, pages 4163–4174. Association for Computational Linguistics.
  18. 18.Tushar Khot, Ashish Sabharwal, and Peter Clark. 2018. Scitail: A textual entailment dataset from science question answering. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18), the 30th innovative Applications of Artificial Intelligence (IAAI-18), and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI-18), New Orleans, Louisiana, USA, February 2-7, 2018, pages 5189–5197. AAAI Press.
  19. 19.Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams. 2021. Dynabench: Rethinking benchmarking in NLP. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, pages 4110–4124. Association for Computational Linguistics.
  20. 20.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  21. 21.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020. ALBERT: A lite BERT for self-supervised learning of language representations. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  22. 22.Hao Li, Hong Zhang, Xiaojuan Qi, Ruigang Yang, and Gao Huang. 2019. Improved techniques for training adaptive deep networks. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019, pages 1891–1900. IEEE.
  23. 23.Lei Li, Yankai Lin, Deli Chen, Shuhuai Ren, Peng Li, Jie Zhou, and Xu Sun. 2021a. Cascadebert: Accelerating inference of pre-trained language models via calibrated complete models cascade. In Findings of EMNLP.
  24. 24.Xiaonan Li, Yunfan Shao, Tianxiang Sun, Hang Yan, Xipeng Qiu, and Xuanjing Huang. 2021b. Accelerating BERT inference for sequence labeling via early-exit. In Proceedings of ACL/IJCNLP 2021, pages 189–199. Association for Computational Linguistics.
  25. 25.Kaiyuan Liao, Yi Zhang, Xuancheng Ren, Qi Su, Xu Sun, and Bin He. 2021. A global past-future early exit method for accelerating inference of pre-trained language models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, pages 2013–2023. Association for Computational Linguistics.
  26. 26.Tianyang Lin, Yuxin Wang, Xiangyang Liu, and Xipeng Qiu. 2021. A survey of transformers. CoRR, abs/2106.04554.
  27. 27.Weijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao, Haotang Deng, and Qi Ju. 2020a. Fastbert: a self-distilling BERT with adaptive inference time. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 6035–6044. Association for Computational Linguistics.
  28. 28.Weijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao, Haotang Deng, and Qi Ju. 2020b. Fastbert: a self-distilling BERT with adaptive inference time. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 6035–6044. Association for Computational Linguistics.
  29. 29.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  30. 30.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  31. 31.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In The 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Proceedings of the Conference, 19-24 June, 2011, Portland, Oregon, USA, pages 142–150. The Association for Computer Linguistics.
  32. 32.Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher. 2018. The natural language decathlon: Multitask learning as question answering. CoRR, abs/1806.08730.
  33. 33.Paul Michel, Omer Levy, and Graham Neubig. 2019. Are sixteen heads really better than one? In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 14014–14024.
  34. 34.Sewon Min, Jordan L. Boyd-Graber, Chris Alberti, Danqi Chen, Eunsol Choi, Michael Collins, Kelvin Guu, Hannaneh Hajishirzi, Kenton Lee, Jennimaria Palomaki, Colin Raffel, Adam Roberts, Tom Kwiatkowski, Patrick S. H. Lewis, Yuxiang Wu, Heinrich Küttler, Linqing Liu, Pasquale Minervini, Pontus Stenetorp, Sebastian Riedel, Sohee Yang, Minjoon Seo, Gautier Izacard, Fabio Petroni, Lucas Hosseini, Nicola De Cao, Edouard Grave, Ikuya Yamada, Sonse Shimaoka, Masatoshi Suzuki, Shumpei Miyawaki, Shun Sato, Ryo Takahashi, Jun Suzuki, Martin Fajcik, Martin Docekal, Karel Ondrej, Pavel Smrz, Hao Cheng, Yelong Shen, Xiaodong Liu, Pengcheng He, Weizhu Chen, Jianfeng Gao, Barlas Oguz, Xilun Chen, Vladimir Karpukhin, Stan Peshterliev, Dmytro Okhonko, Michael Sejr Schlichtkrull, Sonal Gupta, Yashar Mehdad, and Wen-tau Yih. 2020. Neurips 2020 efficientqa competition: Systems, analyses and lessons learned. In NeurIPS 2020 Competition and Demonstration Track, 6-12 December 2020, Virtual Event / Vancouver, BC, Canada, volume 133 of Proceedings of Machine Learning Research, pages 86–111. PMLR.
  35. 35.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 8024–8035.
  36. 36.Xipeng Qiu, Tianxiang Sun, Yige Xu, Yunfan Shao, Ning Dai, and Xuanjing Huang. 2020. Pre-trained models for natural language processing: A survey. SCIENCE CHINA Technological Sciences.
  37. 37.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67.
  38. 38.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. Squad: 100, 000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP 2016, Austin, Texas, USA, November 1-4, 2016, pages 2383–2392. The Association for Computational Linguistics.
  39. 39.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter. CoRR, abs/1910.01108.
  40. 40.Roy Schwartz, Gabriel Stanovsky, Swabha Swayamdipta, Jesse Dodge, and Noah A. Smith. 2020. The right tool for the job: Matching model and instance complexities. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 6640–6651. Association for Computational Linguistics.
  41. 41.Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer. 2020. Q-BERT: hessian based ultra low precision quantization of BERT. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, New York, NY, USA, February 7-12, 2020, pages 8815–8821. AAAI Press.
  42. 42.Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-lm: Training multi-billion parameter language models using model parallelism. CoRR, abs/1909.08053.
  43. 43.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, EMNLP 2013, 18-21 October 2013, Grand Hyatt Seattle, Seattle, Washington, USA, A meeting of SIGDAT, a Special Interest Group of the ACL, pages 1631–1642. ACL.
  44. 44.Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019. Patient knowledge distillation for BERT model compression. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 4322–4331. Association for Computational Linguistics.
  45. 45.Tianxiang Sun, Xiangyang Liu, Xipeng Qiu, and Xuanjing Huang. 2021a. Paradigm shift in natural language processing. CoRR, abs/2109.12575.
  46. 46.Tianxiang Sun, Yunfan Shao, Xipeng Qiu, Qipeng Guo, Yaru Hu, Xuanjing Huang, and Zheng Zhang. 2020. Colake: Contextualized language and knowledge embedding. In Proceedings of the 28th International Conference on Computational Linguistics, COLING 2020, Barcelona, Spain (Online), December 8-13, 2020, pages 3660–3670. International Committee on Computational Linguistics.
  47. 47.Tianxiang Sun, Yunhua Zhou, Xiangyang Liu, Xinyu Zhang, Hao Jiang, Zhao Cao, Xuanjing Huang, and Xipeng Qiu. 2021b. Early exiting with ensemble internal classifiers. CoRR, abs/2105.13792.
  48. 48.Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler. 2021. Long range arena : A benchmark for efficient transformers. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  49. 49.Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Well-read students learn better: The impact of student initialization on knowledge distillation. CoRR, abs/1908.08962.
  50. 50.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008.
  51. 51.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019a. Superglue: A stickier benchmark for general-purpose language understanding systems. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 3261–3275.
  52. 52.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019b. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  53. 53.Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers), pages 1112–1122. Association for Computational Linguistics.
  54. 54.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of EMNLP 2020 - Demos, pages 38–45. Association for Computational Linguistics.
  55. 55.Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin. 2020. Deebert: Dynamic early exiting for accelerating BERT inference. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 2246–2251. Association for Computational Linguistics.
  56. 56.Ji Xin, Raphael Tang, Yaoliang Yu, and Jimmy Lin. 2021. Berxit: Early exiting for BERT with better fine-tuning and extension to regression. In Proceedings of EACL 2021, pages 91–104. Association for Computational Linguistics.
  57. 57.Canwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei, and Ming Zhou. 2020a. Bert-of-theseus: Compressing BERT by progressive module replacing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 7859–7869. Association for Computational Linguistics.
  58. 58.Liang Xu, Hai Hu, Xuanwei Zhang, Lu Li, Chenjie Cao, Yudong Li, Yechen Xu, Kai Sun, Dian Yu, Cong Yu, Yin Tian, Qianqian Dong, Weitang Liu, Bo Shi, Yiming Cui, Junyi Li, Jun Zeng, Rongzhao Wang, Weijian Xie, Yanting Li, Yina Patterson, Zuoyu Tian, Yiwen Zhang, He Zhou, Shaoweihua Liu, Zhe Zhao, Qipeng Zhao, Cong Yue, Xinrui Zhang, Zhengliang Yang, Kyle Richardson, and Zhenzhong Lan. 2020b. CLUE: A chinese language understanding evaluation benchmark. In Proceedings of the 28th International Conference on Computational Linguistics, COLING 2020, Barcelona, Spain (Online), December 8-13, 2020, pages 4762–4772. International Committee on Computational Linguistics.
  59. 59.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019. Xlnet: Generalized autoregressive pretraining for language understanding. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 5754–5764.
  60. 60.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, pages 2369–2380. Association for Computational Linguistics.
  61. 61.Wangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley, Ke Xu, and Furu Wei. 2020a. BERT loses patience: Fast and robust inference with early exit. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  62. 62.Wangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley, Ke Xu, and Furu Wei. 2020b. BERT loses patience: Fast and robust inference with early exit. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  63. 63.Wei Zhu. 2021. Leebert: Learned early exit for BERT with cross-level optimization. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 2968–2980. Association for Computational Linguistics.
  64. 64.Yukun Zhu, Ryan Kiros, Richard S. Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, December 7-13, 2015, pages 19–27. IEEE Computer Society.

Citation

MLA
Liu, X., et al. “Towards Efficient NLP: A Standard Evaluation and A Strong Baseline”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 3288–303, https://doi.org/10.18653/V1/2022.NAACL-MAIN.240.
APA
Liu, X., Sun, T., He, J., Wu, J., Wu, L., Zhang, X., Jiang, H., Cao, Z., Huang, X., & Qiu, X. (2022). Towards Efficient NLP: A Standard Evaluation and A Strong Baseline. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3288–3303. https://doi.org/10.18653/V1/2022.NAACL-MAIN.240
Chicago
Liu, X., T. Sun, J. He, et al. 2022. “Towards Efficient NLP: A Standard Evaluation and A Strong Baseline”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3288–3303. https://doi.org/10.18653/V1/2022.NAACL-MAIN.240.
Harvard
Liu, X. et al. (2022) “Towards Efficient NLP: A Standard Evaluation and A Strong Baseline”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 3288–3303. Available at: https://doi.org/10.18653/V1/2022.NAACL-MAIN.240.
Vancouver
1. Liu X, Sun T, He J, Wu J, Wu L, Zhang X, Jiang H, Cao Z, Huang X, Qiu X (2022) Towards Efficient NLP: A Standard Evaluation and A Strong Baseline. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 3288–3303

BibTeX

@inproceedings{Liu_2022, title={Towards Efficient NLP: A Standard Evaluation and A Strong Baseline}, url={http://dx.doi.org/10.18653/V1/2022.NAACL-MAIN.240}, DOI={10.18653/v1/2022.naacl-main.240}, booktitle={Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies}, publisher={Association for Computational Linguistics}, author={Liu, Xiangyang and Sun, Tianxiang and He, Junliang and Wu, Jiawen and Wu, Lingling and Zhang, Xinyu and Jiang, Hao and Cao, Zhao and Huang, Xuanjing and Qiu, Xipeng}, year={2022}, pages={3288–3303} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/