MPNet: Masked and Permuted Pre-training for Language Understanding

Kaitao SongXu TanTao QinJianfeng LuTie-Yan Liu

article2020NeurIPS1,888 citations

Introduces MPNet, a language model pre-training method that integrates masked and permuted language modeling to capture predicted token dependencies while eliminating position discrepancy, consistently outperforming BERT, XLNet, and RoBERTa on downstream language understanding benchmarks.

Listen

Modern language understanding relies heavily on self-supervised pre-training, where models learn broad language representations from massive text corpora before being fine-tuned on specific tasks. Existing leading frameworks face a fundamental design trade-off. Masked language modeling, used in models like BERT, allows the system to see full-sentence position information but fails to capture dependencies between multiple missing words because it predicts them independently. Conversely, permuted language modeling, used in models like XLNet, captures dependencies among predicted words but deprives the model of full-sentence positional context during pre-training, causing a structural mismatch when transitioning to downstream tasks. The article introduces MPNet (Masked and Permuted Network) to evaluate whether unifying these two approaches into a single training objective eliminates their individual weaknesses and achieves superior language understanding performance.

To demonstrate this, the authors established a unified view of masked and permuted modeling. MPNet incorporates permuted language modeling to learn dependencies among target tokens while adding auxiliary position compensation via special mask symbols, ensuring the model always observes the complete positional structure of the sentence. Credibility is established through extensive pre-training on over 160 gigabytes of standard text corpora over 500,000 steps using 32 high-end GPUs. The resulting model was evaluated across several standard language benchmarks, including the multi-task GLUE benchmark, question-answering datasets (SQuAD v1.1 and v2.0), reading comprehension (RACE), and sentiment classification (IMDB).

The findings show that MPNet consistently outperforms established baselines under identical standard base configurations. On the GLUE benchmark development set, MPNet achieved an average score of 87.9, surpassing BERT by 4.8 points, XLNet by 3.4 points, and RoBERTa by 1.5 points, while also beating ELECTRA by 0.7 points on the GLUE test set. In question answering on SQuAD v2.0, MPNet achieved an exact match and F1 score of 82.8 and 85.8, outperforming previous baselines by 2 to 9 points. On reading comprehension benchmarks, MPNet scored 76.1% accuracy, substantially exceeding comparable baselines. Ablation experiments confirmed that both core innovations—position compensation and permuted dependency modeling—were essential drivers of these accuracy gains.

These results demonstrate that resolving pre-training and fine-tuning discrepancies while preserving inter-token dependencies substantially improves model accuracy across varied language tasks. In terms of resource efficiency, MPNet achieved equal or better performance using fewer overall computational operations than previous approaches; for example, an intermediate 300,000-step MPNet model outperformed RoBERTa while requiring roughly 40% fewer computations. For engineering and product teams, this framework offers a more accurate, computationally competitive foundational architecture for text understanding without introducing deployment overhead, as auxiliary query streams are discarded during fine-tuning.

Organizations developing or deploying natural language processing systems should consider adopting MPNet as a superior baseline backbone over legacy BERT or XLNet architectures for core understanding tasks. As next steps, teams should validate MPNet on domain-specific datasets and test larger-scale variants. A key limitation noted in the article is the substantial computational footprint and time required to pre-train large base models from scratch—requiring over a month of training on multi-GPU clusters. However, given the rigorous ablation tests and consistent gains across diverse standardized benchmarks, confidence in the reported performance advantages remains high.

Cover for MPNet: Masked and Permuted Pre-training for Language Understanding

Abstract

BERT adopts masked language modeling (MLM) for pre-training and is one of the most successful pre-training models. Since BERT neglects dependency among predicted tokens, XLNet introduces permuted language modeling (PLM) for pre-training to address this problem. However, XLNet does not leverage the full position information of a sentence and thus suffers from position discrepancy between pre-training and fine-tuning. In this paper, we propose MPNet, a novel pre-training method that inherits the advantages of BERT and XLNet and avoids their limitations. MPNet leverages the dependency among predicted tokens through permuted language modeling (vs. MLM in BERT), and takes auxiliary position information as input to make the model see a full sentence and thus reducing the position discrepancy (vs. PLM in XLNet). We pre-train MPNet on a large-scale dataset (over 160GB text corpora) and fine-tune on a variety of down-streaming tasks (GLUE, SQuAD, etc). Experimental results show that MPNet outperforms MLM and PLM by a large margin, and achieves better results on these tasks compared with previous state-of-the-art pre-trained methods (e.g., BERT, XLNet, RoBERTa) under the same model setting. The code and the pre-trained models are available at: this https URL.

Table of Contents

  • 1 Introduction
  • 2 MPNet
  • 2.1 Background
  • 2.2 A Unified View of MLM and PLM
  • 2.3 Our Proposed Method
  • 2.4 Discussions
  • 3 Experiments and Results
  • 3.1 Experimental Setup
  • 3.2 Results on GLUE Benchmark
  • 3.3 Results on Question Answering (SQuAD)
  • 3.4 Results on RACE
  • 3.5 Results on IMDB
  • 3.6 Ablation Study
  • 4 Conclusion
  • References
  • A Pre-training Hyper-parameters
  • B Fine-tuning Hyper-parameters
  • C MPNet on Large Setting
  • D Effect of MNLI initialization
  • E Training Efficiency
  • F More Ablation Studies
  • G Training Speedup

Knowls

  1. Knowl 1 — Masked and Permuted Language Modeling (MPNet) Training Objective

    model/method

    Masked and Permuted Pre-training (MPNet) unifies masked language modeling (MLM) and permuted language modeling (PLM) into a single framework. Given a token sequence x=(x1,x2,…,xn)x = (x_1, x_2, \dots, x_n) and a random permutation z∈Znz \in \mathcal{Z}_n of index positions (1,2,…,n)(1, 2, \dots, n), the sequence is partitioned into a non-predicted segment xz≤cx_{z_{\le c}} of length cc (typically c=0.85nc = 0.85n) and a predicted segment xz>cx_{z_{>c}} of length n−cn - c. Mask tokens Mz>c=([M],[M],…,[M])M_{z_{>c}} = ([\text{M}], [\text{M}], \dots, [\text{M}]) are placed at positions z>cz_{>c} alongside the non-predicted tokens.

    The training objective maximizes the log-likelihood of predicted tokens conditioned on both preceding tokens in the permutation order and the position-compensating mask tokens:

    LMPNet(θ)=Ez∈Zn∑t=c+1nlog⁡P(xzt∣xz<t,Mz>c;θ)\mathcal{L}_{\text{MPNet}}(\theta) = \mathbb{E}_{z \in \mathcal{Z}_n} \sum_{t=c+1}^n \log P\left(x_{z_t} \mid x_{z_{<t}}, M_{z_{>c}}; \theta\right)

    where:

    • nn is the sequence length in tokens;
    • Zn\mathcal{Z}_n is the set of all n!n! possible permutations of (1,2,…,n)(1, 2, \dots, n);
    • ztz_t denotes the index at step tt of permutation zz;
    • xz<t=(xz1,…,xzt−1)x_{z_{<t}} = (x_{z_1}, \dots, x_{z_{t-1}}) represents all tokens preceding position ztz_t in the permutation;
    • Mz>cM_{z_{>c}} contains the mask tokens [M][\text{M}] associated with the target positions z>cz_{>c};
    • θ\theta denotes the model parameters.

    This objective enables autoregressive modeling of dependencies among predicted target tokens (addressing MLM's conditional independence assumption) while feeding complete positional information for the entire sequence (addressing PLM's full-sentence position discrepancy).

  2. Knowl 2 — Two-Stream Self-Attention with Position Compensation

    model/method

    MPNet adopts a modified two-stream self-attention mechanism with position compensation to predict target tokens autoregressively while maintaining full-sentence positional awareness. For an input sequence x=(x1,…,xn)x = (x_1, \dots, x_n) permuted according to z=(z1,…,zn)z = (z_1, \dots, z_n) with non-predicted length cc, the input consists of:

    • Non-predicted tokens and masks: (xz≤c,Mz>c)=(xz1,…,xzc,[M],…,[M])(x_{z_{\le c}}, M_{z_{>c}}) = (x_{z_1}, \dots, x_{z_c}, [\text{M}], \dots, [\text{M}]) with position sequence (pz1,…,pzc,pzc+1,…,pzn)(p_{z_1}, \dots, p_{z_c}, p_{z_{c+1}}, \dots, p_{z_n});
    • Predicted tokens: xz>c=(xzc+1,…,xzn)x_{z_{>c}} = (x_{z_{c+1}}, \dots, x_{z_n}) with position sequence (pzc+1,…,pzn)(p_{z_{c+1}}, \dots, p_{z_n}).

    The two streams operate as follows:

    1. Content Stream: Computes contextual representations. The non-predicted segment (xz≤c,Mz>c)(x_{z_{\le c}}, M_{z_{>c}}) uses bidirectional self-attention. For predicted tokens, the content representation at step tt (t>ct > c) attends to all non-predicted tokens, preceding predicted tokens xz<tx_{z_{<t}}, and remaining position-compensating mask tokens Mz≥tM_{z_{\ge t}}.
    2. Query Stream: Generates target token predictions without seeing the target token itself. When predicting token xztx_{z_t}, the query stream takes mask token [M][\text{M}] and position pztp_{z_t} as query input, attending to the content stream representations of xz<tx_{z_{<t}} as well as the position-compensating mask tokens Mz≥t=([M],…,[M])M_{z_{\ge t}} = ([\text{M}], \dots, [\text{M}]) with their positions pz≥tp_{z_{\ge t}}.

    By including Mz≥tM_{z_{\ge t}} and pz≥tp_{z_{\ge t}}, each prediction step attends to exactly nn tokens/positions, ensuring the model retains full structural context of the sequence length and layout during pre-training, mirroring downstream task inputs.

  3. Knowl 3 — Attention Matrix Partitioning for Content Stream Efficiency

    model/method

    To reduce pre-training computational overhead in MPNet, the content stream self-attention matrix over the combined non-predicted segment (xz≤c,Mz>c)(x_{z_{\le c}}, M_{z_{>c}}) of length nn and predicted segment xz>cx_{z_{>c}} of length n−cn-c is partitioned into four sub-matrices:

    • Block A: Queries from the non-predicted segment, Keys/Values from the non-predicted segment (bidirectional attention of size n×nn \times n);
    • Block B: Queries from the non-predicted segment, Keys/Values from the predicted segment (size n×(n−c)n \times (n-c));
    • Block C: Queries from the predicted segment, Keys/Values from the non-predicted segment (size (n−c)×n(n-c) \times n);
    • Block D: Queries from the predicted segment, Keys/Values from the predicted segment (masked autoregressive attention of size (n−c)×(n−c)(n-c) \times (n-c)).

    Because the representations of the non-predicted segment do not require information from the autoregressively predicted targets, Block B is set to un-attended (zero attention weights) and excluded from computation. Computing only Blocks A, C, and D saves approximately 10% of total pre-training computation.

  4. Knowl 4 — Input Information Conditioning Capacity across Pre-training Objectives

    model/method

    Assuming a standard 15% masking/prediction ratio (c=0.85nc = 0.85n), the proportion of token content and positional information available to the model when predicting target tokens varies fundamentally across pre-training objectives:

    • Masked Language Modeling (BERT): Predicts masked tokens conditionally independent of one another given unmasked tokens and [M][\text{M}] tokens. It conditions on 85%85\% of token content and 100%100\% of position embeddings.
    • Permuted Language Modeling (XLNet): Autoregressively factorizes target tokens over permutations without full-sentence position compensation. The ii-th target token sees i−1i-1 preceding targets, yielding on average (k−1)/2≈7.5%(k-1)/2 \approx 7.5\% of target tokens. It conditions on 85%+7.5%=92.5%85\% + 7.5\% = 92.5\% of token content and 92.5%92.5\% of position embeddings, suffering from input position discrepancy relative to fine-tuning.
    • Masked and Permuted Language Modeling (MPNet): Leverages autoregressive factorization over the permutation while restoring full-sentence layout via position-compensating mask tokens. It conditions on 92.5%92.5\% of token content and 100%100\% of position embeddings, maximizing contextual and positional awareness.
  5. Knowl 5 — MPNet Architecture and Pre-training Setup

    experimental setup

    MPNet is evaluated primarily under the BERTBASE\text{BERT}_{\text{BASE}} architecture:

    • Model Dimensions: 12 Transformer layers, hidden size d=768d = 768, intermediate feed-forward filter size 3072, 12 attention heads, and 110M total parameters.
    • Pre-training Corpus: 160GB total text data matching RoBERTa, comprising Wikipedia, BooksCorpus, OpenWebText, CC-News, and Stories. Tokenization uses a 30K byte-pair encoding (BPE) subword vocabulary.
    • Hyperparameters: Adam optimizer with β1=0.9,β2=0.98,ϵ=1×10−6\beta_1 = 0.9, \beta_2 = 0.98, \epsilon = 1\times 10^{-6}, weight decay 0.01, peak learning rate 6×10−46 \times 10^{-4}, warmup ratio 0.06 with linear decay. Sequence length is capped at 512 tokens with a batch size of 8,192 sequences.
    • Training Duration & Compute: Trained for 500K steps using FP16 mixed precision on 32 NVIDIA Tesla V100 32GB GPUs (total duration: 35 days).
    • Additional Enhancements: Employs whole word masking, a layer-shared relative position embedding across attention heads, and BERT's 8:1:1 token replacement strategy (80% [M][\text{M}], 10% random token, 10% original token) on the 15% predicted tokens.
    • Fine-tuning: The query stream is removed; only the content stream hidden states are used for downstream prediction tasks.
  6. Knowl 6 — GLUE Benchmark Evaluation Results for MPNet Base

    data/table

    Under the BERTBASE\text{BERT}_{\text{BASE}} setting without data augmentation, MPNet achieves superior performance across GLUE benchmark tasks compared to BERT, XLNet, RoBERTa, and ELECTRA:

    Model MNLI QNLI QQP RTE SST-2 MRPC CoLA STS-B Avg
    Single model on dev set
    BERT 84.5 91.7 91.3 68.6 93.2 87.3 58.9 89.5 83.1
    XLNet 86.8 91.7 91.4 74.0 94.7 88.2 60.2 89.5 84.5
    RoBERTa 87.6 92.8 91.9 78.7 94.8 90.2 63.6 91.2 86.4
    MPNet 88.5 93.3 91.9 85.8 95.5 91.8 65.0 91.1 87.9
    Single model on test set
    BERT 84.6 90.5 89.2 66.4 93.5 84.8 52.1 87.1 79.9
    ELECTRA 88.5 93.1 89.5 75.2 96.0 88.1 64.6 91.0 85.8
    MPNet 88.5 93.1 89.9 81.0 96.0 89.1 64.0 90.7 86.5

    Metrics: STS-B is evaluated using Pearson correlation, CoLA uses Matthew's correlation, and other tasks report accuracy. On the development set, MPNet outperforms BERT, XLNet, and RoBERTa by 4.8, 3.4, and 1.5 average points, respectively. On the test set, MPNet exceeds ELECTRA by 0.7 points on average, with substantial gains on inference benchmarks such as RTE (81.0 vs. 75.2).

  7. Knowl 7 — Performance on SQuAD, RACE, and IMDB Benchmarks

    empirical result

    Under the BERTBASE\text{BERT}_{\text{BASE}} setting, MPNet demonstrates strong transfer capability on question answering and comprehension benchmarks:

    1. Stanford Question Answering Dataset (SQuAD):

      • SQuAD v1.1 Dev: MPNet achieves 86.9 Exact Match (EM) / 92.7 F1, outperforming BERT (80.8 / 88.5), XLNet (81.3 / -), and RoBERTa (84.6 / 91.5).
      • SQuAD v2.0 Dev: MPNet achieves 82.7 EM / 85.7 F1, outperforming BERT (73.7 / 76.3), XLNet (80.2 / -), and RoBERTa (80.5 / 83.7).
      • SQuAD v2.0 Test: MPNet obtains 82.8 EM / 85.8 F1 compared to BERT's 73.1 / 76.2.
    2. ReAding Comprehension from Examinations (RACE):

      • Pre-trained on 16GB data (Wikipedia + BooksCorpus), MPNet achieves an overall accuracy of 70.4% (Middle school: 76.8%, High school: 67.7%), outperforming BERT (65.0%) and XLNet (66.8%).
      • Pre-trained on full 160GB corpora, MPNet achieves 76.1% overall accuracy (Middle: 79.7%, High: 74.5%).
    3. IMDB Sentiment Classification:

      • On 16GB pre-training data, MPNet achieves an error rate of 4.8% (versus BERT's 5.4% and XLNet's 4.9%). Full 160GB pre-training lowers the error rate to 4.4%.
  8. Knowl 8 — Ablation Analysis of Position Compensation, Permutation, and Architecture Enhancements

    empirical result

    Ablation experiments conducted on BERTBASE\text{BERT}_{\text{BASE}} models pre-trained on 16GB data (Wikipedia and BooksCorpus) for 1M steps (batch size 256, sequence length 512) demonstrate the specific contributions of MPNet's components:

    Model Setting SQuAD v1.1 SQuAD v2.0 MNLI SST-2
    MPNet 85.0 / 91.4 80.5 / 83.3 86.2 94.0
    −- position compensation (= PLM) 83.0 / 89.9 78.5 / 81.0 85.6 93.4
    −- permutation (= MLM + output dep.) 84.1 / 90.6 79.2 / 81.8 85.7 93.5
    −- permutation output dep. (= MLM) 82.0 / 89.5 76.8 / 79.8 85.6 93.3
    −- whole word mask 84.0 / 90.5 79.9 / 82.5 85.6 93.8
    −- relative positional embedding 84.0 / 90.3 79.5 / 82.2 85.3 93.6

    Key takeaways:

    • Removing position compensation degrades MPNet to standard PLM, causing a 0.6 to 2.3 point drop across tasks, proving the value of eliminating pretrain-finetune positional discrepancy.
    • Retaining output dependency without sequence permutation reduces performance by 0.5 to 1.7 points, confirming the benefit of permuted factorization.
    • Removing both permutation and output dependency degenerates the model to standard MLM, resulting in a 0.5 to 3.7 point drop.
    • Both whole word masking and relative positional embeddings provide consistent gains (e.g., +1.0 EM and +0.9/+1.1 F1 on SQuAD v1.1).
  9. Knowl 9 — Training Computation and Efficiency Comparison

    empirical result

    Under the BERTBASE\text{BERT}_{\text{BASE}} setting on 160GB text corpora, MPNet achieves superior performance with equal or fewer training floating-point operations (FLOPs) compared to RoBERTa and XLNet:

    • BERT (16GB data): 6.4×10196.4 \times 10^{19} FLOPs (0.06×0.06\times), GLUE dev average: 83.1.
    • XLNet (160GB data): 1.3×10211.3 \times 10^{21} FLOPs (1.10×1.10\times), GLUE dev average: 84.5.
    • RoBERTa (160GB data): 1.1×10211.1 \times 10^{21} FLOPs (0.92×0.92\times), GLUE dev average: 86.4.
    • MPNet-300K (300K training steps): 7.1×10207.1 \times 10^{20} FLOPs (0.60×0.60\times), GLUE dev average: 87.7.
    • MPNet-500K (500K training steps): 1.2×10211.2 \times 10^{21} FLOPs (1.00×1.00\times), GLUE dev average: 87.9.

    At 300K steps (0.60×0.60\times training FLOPs), MPNet already outperforms RoBERTa by 1.3 points on the GLUE benchmark average while requiring roughly 35%35\% fewer FLOPs.

Coverage note — None was omitted; all key contributions including mathematical formulation, architecture modifications (position compensation, matrix partitioning), experimental configurations, benchmark evaluations, training efficiency, and ablation studies are covered.

References

  1. 1.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training. URL https://s3-us-west-2. amazonaws. com/openai-assets/researchcovers/languageunsupervised/language understanding paper. pdf, 2018.
  2. 2.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL, pages 4171–4186, 2019.
  3. 3.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. OpenAI Blog, 1(8), 2019.
  4. 4.Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. MASS: Masked sequence to sequence pre-training for language generation. In ICML, pages 5926–5936, 2019.
  5. 5.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. Xlnet: Generalized autoregressive pretraining for language understanding. In NIPS, pages 5754–5764, 2019.
  6. 6.Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. Unified language model pre-training for natural language understanding and generation. In NIPS, pages 13042–13054, 2019.
  7. 7.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  8. 8.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683, 2019.
  9. 9.Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy. Spanbert: Improving pre-training by representing and predicting spans. arXiv preprint arXiv:1907.10529, 2019.
  10. 10.Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. Electra: Pre-training text encoders as discriminators rather than generators. arXiv preprint arXiv:2003.10555, 2020.
  11. 11.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NIPS, pages 5998–6008, 2017.
  12. 12.Yiming Cui, Wanxiang Che, Ting Liu, Bing Qin, Ziqing Yang, Shijin Wang, and Guoping Hu. Pre-training with whole word masking for chinese BERT. CoRR, abs/1906.08101, 2019.
  13. 13.Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. Self-attention with relative position representations. In NAACL, pages 464–468, June 2018.
  14. 14.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv e-prints, 2019.
  15. 15.Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In ICCV, pages 19–27, December 2015.
  16. 16.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. 2019.
  17. 17.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692, 2019.
  18. 18.Trieu H. Trinh and Quoc V. Le. A simple method for commonsense reasoning. CoRR, abs/1806.02847, 2018.
  19. 19.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization, 2014.
  20. 20.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In ICLR, 2019.
  21. 21.Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. Neural network acceptability judgments. CoRR, abs/1805.12471, 2018.
  22. 22.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. Recursive deep models for semantic compositionality over a sentiment treebank. In EMNLP, pages 1631–1642, October 2013.
  23. 23.William B. Dolan and Chris Brockett. Automatically constructing a corpus of sentential paraphrases. In Proceedings of the Third International Workshop on Paraphrasing (IWP2005), 2005.
  24. 24.Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation. In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), pages 1–14, August 2017.
  25. 25.Adina Williams, Nikita Nangia, and Samuel Bowman. A broad-coverage challenge corpus for sentence understanding through inference. In NAACL, pages 1112–1122, June 2018.
  26. 26.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. SQuAD: 100,000+ questions for machine comprehension of text. In EMNLP, pages 2383–2392, 2016.
  27. 27.Ido Dagan, Oren Glickman, and Bernardo Magnini. The pascal recognising textual entailment challenge. In Machine Learning Challenges. Evaluating Predictive Uncertainty, Visual Object Classification, and Recognising Tectual Entailment, pages 177–190, 2006.
  28. 28.Hector J. Levesque, Ernest Davis, and Leora Morgenstern. The Winograd Schema Challenge. In Proceedings of the Thirteenth International Conference on Principles of Knowledge Representation and Reasoning, pages 552–561. Rome, Italy, 2012.
  29. 29.Pranav Rajpurkar, Robin Jia, and Percy Liang. Know what you don’t know: Unanswerable questions for SQuAD. In ACL, pages 784–789, 2018.
  30. 30.Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. RACE: Large-scale ReAding comprehension dataset from examinations. In EMNLP, pages 785–794, 2017.
  31. 31.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 142–150, June 2011.
  32. 32.Chi Sun, Xipeng Qiu, Yige Xu, and Xuanjing Huang. How to fine-tune BERT for text classification? CoRR, abs/1905.05583, 2019.

Citation

MLA
Song, K., et al. “MPNet: Masked and Permuted Pre-training for Language Understanding”. Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 16857–67, https://proceedings.neurips.cc/paper_files/paper/2020/file/c3a690be93aa602ee2dc0ccab5b7b67e-Paper.pdf.
APA
Song, K., Tan, X., Qin, T., Lu, J., & Liu, T.-Y. (2020). MPNet: Masked and Permuted Pre-training for Language Understanding. Advances in Neural Information Processing Systems, 33, 16857–16867. https://proceedings.neurips.cc/paper_files/paper/2020/file/c3a690be93aa602ee2dc0ccab5b7b67e-Paper.pdf
Chicago
Song, K., X. Tan, T. Qin, J. Lu, and T.-Y. Liu. 2020. “MPNet: Masked and Permuted Pre-training for Language Understanding”. Advances in Neural Information Processing Systems 33: 16857–67. https://proceedings.neurips.cc/paper_files/paper/2020/file/c3a690be93aa602ee2dc0ccab5b7b67e-Paper.pdf.
Harvard
Song, K. et al. (2020) “MPNet: Masked and Permuted Pre-training for Language Understanding”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 16857–16867. Available at: https://proceedings.neurips.cc/paper_files/paper/2020/file/c3a690be93aa602ee2dc0ccab5b7b67e-Paper.pdf.
Vancouver
1. Song K, Tan X, Qin T, Lu J, Liu T-Y (2020) MPNet: Masked and Permuted Pre-training for Language Understanding. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 16857–16867

BibTeX

@inproceedings{song2020mpnet,
  title = {MPNet: Masked and Permuted Pre-training for Language Understanding},
  author = {Song, Kaitao and Tan, Xu and Qin, Tao and Lu, Jianfeng and Liu, Tie-Yan},
  year = {2020},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {33},
  pages = {16857-16867},
  url = {https://proceedings.neurips.cc/paper_files/paper/2020/file/c3a690be93aa602ee2dc0ccab5b7b67e-Paper.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Published with permission