XLM-E: Cross-lingual Language Model Pre-training via ELECTRA

Zewen ChiShaohan HuangLi DongShuming MaBo ZhengSaksham SinghalPayal BajajXia SongXian-Ling MaoHeyan Huang

article2022ACL142 citations

Proposes a compute-efficient cross-lingual pre-training framework that adapts ELECTRA-style discriminative tasks to multilingual and parallel corpora, achieving superior cross-lingual transferability over strong baselines with only a fraction of the computation cost.

Listen

Building artificial intelligence models that understand multiple languages typically requires immense computational power and financial cost. Conventional approaches rely on generative masked language modeling, which trains systems by predicting hidden words in text sentences. Because this training process evaluates only a small fraction of tokens per pass, training state-of-the-art multilingual models demands weeks or months of expensive compute infrastructure. As global organizations seek to deploy scalable natural language applications across diverse languages, reducing this computational and energetic footprint is a pressing priority.

The article introduces and evaluates XLM-E, a cross-lingual language model trained with discriminative tasks rather than traditional masked language generation. The objective is to demonstrate that detecting whether words in multilingual and translated sentences have been replaced by plausible alternatives can significantly lower computational resource requirements while achieving superior cross-lingual transfer performance.

To accomplish this, the authors implemented two core training tasks: multilingual replaced token detection, which operates on single-language texts across 100 languages, and translation replaced token detection, which processes parallel translated pairs across 100 languages. In addition, the architecture integrates a gated relative position bias to adaptively handle varying word order and distance patterns across languages. The model was trained using large-scale CommonCrawl and parallel web datasets across 125,000 steps on 64 graphics processing units over 1.7 days, and evaluated across seven distinct multilingual benchmarks including question answering, classification, and structured prediction tasks.

The analysis reveals that XLM-E delivers a drastic reduction in computational expenditure, achieving up to a 130-fold speedup and utilizing approximately 1% of the total floating-point operations required by standard baseline models such as XLM-R. Despite this fraction of compute, XLM-E achieved an average score of 69.3 across the multilingual benchmark suite, outperforming existing baselines. Furthermore, the model exhibited stronger cross-lingual alignment at deeper layers and scaled effectively: when enlarged to 2.2 billion parameters, it outperformed competing architectures that contained up to 3.7 billion parameters on reading comprehension and inference tasks.

These findings indicate that discriminative pre-training provides a far more compute-efficient and sample-efficient foundation for multilingual AI systems. Organizations can achieve state-of-the-art cross-lingual transferability at a fraction of standard cloud compute costs and training timelines, substantially mitigating operational expense and environmental impact. The ability of the model to align representations across languages without massive parallel data processing challenges the prevailing assumption that large-scale multilingual mastery necessitates prohibitive compute budgets.

Decision-makers should consider adopting discriminative pre-training frameworks like XLM-E when building or fine-tuning multilingual language pipelines, particularly for enterprise translation, multilingual classification, and cross-lingual question answering. Before enterprise-wide rollout, teams should conduct internal pilots to evaluate performance on specialized proprietary vocabularies and explore scaling beyond base configurations. However, leaders should note that the evaluation was confined to standard open benchmarks and high-to-medium resource language pairs; validation on highly specialized domain data and extremely low-resource languages remains recommended to ensure robust operational deployment.

arXiv: 2106.16138
Cover for XLM-E: Cross-lingual Language Model Pre-training via ELECTRA

Abstract

In this paper, we introduce ELECTRA-style tasks (Clark et al., 2020b) to cross-lingual language model pre-training. Specifically, we present two pre-training tasks, namely multilingual replaced token detection, and translation replaced token detection. Besides, we pretrain the model, named as XLM-E, on both multilingual and parallel corpora. Our model outperforms the baseline models on various cross-lingual understanding tasks with much less computation cost. Moreover, analysis shows that XLM-E tends to obtain better cross-lingual transferability.

Table of Contents

  • 1 Introduction
  • 2 Background: ELECTRA
  • 3 Methods
  • 3.1 Pre-training Tasks
  • 3.2 Pre-training XLM-E
  • 3.3 Gated Relative Position Bias
  • 4 Experiments
  • 4.1 Setup
  • 4.2 Cross-lingual Understanding
  • 4.3 Ablation Studies
  • 4.4 Scaling-up Results
  • 4.5 Training Efficiency
  • 4.6 Cross-lingual Alignment
  • 4.7 Universal Layer Across Languages
  • 4.8 Cross-lingual Transfer Gap
  • 5 Related Work
  • 6 Conclusion
  • 7 Ethical Considerations
  • Acknowledgements
  • References
  • Appendix
  • A Model Hyperparameters
  • B Hyperparameters for Pre-Training
  • C Hyperparameters for Fine-Tuning

Knowls

  1. Knowl 1 — XLM-E joint generator–discriminator objective

    model/method

    XLM-E is a cross-lingual Transformer encoder pretrained in the ELECTRA style with two jointly trained components: a generator with parameters θG\theta_G and a discriminator with parameters θD\theta_D. The generator is trained with masked language modeling on multilingual sentences and translation language modeling on parallel sentence pairs; the discriminator is trained with multilingual replaced token detection and translation replaced token detection. For a multilingual corpus XX and a parallel corpus PP, the joint objective is

    L=LMLM(x;θG)+LTLM(e,f;θG)+λMRTDLMRTD(x;θD)+λTRTDLTRTD(e,f;θD),\mathcal{L}=\mathcal{L}_{\mathrm{MLM}}(x;\theta_G)+\mathcal{L}_{\mathrm{TLM}}(e,f;\theta_G)+\lambda_{\mathrm{MRTD}}\mathcal{L}_{\mathrm{MRTD}}(x;\theta_D)+\lambda_{\mathrm{TRTD}}\mathcal{L}_{\mathrm{TRTD}}(e,f;\theta_D),

    where x∈Xx\in X is a multilingual sentence, (e,f)∈P(e,f)\in P is a translation pair, and the four loss terms are defined by the corresponding pre-training tasks. In the reported training, both discriminator-loss weights are set to 5050. Only the discriminator is used as the downstream language encoder after pre-training.

  2. Knowl 2 — Multilingual replaced token detection

    model/method

    Multilingual replaced token detection (MRTD) applies ELECTRA’s replaced-token objective to sentences from multiple languages. Let x=(x1,…,xn)x=(x_1,\ldots,x_n) be a sentence of nn tokens and let M⊆{1,…,n}M\subseteq\{1,\ldots,n\} be a uniformly sampled set of masked positions. The generator receives xmaskedx^{\mathrm{masked}}, obtained by replacing tokens at positions in MM with a mask symbol, and produces a distribution over replacement tokens. The corrupted sentence xcorruptx^{\mathrm{corrupt}} is formed by

    xicorrupt∼pG(⋅∣xmasked)(i∈M),xicorrupt=xi(i∉M).x_i^{\mathrm{corrupt}}\sim p_G(\cdot\mid x^{\mathrm{masked}})\quad(i\in M),\qquad x_i^{\mathrm{corrupt}}=x_i\quad(i\notin M).

    The discriminator predicts for every position whether the corrupted token is original or generator-sampled. If zi∈{0,1}z_i\in\{0,1\} is this binary label, the discriminator loss is

    LMRTD(x;θD)=−∑i=1nlog⁡pD(zi∣xcorrupt).\mathcal{L}_{\mathrm{MRTD}}(x;\theta_D)=-\sum_{i=1}^{n}\log p_D(z_i\mid x^{\mathrm{corrupt}}).

    The generator and discriminator are shared across languages, as is the vocabulary. The authors report that span masking reduced generator prediction accuracy and consequently harmed pre-training relative to uniform masking.

  3. Knowl 3 — Translation replaced token detection

    model/method

    Translation replaced token detection (TRTD) uses parallel sentences as the corrupted context. Let e=(e1,…,ene)e=(e_1,\ldots,e_{n_e}) and f=(f1,…,fnf)f=(f_1,\ldots,f_{n_f}) be translations, and let [e;f][e;f] denote their concatenation. Independent masked-position sets MeM_e and MfM_f are selected in the two languages. The generator predicts masked tokens in both languages using the concatenated pair, with loss

    LGTRTD(e,f;θG)=−∑i∈Melog⁡pG(ei∣[e;f]masked)−∑j∈Mflog⁡pG(fj∣[e;f]masked).\mathcal{L}_{G}^{\mathrm{TRTD}}(e,f;\theta_G)= -\sum_{i\in M_e}\log p_G(e_i\mid [e;f]^{\mathrm{masked}}) -\sum_{j\in M_f}\log p_G(f_j\mid [e;f]^{\mathrm{masked}}).

    Each masked token is then replaced by a token sampled from the generator, while unmasked tokens remain unchanged, producing ecorrupte^{\mathrm{corrupt}} and fcorruptf^{\mathrm{corrupt}}. If rk∈{0,1}r_k\in\{0,1\} labels whether the kkth token in the concatenated corrupted pair is original or replaced, the discriminator loss is

    LDTRTD(e,f;θD)=−∑k=1ne+nflog⁡pD(rk∣[ecorrupt;fcorrupt]).\mathcal{L}_{D}^{\mathrm{TRTD}}(e,f;\theta_D)= -\sum_{k=1}^{n_e+n_f}\log p_D(r_k\mid [e^{\mathrm{corrupt}};f^{\mathrm{corrupt}}]).

    TRTD therefore retains the translation-pair context of translation language modeling while training the main encoder to classify every token as real or replaced rather than to reconstruct only masked tokens.

  4. Knowl 4 — Gated relative position bias

    model/method

    XLM-E modifies Transformer self-attention with a content-dependent gated relative position bias. For token hidden states hi∈Rdhh_i\in\mathbb{R}^{d_h}, query, key, and value vectors are

    qi=hiWQ,ki=hiWK,vi=hiWV,q_i=h_iW^Q,\qquad k_i=h_iW^K,\qquad v_i=h_iW^V,

    where WQ,WK,WV∈Rdh×dkW^Q,W^K,W^V\in\mathbb{R}^{d_h\times d_k} and dkd_k is the attention-key dimension. The attention weights and output are

    aij=exp⁡(qi⋅kj/dk+ri−j)∑ℓexp⁡(qi⋅kℓ/dk+ri−ℓ),h~i=∑jaijvj,a_{ij}=\frac{\exp\left(q_i\cdot k_j/\sqrt{d_k}+r_{i-j}\right)}{\sum_{\ell}\exp\left(q_i\cdot k_\ell/\sqrt{d_k}+r_{i-\ell}\right)},\qquad \widetilde h_i=\sum_j a_{ij}v_j,

    where ri−jr_{i-j} is the relative-position bias for the distance between positions ii and jj. Let di−jd_{i-j} be a learnable scalar relative-position bias, uu and vv be learnable vectors in Rdk\mathbb{R}^{d_k}, ww be a learnable scalar, and σ\sigma be the sigmoid function. XLM-E computes

    giupdate=σ(qi⋅u),gireset=σ(qi⋅v),g_i^{\mathrm{update}}=\sigma(q_i\cdot u),\qquad g_i^{\mathrm{reset}}=\sigma(q_i\cdot v), r~i−j=w giresetdi−j,ri−j=di−j+giupdatedi−j+(1−giupdate)r~i−j.\widetilde r_{i-j}=w\,g_i^{\mathrm{reset}}d_{i-j},\qquad r_{i-j}=d_{i-j}+g_i^{\mathrm{update}}d_{i-j}+(1-g_i^{\mathrm{update}})\widetilde r_{i-j}.

    Unlike a fixed relative-position bias, this mechanism conditions positional information on the token content, allowing the same relative distance to receive different treatment in different linguistic contexts.

  5. Knowl 5 — Pre-training data, architecture, and optimization

    experimental setup

    XLM-E is pretrained on CC-100 multilingual text covering 100 languages and on parallel data from MultiUN, IIT Bombay, OPUS, WikiMatrix, and CCAligned. If language jj has mjm_j examples among NN languages, an example is sampled from that language with probability

    pj=mjα∑k=1Nmkα,α=0.7,p_j=\frac{m_j^{\alpha}}{\sum_{k=1}^{N}m_k^{\alpha}},\qquad \alpha=0.7,

    which increases sampling probability for lower-resource languages relative to proportional sampling. The Base discriminator is a 12-layer Transformer with hidden size 768768, feed-forward size 3,0723{,}072, and 12 attention heads; its generator is a 4-layer Transformer with the same hidden size, feed-forward size, and number of heads. The Base model uses a shared 250K SentencePiece subword vocabulary.

    The generator and discriminator are initialized from scratch and trained jointly for 125K steps with Adam using learning rate 5×10−45\times10^{-4}, ϵ=10−6\epsilon=10^{-6}, β=(0.9,0.98)\beta=(0.9,0.98), a linear schedule with 10,000 warm-up steps, gradient clipping at 2.02.0, and weight decay 0.010.01. Each pre-training task uses a dynamic batch of approximately 1M tokens; MRTD batches contain 2,048 sequences of length 512, while TRTD sequences use the lengths of the original translation pairs. The reported Base pre-training takes approximately 1.7 days on 64 Nvidia A100 GPUs.

  6. Knowl 6 — Cross-lingual understanding performance

    data/table

    In zero-shot cross-lingual transfer, each model is fine-tuned only on English training data and evaluated across the target languages of the seven-task XTREME benchmark. Scores are averaged over target languages; XLM-E and XLM-R results are averaged over five random seeds. Higher values are better, with F1 for POS and NER, F1/EM for question answering, and accuracy for XNLI and PAWS-X. XLM-E reaches an average score of 69.369.3, compared with 68.968.9 for XLM-ALIGN and 66.466.4 for XLM-R; the multilingual-only XLM-E variant without TRTD reaches 67.667.6, which is 1.21.2 points above XLM-R.

    Model POS F1 NER F1 XQuAD F1/EM MLQA F1/EM TyDiQA F1/EM XNLI Acc. PAWS-X Acc. Avg.
    MBERT 70.3 62.2 64.5/49.4 61.4/44.2 59.7/43.9 65.4 81.9 63.1
    MT5 – 55.7 67.0/49.0 64.6/45.0 57.2/41.2 75.4 86.4 –
    XLM-R 75.6 61.8 71.9/56.4 65.1/47.2 55.4/38.3 75.0 84.9 66.4
    XLM-E (without TRTD) 74.2 62.7 74.3/58.2 67.8/49.7 57.8/40.6 75.1 87.1 67.6
    XLM 70.1 61.2 59.8/44.3 48.5/32.6 43.6/29.1 69.1 80.9 58.6
    INFOXLM – – –/– 68.1/49.6 –/– 76.5 – –
    XLM-ALIGN 76.0 63.7 74.7/59.0 68.1/49.8 62.1/44.8 76.2 86.8 68.9
    XLM-E 75.6 63.5 76.2/60.2 68.3/49.8 62.4/45.7 76.6 88.3 69.3
  7. Knowl 7 — Ablation evidence for TRTD and gated attention

    empirical result

    Under matched pre-training settings, removing TRTD from XLM-E reduces XNLI accuracy from 76.676.6 to 75.175.1 and reduces MLQA from 68.3/49.868.3/49.8 F1/EM to 67.8/49.767.8/49.7. Removing the gated relative position bias as well changes the scores to 75.275.2 on XNLI and 67.4/49.267.4/49.2 on MLQA. A reimplementation of XLM with TLM obtains 73.473.4 on XNLI and 66.2/47.866.2/47.8 on MLQA, while removing TLM reduces it to 70.670.6 and 64.0/46.064.0/46.0, respectively.

    Model XNLI accuracy MLQA F1/EM
    XLM (reimplementation) 73.4 66.2/47.8
    XLM (without TLM) 70.6 64.0/46.0
    XLM-E 76.6 68.3/49.8
    XLM-E (without TRTD) 75.1 67.8/49.7
    XLM-E (without TRTD or gated bias) 75.2 67.4/49.2

    These controlled comparisons attribute gains to both the parallel-data TRTD objective and the gated relative-position mechanism, while also showing that XLM-E outperforms a matched XLM-style generative model.

  8. Knowl 8 — Compute efficiency and model scaling

    data/table

    XLM-E achieves similar or better cross-lingual understanding with substantially fewer pre-training FLOPs. The reported FLOP counts are total pre-training computation; models marked with an asterisk were continued from XLM-R, so their listed costs include the XLM-R pre-training cost. XLM-E uses 9.5×10199.5\times10^{19} FLOPs, approximately 1% of XLM-R’s 9.6×10219.6\times10^{21} FLOPs, while obtaining a higher XTREME average score. The multilingual-only XLM-E variant uses 6.3×10196.3\times10^{19} FLOPs and still exceeds XLM-R’s average score.

    Model XTREME average Parameters FLOPs
    MBERT 63.1 167M 6.4e19
    XLM-R 66.4 279M 9.6e21
    INFOXLM* – 279M 9.6e21 + 1.7e20
    XLM-ALIGN* 68.9 279M 9.6e21 + 9.6e19
    XLM-E 69.3 279M 9.5e19
    XLM-E (without TRTD) 67.6 279M 6.3e19

    Scaling the XLM-E discriminator and generator consistently improves XNLI and MLQA performance. The XL XLM-E reaches 83.783.7 XNLI accuracy and 76.2/57.976.2/57.9 MLQA F1/EM with 2.2B parameters, exceeding the larger XLM-R XL and MT5 XL baselines despite using fewer parameters.

    Model Parameters XNLI accuracy MLQA F1/EM
    XLM-E Base 279M 76.6 68.3/49.8
    XLM-E Large 840M 81.3 72.7/54.2
    XLM-E XL 2.2B 83.7 76.2/57.9
    XLM-R XL 3.5B 82.3 73.4/55.3
    MT5 XL 3.7B 82.9 73.5/54.5
  9. Knowl 9 — Cross-lingual sentence and word alignment

    empirical result

    XLM-E improves both sentence-level and word-level cross-lingual alignment, particularly when TRTD is retained. For Tatoeba sentence retrieval, sentence representations are obtained by mean pooling hidden states and translation candidates are retrieved by cosine-similarity nearest neighbors. XLM-E uses layer 9, whereas XLM-R uses layer 7. Accuracy@1 is averaged over English-to-other-language and other-language-to-English directions.

    Model Tatoeba-14 en→\toxx Tatoeba-14 xx→\toen Tatoeba-36 en→\toxx Tatoeba-36 xx→\toen
    XLM-R 59.5 57.6 55.5 53.4
    INFOXLM 80.6 77.8 68.6 67.3
    XLM-E 74.4 72.3 65.0 62.3
    XLM-E (without TRTD) 55.8 55.1 46.4 44.6

    For word alignment, predicted alignments are produced with optimal transport from layer-9 representations. The alignment error rate is

    AER=1−∣A∩S∣+∣A∩P∣∣A∣+∣S∣,\mathrm{AER}=1-\frac{|A\cap S|+|A\cap P|}{|A|+|S|},

    where AA is the predicted alignment set, SS is the annotated sure-alignment set, and PP is the annotated possible-alignment set. Lower values are better. XLM-E obtains the lowest average AER, 19.3219.32, across English–German, English–French, English–Hindi, and English–Romanian, improving over XLM-ALIGN’s 21.0521.05 and XLM-R’s 22.6422.64.

    Model en-de en-fr en-hi en-ro Average
    fast align 32.14 19.46 59.90 – –
    XLM-R 17.74 7.54 37.79 27.49 22.64
    XLM-ALIGN 16.63 6.61 33.98 26.97 21.05
    XLM-E 16.49 6.19 30.20 24.41 19.32
    XLM-E (without TRTD) 17.87 6.29 35.02 30.22 22.35
  10. Knowl 10 — Universal layer behavior and transfer gaps

    empirical result

    Layer-wise alignment analyses show that XLM-E produces more language-universal representations in upper Transformer layers than XLM-R. Sentence-retrieval accuracy follows a parabolic pattern across layers: XLM-R peaks at 54.4254.42 on layer 7, whereas XLM-E peaks at 63.6663.66 on layer 9. At layer 10, XLM-R falls to 43.3443.34, while XLM-E retains 57.1457.14. Word alignment exhibits the same qualitative pattern, with XLM-E’s best layer shifted to layer 9 rather than XLM-R’s lower best-performing layer. The authors interpret the stronger upper-layer alignment as evidence that discriminative pre-training encourages universal representations, whereas masked-language-modeling objectives tend to make upper layers more language-specific.

    The cross-lingual transfer gap is defined as English test performance minus the average performance on non-English test languages; lower values indicate better transfer. XLM-E has the lowest gap among the compared models on PAWS-X and remains reasonably competitive on the other tasks.

    Model XQuAD MLQA TyDiQA XNLI PAWS-X
    MBERT 25.0 27.5 22.2 16.5 14.1
    XLM-R 15.9 20.3 15.2 10.4 11.4
    INFOXLM – 18.8 – 10.3 –
    XLM-ALIGN 14.6 18.7 10.6 11.2 9.7
    XLM-E 14.9 19.2 13.1 11.2 8.8
    XLM-E (without TRTD) 16.3 18.6 16.3 11.5 9.6

Coverage note — Detailed per-task fine-tuning hyperparameters and the preliminary span-masking experiment were omitted as implementation and supporting details rather than load-bearing contributions.

References

  1. 1.Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020. On the cross-lingual transferability of monolingual representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4623–4637, Online. Association for Computational Linguistics.
  2. 2.Mikel Artetxe and Holger Schwenk. 2019. Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond. Transactions of the Association for Computational Linguistics, 7(0):597–610.
  3. 3.Hangbo Bao, Li Dong, Furu Wei, Wenhui Wang, Nan Yang, Xiaodong Liu, Yu Wang, Jianfeng Gao, Songhao Piao, Ming Zhou, and Hsiao-Wuen Hon. 2020. UniLMv2: Pseudo-masked language models for unified language model pre-training. In Proceedings of the 37th International Conference on Machine Learning, pages 7006–7016.
  4. 4.Steven Cao, Nikita Kitaev, and Dan Klein. 2020. Multilingual alignment of contextual word representations. In International Conference on Learning Representations.
  5. 5.Zewen Chi, Li Dong, Shuming Ma, Shaohan Huang, Xian-Ling Mao, Heyan Huang, and Furu Wei. 2021a. mT6: Multilingual pretrained text-to-text transformer with translation pairs. arXiv preprint arXiv:2104.08692.
  6. 6.Zewen Chi, Li Dong, Furu Wei, Wenhui Wang, Xian-Ling Mao, and Heyan Huang. 2020. Cross-lingual natural language generation via pre-training. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, New York, NY, USA, February 7-12, 2020, pages 7570–7577. AAAI Press.
  7. 7.Zewen Chi, Li Dong, Furu Wei, Nan Yang, Saksham Singhal, Wenhui Wang, Xia Song, Xian-Ling Mao, Heyan Huang, and Ming Zhou. 2021b. InfoXLM: An information-theoretic framework for cross-lingual language model pre-training. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3576–3588, Online. Association for Computational Linguistics.
  8. 8.Zewen Chi, Li Dong, Bo Zheng, Shaohan Huang, Xian-Ling Mao, Heyan Huang, and Furu Wei. 2021c. Improving pretrained cross-lingual language models via self-labeled word alignment. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3418–3430, Online. Association for Computational Linguistics.
  9. 9.Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. Learning phrase representations using RNN encoder–decoder for statistical machine translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1724–1734, Doha, Qatar. Association for Computational Linguistics.
  10. 10.Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020a. TyDi QA: A benchmark for information-seeking question answering in typologically diverse languages. Transactions of the Association for Computational Linguistics, 8:454–470.
  11. 11.Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020b. Electra: Pre-training text encoders as discriminators rather than generators. In International Conference on Learning Representations.
  12. 12.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzman, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  13. 13.Alexis Conneau and Guillaume Lample. 2019. Cross-lingual language model pretraining. In Advances in Neural Information Processing Systems, pages 7057–7067. Curran Associates, Inc.
  14. 14.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2475–2485, Brussels, Belgium. Association for Computational Linguistics.
  15. 15.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  16. 16.Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019. Unified language model pre-training for natural language understanding and generation. In Advances in Neural Information Processing Systems, pages 13063–13075. Curran Associates, Inc.
  17. 17.Chris Dyer, Victor Chahuneau, and Noah A Smith. 2013. A simple, fast, and effective reparameterization of ibm model 2. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 644–648.
  18. 18.Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzman, and Philipp Koehn. 2020. CCAligned: A massive collection of cross-lingual web-document pairs. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5960–5969, Online. Association for Computational Linguistics.
  19. 19.Naman Goyal, Jingfei Du, Myle Ott, Giri Anantharaman, and Alexis Conneau. 2021. Larger-scale transformers for multilingual masked language modeling. arXiv preprint arXiv:2105.00572.
  20. 20.Junjie Hu, Melvin Johnson, Orhan Firat, Aditya Siddhant, and Graham Neubig. 2020a. Explicit alignment objectives for multilingual bidirectional encoders. arXiv preprint arXiv:2010.07972.
  21. 21.Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020b. XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization. arXiv preprint arXiv:2003.11080.
  22. 22.Masoud Jalili Sabet, Philipp Dufter, François Yvon, and Hinrich Schutze. 2020. SimAlign: High quality word alignments without parallel training data using static and contextualized embeddings. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1627–1643, Online. Association for Computational Linguistics.
  23. 23.Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy. 2019. SpanBERT: Improving pre-training by representing and predicting spans. arXiv preprint arXiv:1907.10529.
  24. 24.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, San Diego, CA.
  25. 25.Taku Kudo and John Richardson. 2018. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 66–71, Brussels, Belgium. Association for Computational Linguistics.
  26. 26.Anoop Kunchukuttan, Pratik Mehta, and Pushpak Bhattacharyya. 2018. The IIT Bombay English-Hindi parallel corpus. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation, Miyazaki, Japan. European Language Resources Association.
  27. 27.Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2020. MLQA: Evaluating cross-lingual extractive question answering. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7315–7330, Online. Association for Computational Linguistics.
  28. 28.Fuli Luo, Wei Wang, Jiahao Liu, Yijia Liu, Bin Bi, Songfang Huang, Fei Huang, and Luo Si. 2020. VECO: Variable encoder-decoder pre-training for cross-lingual understanding and generation. arXiv preprint arXiv:2010.16046.
  29. 29.Shuming Ma, Li Dong, Shaohan Huang, Dongdong Zhang, Alexandre Muzio, Saksham Singhal, Hany Hassan Awadalla, Xia Song, and Furu Wei. 2021. DeltaLM: Encoder-decoder pre-training for language generation and translation by augmenting pretrained multilingual encoders. arXiv preprint arXiv:2106.13736.
  30. 30.Yu Meng, Chenyan Xiong, Payal Bajaj, Saurabh Tiwary, Paul Bennett, Jiawei Han, and Xia Song. 2021. COCO-LM: Correcting and contrasting text sequences for language model pretraining. arXiv preprint arXiv:2102.08473.
  31. 31.Franz Josef Och and Hermann Ney. 2003. A systematic comparison of various statistical alignment models. Computational linguistics, 29(1):19–51.
  32. 32.Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017. Cross-lingual name tagging and linking for 282 languages. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1946–1958, Vancouver, Canada. Association for Computational Linguistics.
  33. 33.Ankur Parikh, Oscar Tackström, Dipanjan Das, and Jakob Uszkoreit. 2016. A decomposable attention model for natural language inference. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2249–2255, Austin, Texas. Association for Computational Linguistics.
  34. 34.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  35. 35.Afshin Rahimi, Yuan Li, and Trevor Cohn. 2019. Massively multilingual transfer for NER. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 151–164, Florence, Italy. Association for Computational Linguistics.
  36. 36.Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzman. 2019. WikiMatrix: Mining 135M parallel sentences in 1620 language pairs from wikipedia. arXiv preprint arXiv:1907.05791.
  37. 37.Jorg Tiedemann. 2012. Parallel data, tools and interfaces in OPUS. In Proceedings of the Eighth International Conference on Language Resources and Evaluation, pages 2214–2218, Istanbul, Turkey. European Language Resources Association.
  38. 38.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, pages 5998–6008. Curran Associates, Inc.
  39. 39.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online. Association for Computational Linguistics.
  40. 40.Jian Yang, Shuming Ma, Dongdong Zhang, Shuangzhi Wu, Zhoujun Li, and Ming Zhou. 2020. Alternating language modeling for cross-lingual pre-training. In Thirty-Fourth AAAI Conference on Artificial Intelligence.
  41. 41.Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. 2019a. PAWS-X: A cross-lingual adversarial dataset for paraphrase identification. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3687–3692, Hong Kong, China. Association for Computational Linguistics.
  42. 42.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ Russ Salakhutdinov, and Quoc V Le. 2019b. Xlnet: Generalized autoregressive pretraining for language understanding. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  43. 43.Daniel Zeman, Joakim Nivre, Mitchell Abrams, and et al. 2019. Universal dependencies 2.5. LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (UFAL), Faculty of Mathematics and Physics, Charles University.
  44. 44.Wei Zhao, Steffen Eger, Johannes Bjerva, and Isabelle Augenstein. 2021. Inducing language-agnostic multilingual representations. In Proceedings of *SEM 2021: The Tenth Joint Conference on Lexical and Computational Semantics, pages 229–240, Online. Association for Computational Linguistics.
  45. 45.Bo Zheng, Li Dong, Shaohan Huang, Saksham Singhal, Wanxiang Che, Ting Liu, Xia Song, and Furu Wei. 2021. Allocating large vocabulary capacity for cross-lingual language model pre-training. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3203–3215, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  46. 46.Michał Ziemski, Marcin Junczys-Dowmunt, and Bruno Pouliquen. 2016. The united nations parallel corpus v1. 0. In LREC, pages 3530–3534.

Citation

MLA
Chi, Z., et al. “XLM-E: Cross-lingual Language Model Pre-training via ELECTRA”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 6170–82, https://doi.org/10.18653/v1/2022.acl-long.427.
APA
Chi, Z., Huang, S., Dong, L., Ma, S., Zheng, B., Singhal, S., Bajaj, P., Song, X., Mao, X.-L., (黄河燕), H.-Y. H., & Wei, F. (2022). XLM-E: Cross-lingual Language Model Pre-training via ELECTRA. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6170–6182. https://doi.org/10.18653/v1/2022.acl-long.427
Chicago
Chi, Z., S. Huang, L. Dong, et al. 2022. “XLM-E: Cross-lingual Language Model Pre-training via ELECTRA”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6170–82. https://doi.org/10.18653/v1/2022.acl-long.427.
Harvard
Chi, Z. et al. (2022) “XLM-E: Cross-lingual Language Model Pre-training via ELECTRA”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 6170–6182. Available at: https://doi.org/10.18653/v1/2022.acl-long.427.
Vancouver
1. Chi Z, Huang S, Dong L, et al (2022) XLM-E: Cross-lingual Language Model Pre-training via ELECTRA. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 6170–6182

BibTeX

@inproceedings{chi-etal-2022-xlm,
    title = "{XLM}-{E}: Cross-lingual Language Model Pre-training via {ELECTRA}",
    author = "Chi, Zewen  and
      Huang, Shaohan  and
      Dong, Li  and
      Ma, Shuming  and
      Zheng, Bo  and
      Singhal, Saksham  and
      Bajaj, Payal  and
      Song, Xia  and
      Mao, Xian-Ling  and
      Huang, Heyan  and
      Wei, Furu",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.427/",
    doi = "10.18653/v1/2022.acl-long.427",
    pages = "6170--6182"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/