WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models

Benjamin MinixhoferFabian PaischerNavid Rekabsaz

article2022NAACL131 citations

Presents WECHSEL, an efficient method that transfers English language models to new languages by initializing target subword embeddings via multilingual static word vectors, outperforming models trained from scratch while reducing training compute by up to 64x.

Listen

State-of-the-art language models are critical components across modern artificial intelligence applications, yet the vast majority are built primarily for English. Pretraining these massive systems from scratch for other languages requires immense computing power, financial expenditure, and environmental cost. While multilingual models attempt to bridge this gap, they often suffer performance degradation as more languages are added, leaving a clear need for high-performing monolingual models in non-English languages.

The article demonstrates a novel parameter transfer method called WECHSEL, which efficiently transfers pretrained English language models into new target languages. WECHSEL retains the deep internal representations of an existing English model and replaces the tokenizer. It then initializes the new target-language token embeddings to be semantically aligned with the English embeddings using cross-lingual static word dictionaries, targeting the roughly one-third of model parameters that are usually discarded and randomized during transfer.

The researchers evaluated this approach on two standard architectures—an encoder model (RoBERTa) and a decoder model (GPT-2)—across diverse languages. They transferred models into four medium-resource languages (French, German, Chinese, and Swahili) and four very low-resource languages (Sundanese, Scottish Gaelic, Uyghur, and Malagasy). The transferred models were compared against models trained entirely from scratch, baseline transfer techniques that initialize new token embeddings randomly, and established native monolingual and multilingual models.

The evaluation produced four key findings. First, models initialized with WECHSEL consistently outperformed both randomly initialized models and baseline transfer methods across all evaluated languages and tasks. Second, the transferred models surpassed prior native monolingual models while requiring substantially less training effort—outperforming the French model CamemBERT with 64 times less training compute and the German model GBERT with 39 times less compute. Third, WECHSEL achieved an average improvement over the high-performing multilingual model XLM-R by 3.54% accuracy in natural language inference and 1.14% in named entity recognition. Finally, the relative performance advantage of WECHSEL grew even larger in data-constrained scenarios, yielding significant quality improvements in low-resource language modeling.

These findings indicate that deep language models learn fundamental structural abstractions that generalize across human languages. Practitioners can cut pretraining timelines, computational costs, and carbon footprints dramatically by transferring existing English models rather than training new models from scratch. Organizations seeking to deploy language models in new or under-resourced languages should adopt WECHSEL as an effective initialization strategy, which also eliminates the need to freeze internal model layers during early training.

Decision-makers should note certain limitations: the method was evaluated across eight languages and focused primarily on two language understanding tasks alongside language modeling perplexity, so performance across every linguistic family or specialized downstream task cannot be guaranteed. Furthermore, because WECHSEL transfers representations directly from English source models, it risks inheriting and propagating societal biases present in the original data, warranting responsible governance and targeted audits prior to deployment.

arXiv: 2112.06598CPJKU/wechsel
  • Paper: Enriching Word Vectors with Subword Information, Piotr Bojanowski et al. (2017). Its character n-gram word vectors provide the subword-aware static representations that make WECHSEL’s embedding initialization work for rare and unseen word forms.
  • Paper: Word Translation Without Parallel Data, Alexis Conneau et al. (2017). Its unsupervised alignment of monolingual word-vector spaces supplies essential background for WECHSEL’s use of cross-lingually aligned static embeddings.
  • Paper: Learning Word Vectors for 157 Languages, Edouard Grave et al. (2018). Its multilingual word vectors establish the broad-coverage static embedding resource needed to initialize a model’s target-language subword embeddings.
Cover for WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models

Abstract

Large pretrained language models (LMs) have become the central building block of many NLP applications. Training these models requires ever more computational resources and most of the existing models are trained on English text only. It is exceedingly expensive to train these models in other languages. To alleviate this problem, we introduce a novel method – called WECHSEL – to efficiently and effectively transfer pretrained LMs to new languages. WECHSEL can be applied to any model which uses subword-based tokenization and learns an embedding for each subword. The tokenizer of the source model (in English) is replaced with a tokenizer in the target language and token embeddings are initialized such that they are semantically similar to the English tokens by utilizing multilingual static word embeddings covering English and the target language. We use WECHSEL to transfer the English RoBERTa and GPT-2 models to four languages (French, German, Chinese and Swahili). We also study the benefits of our method on very low-resource languages. WECHSEL improves over proposed methods for cross-lingual parameter transfer and outperforms models of comparable size trained from scratch with up to 64x less training effort. Our method makes training large language models for new languages more accessible and less damaging to the environment. We make our code and models publicly available.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Subword Embedding Computation
  • 3.2 Subword similarity-based Transfer
  • 4 Experiment Design
  • 5 Results
  • 5.1 Transferring RoBERTa
  • 5.2 Transferring GPT-2
  • 5.2.1 To Medium-Resource Languages
  • 5.2.2 To Low-Resource Languages
  • 5.3 Is freezing necessary?
  • 6 Limitations and Potential Risks
  • 6.1 Limitations
  • 6.2 Risks
  • 7 Conclusion
  • 8 Acknowledgments
  • References
  • A Grid search over k and τ
  • B Hyperparameters
  • C Qualitative subword correspondence
  • D Using Word Embeddings without subword information
  • E Choosing a transfer baseline
  • F Sensitivity Analysis w. r. t. training data size

Knowls

  1. Knowl 1 — WECHSEL transfers a monolingual model by replacing its lexical interface

    model/method

    WECHSEL converts a pretrained source-language language model into a target-language model by replacing the source tokenizer with a target-language tokenizer, copying the source model’s non-token-embedding parameters, and initializing the target token embeddings from cross-lingual semantic correspondences. It then continues training on target-language text using the model’s original objective: masked language modeling for RoBERTa or causal language modeling for GPT-2. The method is intended to create a target-language monolingual model, rather than preserve the model’s source-language capabilities.

  2. Knowl 2 — Subword vectors are composed from aligned word-embedding n-grams

    model/method

    For each source- or target-language tokenizer token xx, WECHSEL constructs a static subword vector in the aligned multilingual word-embedding space by summing the vectors of its known character n-grams, following the fastText composition scheme:

    ux=∑g∈G(x)wg.u_x = \sum_{g \in G(x)} w_g.

    Here, G(x)G(x) is the set of character n-grams occurring in token xx, wgw_g is the static embedding of n-gram gg, and uxu_x is the resulting subword vector. The computation is performed independently for the source and target vocabularies. If no n-gram for a token is known, its static subword vector is set to zero.

  3. Knowl 3 — Target token embeddings are initialized from similar source tokens

    equation

    WECHSEL compares target and source tokenizer tokens using cosine similarity between their static subword vectors, then initializes each target model embedding as a softmax-weighted average of the embeddings of its kk most similar source tokens:

    sx,y=uxt(uys)T∥uxt∥ ∥uys∥,ext=∑y∈Jxexp⁡(sx,y/τ) eys∑y′∈Jxexp⁡(sx,y′/τ).s_{x,y} = \frac{u_x^t (u_y^s)^T}{\lVert u_x^t \rVert\,\lVert u_y^s \rVert}, \qquad e_x^t = \frac{\sum_{y \in J_x} \exp(s_{x,y}/\tau)\,e_y^s}{\sum_{y' \in J_x} \exp(s_{x,y'}/\tau)}.

    Here, xx is a target tokenizer token and yy is a source tokenizer token; uxtu_x^t and uysu_y^s are their static subword vectors in the shared multilingual embedding space; sx,ys_{x,y} is their cosine similarity; exte_x^t and eyse_y^s are their target- and source-model token embeddings; and JxJ_x is the set of the kk source tokens most similar to xx. The temperature is τ\tau. The paper’s experiments use k=10k=10 and τ=0.1\tau=0.1. If a target token’s static vector is zero, its model embedding is instead sampled from a normal distribution whose mean and variance are those of the source model’s token-embedding values.

  4. Knowl 4 — Transfer experiments use aligned fastText vectors and matched training budgets

    experimental setup

    The experiments transfer 125-million-parameter English RoBERTa and 117-million-parameter English GPT-2 models. RoBERTa is transferred to French, German, Chinese, and Swahili; GPT-2 is transferred to those languages and to Sundanese, Scottish Gaelic, Uyghur, and Malagasy. Target tokenizers are byte-level BPE tokenizers with 50,000 tokens. Monolingual fastText vectors are aligned with orthogonal Procrustes mappings using bilingual dictionaries: MUSE dictionaries for French, German, and Chinese, FreeDict for Swahili, and dictionaries collected from Wiktionary for the four low-resource languages.

    French, German, and Chinese training text comes from 4 GiB subsets of OSCAR. CC-100 supplies 1.6 GiB for Swahili, 0.1 GiB each for Sundanese and Scottish Gaelic, 0.4 GiB for Uyghur, and 0.2 GiB for Malagasy. Main model training uses 250,000 steps, batch size 512, sequence length 512, and a linear learning-rate warmup over the first 10% of steps followed by linear decay; the peak learning rates are 10−410^{-4} for RoBERTa and 5×10−45\times10^{-4} for GPT-2. Training one model takes about four days on a TPUv3-8. The comparisons are WECHSEL, FullRand (training a model from random initialization), and TransInner (copying non-token-embedding parameters but randomly initializing token embeddings). TransInner trains embeddings alone for its first 25,000 steps before training the whole model.

    RoBERTa is evaluated after target-language pretraining by fine-tuning on XNLI and WikiANN, measuring NLI accuracy and NER micro-F1, respectively; reported downstream scores average three runs. GPT-2 is evaluated by perplexity on held-out text. Low-resource GPT-2 runs are assessed at 2,500-step intervals and stopped when held-out perplexity fails to improve for 10,000 steps.

  5. Knowl 5 — WECHSEL improves RoBERTa transfer across four languages and is competitive with much longer-trained models

    empirical result

    After 250,000 target-language pretraining steps, WECHSEL-RoBERTa exceeds both TransInner-RoBERTa and FullRand-RoBERTa in the reported NLI-accuracy and NER-micro-F1 averages for each language. The tuples below give (NLI accuracy, NER micro-F1) in percent, in the order WECHSEL, TransInner, FullRand:

    • French: (82.43, 90.88), (81.75, 90.34), (75.28, 89.30).
    • German: (81.79, 89.72), (80.75, 89.30), (75.48, 88.36).
    • Chinese: (78.32, 80.55), (76.99, 80.00), (71.38, 78.35).
    • Swahili: (75.05, 87.39), (74.10, 87.05), (70.34, 87.34).

    WECHSEL also reaches strong performance early. After 25,000 pretraining steps, its reported average of the two task scores is 85.95 for French, 85.08 for German, 78.13 for Chinese, and 80.75 for Swahili. In French and German, those scores exceed the corresponding full-training averages of the earlier monolingual models CamemBERT and GBERTBase, respectively, despite using about 64 times and 39 times fewer target-language training tokens. At 250,000 steps, WECHSEL’s average score exceeds XLM-RBase by 3.54 percentage points on NLI and 1.14 points on NER. It also beats the prior monolingual models on NLI for French, German, and Chinese, and on NER for French and German; BERTBase-Chinese remains 1.50 points higher on NER than WECHSEL for Chinese.

  6. Knowl 6 — WECHSEL improves GPT-2 perplexity on medium-resource languages

    empirical result

    For French, German, Chinese, and Swahili, WECHSEL-GPT-2 obtains lower held-out perplexity after 250,000 steps than both TransInner-GPT-2 and FullRand-GPT-2. Each tuple below gives perplexity before target-language training, after 25,000 steps, and after 250,000 steps, in the order WECHSEL, TransInner, FullRand:

    • French: WECHSEL (1.7e+3, 23.47, 19.71); TransInner (1.4e+5, 67.97, 20.13); FullRand (5.9e+4, 25.99, 20.47).
    • German: WECHSEL (3.7e+3, 34.35, 26.80); TransInner (1.5e+5, 121.67, 27.76); FullRand (5.8e+4, 37.29, 27.63).
    • Chinese: WECHSEL (2.4e+4, 71.02, 51.97); TransInner (1.5e+5, 231.05, 56.17); FullRand (5.8e+4, 69.29, 52.98).
    • Swahili: WECHSEL (1.4e+5, 13.02, 10.14); TransInner (1.4e+5, 42.95, 10.28); FullRand (5.8e+4, 13.22, 10.58).

    Lower perplexity is better. The training curves show that WECHSEL is consistently better than the two baselines through training for French and German. For Chinese, FullRand has lower perplexity at 25,000 steps, but WECHSEL finishes with lower perplexity.

  7. Knowl 7 — GPT-2 transfer remains effective in low-resource settings, with larger gains at smaller data sizes

    empirical result

    On four low-resource languages, the best held-out perplexities are lower for WECHSEL-GPT-2 than for either baseline. The results are: Sundanese, 111.72 versus 151.86 for TransInner and 149.46 for FullRand; Scottish Gaelic, 16.43 versus 18.62 and 19.53; Uyghur, 34.33 versus 39.06 and 42.82; Malagasy, 14.01 versus 14.85 and 15.93. Low-resource models can overfit early, so the reported values are the best observed during training rather than scores at a fixed final step.

    A French data-size analysis likewise finds that WECHSEL’s advantage over the baselines grows as training text shrinks. Best perplexities at 16, 64, 256, and 1,024 MiB, respectively, are: WECHSEL using the original fastText vectors, 78.33, 44.75, 31.63, 24.66; WECHSEL using fastText vectors trained only on the matching text subsample, 97.42, 49.50, 32.88, 24.75; FullRand, 281.46, 83.43, 43.08, 27.09; and TransInner, 216.37, 77.71, 35.27, 25.15. Thus, reducing the data used to train the fastText vectors worsens WECHSEL’s results, but its perplexity remains lower than the two baselines at each tested training-text size.

  8. Knowl 8 — WECHSEL does not require an initial freeze of non-embedding parameters

    empirical result

    In a German GPT-2 experiment trained for 75,000 steps, the authors compare WECHSEL and TransInner with and without freezing non-token-embedding parameters for the first 10% of training. The perplexity curves indicate that freezing is needed for TransInner: its no-freeze variant performs worse than its frozen variant. WECHSEL’s frozen and unfrozen curves are comparable, leading the authors to conclude that this initial freeze is unnecessary for WECHSEL. This finding is specific to the tested German GPT-2 setup.

  9. Knowl 9 — Frequency-weighted word vectors provide a fallback when n-gram vectors are unavailable

    model/method

    The paper’s WECHSEL-TFR variant computes tokenizer-token vectors without requiring n-gram information in the static embeddings, but it requires word frequencies. It tokenizes each word vv in the static word-embedding vocabulary, gathers the words V(x)V(x) whose tokenization contains tokenizer token xx, and averages their word vectors with frequency weights:

    ux=∑v∈V(x)wvfv∑v∈V(x)fv.u_x = \frac{\sum_{v \in V(x)} w_v f_v}{\sum_{v \in V(x)} f_v}.

    Here, wvw_v is the static embedding of word vv, fvf_v is its frequency, and uxu_x is the resulting vector for token xx. In experiments, WECHSEL-TFR is broadly on par with the n-gram-based method: final GPT-2 perplexities (WECHSEL, WECHSEL-TFR) are 19.71 and 19.70 for French, 26.80 and 26.82 for German, 51.97 and 52.07 for Chinese, and 10.14 and 10.06 for Swahili.

  10. Knowl 10 — Language and task coverage limits the scope of the evidence, and transferred models may inherit biases

    limitation

    The experiments cover at most eight target languages and evaluate extrinsic language understanding on only NLI and NER; the authors note that computational and multilingual-data constraints prevent establishing whether similar gains hold for other languages or tasks. They also warn that because WECHSEL transfers parameters from English models, biases and stereotypes encoded in those source models are likely to carry over to the target-language models. The paper therefore recommends responsible use of transferred models.

Coverage note — Detailed token-correspondence examples and the full linear-probe grid-search diagnostics are omitted because they support the method and hyperparameter choice rather than constitute standalone contributions.

References

  1. 1.Wissam Antoun, Fady Baly, and Hazem Hajj. 2020. AraBERT: Transformer-based model for Arabic language understanding. In Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools, with a Shared Task on Offensive Language Detection, pages 9–15, Marseille, France. European Language Resource Association.
  2. 2.Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2016. Learning principled bilingual mappings of word embeddings while preserving monolingual invariance. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2289–2294, Austin, Texas. Association for Computational Linguistics.
  3. 3.Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2017. Learning bilingual word embeddings with (almost) no bilingual data. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 451–462, Vancouver, Canada. Association for Computational Linguistics.
  4. 4.Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2018. A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 789–798, Melbourne, Australia. Association for Computational Linguistics.
  5. 5.Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020. On the cross-lingual transferability of monolingual representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4623–4637, Online. Association for Computational Linguistics.
  6. 6.Piotr Banski and Beata Wójtowicz. 2009. Freedict: an ´ open source repository of tei-encoded bilingual dictionaries.
  7. 7.Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 610–623.
  8. 8.Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017. Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5:135–146.
  9. 9.Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. In Proc. of NeurIPS.
  10. 10.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165.
  11. 11.Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science.
  12. 12.Branden Chan, Stefan Schweter, and Timo Möller. 2020. German’s next language model. In Proceedings of the 28th International Conference on Computational Linguistics, pages 6788–6796, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  13. 13.Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. 2020. Electra: Pre-training text encoders as discriminators rather than generators. arXiv preprint arXiv:2003.10555.
  14. 14.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  15. 15.Alexis Conneau, Guillaume Lample, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2017. Word translation without parallel data. arXiv preprint arXiv:1710.04087.
  16. 16.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel R. Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. Xnli: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics.
  17. 17.Wietse de Vries and Malvina Nissim. 2021. As good as new. how to successfully recycle English GPT-2 to make models for other languages. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 836–846, Online. Association for Computational Linguistics.
  18. 18.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  19. 19.Long Duong, Hiroshi Kanayama, Tengfei Ma, Steven Bird, and Trevor Cohn. 2016. Learning crosslingual word embeddings without bilingual corpora. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1285–1295, Austin, Texas. Association for Computational Linguistics.
  20. 20.Yanai Elazar and Yoav Goldberg. 2018. Adversarial removal of demographic attributes from text data. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing.
  21. 21.William Fedus, Barret Zoph, and Noam Shazeer. 2021. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. arXiv preprint arXiv:2101.03961.
  22. 22.Christian Ganhör, David Penz, Navid Rekabsaz, Oleg Lesota, and Markus Schedl. 2022. Mitigating consumer biases in recommendations with adversarial training. In Proceedings of the 45th International ACM SIGIR conference on research and development in Information Retrieval, SIGIR 2022. ACM.
  23. 23.Armand Joulin, Piotr Bojanowski, Tomas Mikolov, Hervé Jégou, and Edouard Grave. 2018. Loss in translation: Learning bilingual word mapping with a retrieval criterion. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2979–2984, Brussels, Belgium. Association for Computational Linguistics.
  24. 24.Klara Krieg, Emilia Parada-Cabaleiro, Markus Schedl, and Navid Rekabsaz. 2022. Do perceived gender biases in retrieval results affect relevance judgements? In Proceedings of the Workshop on Algorithmic Bias in Search and Recommendation at the European Conference on Information Retrieval (ECIR-BIAS 2022).
  25. 25.Guillaume Lample, Alexis Conneau, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2018. Word translation without parallel data. In International Conference on Learning Representations.
  26. 26.Mike Lewis, Marjan Ghazvininejad, Gargi Ghosh, Armen Aghajanyan, Sida Wang, and Luke Zettlemoyer. 2020. Pre-training via paraphrasing. In Advances in Neural Information Processing Systems, volume 33, pages 18470–18481. Curran Associates, Inc.
  27. 27.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  28. 28.Antoine Louis. 2020. BelGPT-2: a GPT-2 model pre-trained on French corpora. https://github.com/antoiloui/belgpt2.
  29. 29.Thang Luong, Hieu Pham, and Christopher D. Manning. 2015. Bilingual word representations with monolingual quality in mind. In Proceedings of the 1st Workshop on Vector Space Modeling for Natural Language Processing, pages 151–159, Denver, Colorado. Association for Computational Linguistics.
  30. 30.Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric de la Clergerie, Djamé Seddah, and Benoît Sagot. 2020. CamemBERT: a tasty French language model. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7203–7219, Online. Association for Computational Linguistics.
  31. 31.Alessandro B. Melchiorre, Navid Rekabsaz, Emilia Parada-Cabaleiro, Stefan Brandl, Oleg Lesota, and Markus Schedl. 2021. Investigating gender fairness of recommendation algorithms in the music domain. Information Processing and Management, 58(5):102666.
  32. 32.Toan Q. Nguyen and David Chiang. 2017. Transfer learning across low-resource, related languages for neural machine translation. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 296–301, Taipei, Taiwan. Asian Federation of Natural Language Processing.
  33. 33.Debora Nozza, Federico Bianchi, and Dirk Hovy. 2020. What the [mask]? making sense of language-specific bert models. arXiv preprint arXiv:2003.02912.
  34. 34.Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017. Cross-lingual name tagging and linking for 282 languages. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1946–1958, Vancouver, Canada. Association for Computational Linguistics.
  35. 35.Telmo Pires, Eva Schlinger, and Dan Garrette. 2019. How multilingual is multilingual BERT? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4996–5001, Florence, Italy. Association for Computational Linguistics.
  36. 36.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding with unsupervised learning.
  37. 37.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  38. 38.Afshin Rahimi, Yuan Li, and Trevor Cohn. 2019. Massively multilingual transfer for NER. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 151–164, Florence, Italy. Association for Computational Linguistics.
  39. 39.Ori Ram, Yuval Kirstain, Jonathan Berant, Amir Globerson, and Omer Levy. 2021. Few-shot question answering by pretraining span selection. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3066–3079, Online. Association for Computational Linguistics.
  40. 40.Navid Rekabsaz, Simone Kopeinik, and Markus Schedl. 2021a. Societal biases in retrieved contents: Measurement framework and adversarial mitigation of bert rankers. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, page 306–316.
  41. 41.Navid Rekabsaz, Nikolaos Pappas, James Henderson, Banriskhem K Khonglah, and Srikanth Madikeri. 2019. Regularization advantages of multilingual neural language models for low resource domains. arXiv preprint arXiv:1906.01496.
  42. 42.Navid Rekabsaz and Markus Schedl. 2020. Do neural ranking models intensify gender bias? In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 2065–2068.
  43. 43.Navid Rekabsaz, Robert West, James Henderson, and Allan Hanbury. 2021b. Measuring societal biases from text corpora with smoothed first-order co-occurrence. In Proceedings of the Fifteenth International AAAI Conference on Web and Social Media, ICWSM 2021, held virtually, June 7-10, 2021, pages 549–560. AAAI Press.
  44. 44.Phillip Rust, Jonas Pfeiffer, Ivan Vulic, Sebastian Ruder, and Iryna Gurevych. 2021. How good is your tokenizer? on the monolingual performance of multilingual language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3118–3135, Online. Association for Computational Linguistics.
  45. 45.Peter H Schönemann. 1966. A generalized solution of the orthogonal procrustes problem. Psychometrika, 31(1):1–10.
  46. 46.Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019. Energy and policy considerations for deep learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3645–3650, Florence, Italy. Association for Computational Linguistics.
  47. 47.Ke Tran. 2020. From english to foreign languages: Transferring pre-trained language models. arXiv preprint arXiv:2002.07306.
  48. 48.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  49. 49.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  50. 50.Shijie Wu and Mark Dredze. 2020. Are all languages created equal in multilingual BERT? In Proceedings of the 5th Workshop on Representation Learning for NLP, pages 120–130, Online. Association for Computational Linguistics.
  51. 51.Chao Xing, Dong Wang, Chao Liu, and Yiye Lin. 2015. Normalized word embedding and orthogonal transform for bilingual word translation. In Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1006–1011, Denver, Colorado. Association for Computational Linguistics.
  52. 52.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online. Association for Computational Linguistics.
  53. 53.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019. Xlnet: Generalized autoregressive pretraining for language understanding. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  54. 54.George Zerveas, Navid Rekabsaz, Daniel Cohen, and Carsten Eickhoff. 2022. Mitigating bias in search results through set-based document reranking and neutrality regularization. In Proceedings of the 45th International ACM SIGIR conference on research and development in Information Retrieval, SIGIR 2022. ACM.
  55. 55.Barret Zoph, Deniz Yuret, Jonathan May, and Kevin Knight. 2016. Transfer learning for low-resource neural machine translation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1568–1575, Austin, Texas. Association for Computational Linguistics.

Citation

MLA
Minixhofer, B., et al. “WECHSEL: Effective Initialization of Subword Embeddings for Cross-lingual Transfer of Monolingual Language Models”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 3992–4006, https://doi.org/10.18653/v1/2022.naacl-main.293.
APA
Minixhofer, B., Paischer, F., & Rekabsaz, N. (2022). WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3992–4006. https://doi.org/10.18653/v1/2022.naacl-main.293
Chicago
Minixhofer, B., F. Paischer, and N. Rekabsaz. 2022. “WECHSEL: Effective Initialization of Subword Embeddings for Cross-lingual Transfer of Monolingual Language Models”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3992–4006. https://doi.org/10.18653/v1/2022.naacl-main.293.
Harvard
Minixhofer, B., Paischer, F. and Rekabsaz, N. (2022) “WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 3992–4006. Available at: https://doi.org/10.18653/v1/2022.naacl-main.293.
Vancouver
1. Minixhofer B, Paischer F, Rekabsaz N (2022) WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 3992–4006

BibTeX

@inproceedings{minixhofer-etal-2022-wechsel,
    title = "{WECHSEL}: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models",
    author = "Minixhofer, Benjamin  and
      Paischer, Fabian  and
      Rekabsaz, Navid",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.293/",
    doi = "10.18653/v1/2022.naacl-main.293",
    pages = "3992--4006"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/