WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models
Benjamin MinixhoferFabian PaischerNavid Rekabsaz
Presents WECHSEL, an efficient method that transfers English language models to new languages by initializing target subword embeddings via multilingual static word vectors, outperforming models trained from scratch while reducing training compute by up to 64x.
State-of-the-art language models are critical components across modern artificial intelligence applications, yet the vast majority are built primarily for English. Pretraining these massive systems from scratch for other languages requires immense computing power, financial expenditure, and environmental cost. While multilingual models attempt to bridge this gap, they often suffer performance degradation as more languages are added, leaving a clear need for high-performing monolingual models in non-English languages.
The article demonstrates a novel parameter transfer method called WECHSEL, which efficiently transfers pretrained English language models into new target languages. WECHSEL retains the deep internal representations of an existing English model and replaces the tokenizer. It then initializes the new target-language token embeddings to be semantically aligned with the English embeddings using cross-lingual static word dictionaries, targeting the roughly one-third of model parameters that are usually discarded and randomized during transfer.
The researchers evaluated this approach on two standard architectures—an encoder model (RoBERTa) and a decoder model (GPT-2)—across diverse languages. They transferred models into four medium-resource languages (French, German, Chinese, and Swahili) and four very low-resource languages (Sundanese, Scottish Gaelic, Uyghur, and Malagasy). The transferred models were compared against models trained entirely from scratch, baseline transfer techniques that initialize new token embeddings randomly, and established native monolingual and multilingual models.
The evaluation produced four key findings. First, models initialized with WECHSEL consistently outperformed both randomly initialized models and baseline transfer methods across all evaluated languages and tasks. Second, the transferred models surpassed prior native monolingual models while requiring substantially less training effort—outperforming the French model CamemBERT with 64 times less training compute and the German model GBERT with 39 times less compute. Third, WECHSEL achieved an average improvement over the high-performing multilingual model XLM-R by 3.54% accuracy in natural language inference and 1.14% in named entity recognition. Finally, the relative performance advantage of WECHSEL grew even larger in data-constrained scenarios, yielding significant quality improvements in low-resource language modeling.
These findings indicate that deep language models learn fundamental structural abstractions that generalize across human languages. Practitioners can cut pretraining timelines, computational costs, and carbon footprints dramatically by transferring existing English models rather than training new models from scratch. Organizations seeking to deploy language models in new or under-resourced languages should adopt WECHSEL as an effective initialization strategy, which also eliminates the need to freeze internal model layers during early training.
Decision-makers should note certain limitations: the method was evaluated across eight languages and focused primarily on two language understanding tasks alongside language modeling perplexity, so performance across every linguistic family or specialized downstream task cannot be guaranteed. Furthermore, because WECHSEL transfers representations directly from English source models, it risks inheriting and propagating societal biases present in the original data, warranting responsible governance and targeted audits prior to deployment.
- Paper: Enriching Word Vectors with Subword Information, Piotr Bojanowski et al. (2017). Its character n-gram word vectors provide the subword-aware static representations that make WECHSEL’s embedding initialization work for rare and unseen word forms.
- Paper: Word Translation Without Parallel Data, Alexis Conneau et al. (2017). Its unsupervised alignment of monolingual word-vector spaces supplies essential background for WECHSEL’s use of cross-lingually aligned static embeddings.
- Paper: Learning Word Vectors for 157 Languages, Edouard Grave et al. (2018). Its multilingual word vectors establish the broad-coverage static embedding resource needed to initialize a model’s target-language subword embeddings.
- Paper: Teaching Old Tokenizers New Words: Efficient Tokenizer Adaptation for Pre-trained Models, Taido Purason et al. (2025). It carries WECHSEL’s tokenizer-adaptation problem forward with continued BPE training and vocabulary pruning, offering a later approach to making pretrained tokenizers more effective across languages.
