Making Monolingual Sentence Embeddings Multilingual Using Knowledge Distillation

Nils ReimersIryna Gurevych

article2020EMNLP1,416 citations

Proposes a lightweight knowledge distillation technique that transfers monolingual sentence embedding models to over 50 languages by aligning translated sentences into a shared vector space with minimal computational cost.

Listen

Generating dense sentence embeddings—vector representations that place semantically similar sentences near each other in a mathematical space—is essential for search, clustering, and text retrieval. However, most existing models are monolingual (primarily English) because annotating high-quality training datasets for non-English languages is difficult and expensive. While existing cross-lingual approaches exist, they often require complex multi-task setups, demand heavy computational resources, or exhibit systematic biases toward specific languages.

The article evaluates an efficient method called multilingual knowledge distillation to extend existing monolingual sentence embedding systems to more than 50 languages. The primary objective is to demonstrate that a multilingual "student" model can be trained simply to replicate the output vectors of a strong monolingual "teacher" model using parallel translated sentences, successfully transferring semantic properties across languages.

To test this approach, the researchers used an English Sentence-BERT teacher model and trained a multilingual student model based on XLM-RoBERTa using standard translation datasets (such as TED talk subtitles and news commentary). The systems were evaluated across three primary tasks: cross-lingual semantic textual similarity, parallel sentence extraction from large mixed corpora (bitext retrieval), and cross-lingual similarity search on lower-resource languages.

The analysis produced several key findings. First, the student models achieved state-of-the-art results on semantic textual similarity, obtaining an average correlation score of 83.7 on cross-lingual tasks, outperforming established baselines such as LASER (67.0) and LaBSE (73.5). Second, for lower-resource languages such as Georgian, Swahili, Tagalog, and Tatar, the distillation technique improved sentence matching accuracy by roughly 30 to 40 percentage points over LASER. Third, the distilled models demonstrated virtually no language bias (a statistically insignificant drop of only 0.11 on mixed-language datasets), whereas competing models showed significant drops due to clustering sentences by language rather than meaning. Finally, experiments showed that relatively small datasets (10,000 to 25,000 sentence pairs) are sufficient to align vector spaces for similar language pairs.

These findings have immediate practical implications for engineering, computational cost, and multilingual system performance. Decoupling the creation of semantic properties from multilingual expansion allows organizations to first refine a model in high-resource languages and then port it to new languages efficiently using basic translation pairs. This removes the risk of catastrophic forgetting common in complex multi-task pipelines and reduces the hardware overhead needed to build production-grade cross-lingual search engines.

For operational decision-making, teams building multilingual search, semantic matching, or clustering pipelines should adopt multilingual distillation as a lightweight, low-risk deployment framework. However, if an organization's primary objective is strictly the mining of exact literal translations rather than assessing semantic similarity, dedicated dual-encoder systems like LASER or LaBSE remain the stronger choice, as the distilled models intentionally prioritize broad conceptual similarity over word-for-word translation equivalence.

Confidence in these findings is high for languages backed by parallel sentence training pairs. The primary limitation is that performance degrades significantly when attempting to embed languages lacking parallel training data during the distillation phase, confirming that while the technique generalizes well with minimal parallel text, it cannot reliably align entirely zero-shot languages without paired examples.

Cover for Making Monolingual Sentence Embeddings Multilingual Using Knowledge Distillation

Abstract

We present an easy and efficient method to extend existing sentence embedding models to new languages. This allows to create multilingual versions from previously monolingual models. The training is based on the idea that a translated sentence should be mapped to the same location in the vector space as the original sentence. We use the original (monolingual) model to generate sentence embeddings for the source language and then train a new system on translated sentences to mimic the original model. Compared to other methods for training multilingual sentence embeddings, this approach has several advantages: It is easy to extend existing models with relatively few samples to new languages, it is easier to ensure desired properties for the vector space, and the hardware requirements for training is lower. We demonstrate the effectiveness of our approach for 50+ languages from various language families. Code to extend sentence embeddings models to more than 400 languages is publicly available.

Table of Contents

  • 1 Introduction
  • 2 Training
  • 3 Training Data
  • 4 Experiments
  • 4.1 Multilingual Semantic Textual Similarity
  • 4.2 BUCC: Bitext Retrieval
  • 4.3 Tatoeba: Similarity Search
  • 5 Evaluation of Training Datasets
  • 6 Target Language Training
  • 7 Language Bias
  • 8 Related Work
  • 9 Conclusion
  • References
  • A Tatoeba Similarity Search

Knowls

  1. Knowl 1 — Multilingual Knowledge Distillation for Sentence Embeddings

    model/method

    Multilingual knowledge distillation is a training method designed to extend existing monolingual sentence embedding models to new languages while aligning cross-lingual vector spaces. The approach requires a pre-trained teacher model MM that produces dense sentence representations in one or more source languages ss (e.g., English), and a set of parallel sentence pairs ((s1,t1),(s2,t2),au,(sn,tn))((s_1, t_1), (s_2, t_2), au, (s_n, t_n)), where tit_i is the translation of source sentence sis_i in a target language.

    A student network M^\hat{M} (typically initialized from a multilingual pre-trained transformer with a large shared vocabulary, such as XLM-RoBERTa with a 250,000-token SentencePiece vocabulary) is trained to map both the source sentence sis_i and the translated sentence tit_i to the embedding generated by the teacher for the source sentence, i.e., M^(si)≈M(si)\hat{M}(s_i) \approx M(s_i) and M^(ti)≈M(si)\hat{M}(t_i) \approx M(s_i).

    This training framework achieves two core properties:

    1. Cross-lingual vector space alignment: Semantically equivalent sentences across different languages are mapped to proximate vector locations.
    2. Vector space property transfer: The geometric and semantic properties of the source teacher model (such as clustering behavior or paraphrase sensitivity) are transferred to target languages without requiring target-language task annotations or multi-task negative sampling.
  2. Knowl 2 — Mean Squared Error Objective for Multilingual Knowledge Distillation

    equation

    The training objective for multilingual knowledge distillation minimizes the mean squared error (MSE) between the teacher embedding of the source sentence and the student embeddings of both the source and target sentences over a mini-batch BB of parallel sentence pairs (sj,tj)(s_j, t_j):

    L(B)=1∣B∣∑j∈B[(M(sj)−M^(sj))2+(M(sj)−M^(tj))2]\mathcal{L}(B) = \frac{1}{|B|} \sum_{j \in B} \left[ \left( M(s_j) - \hat{M}(s_j) \right)^2 + \left( M(s_j) - \hat{M}(t_j) \right)^2 \right]

    where:

    • BB denotes the mini-batch of parallel sentence pairs.
    • ∣B∣|B| is the number of sentence pairs in the mini-batch.
    • sjs_j is the sentence in the source language (e.g., English).
    • tjt_j is the corresponding translated sentence in the target language.
    • M(sj)∈RdM(s_j) \in \mathbb{R}^d is the fixed sentence embedding generated by the teacher model for source sentence sjs_j.
    • M^(sj)∈Rd\hat{M}(s_j) \in \mathbb{R}^d is the student model's predicted sentence embedding for source sentence sjs_j.
    • M^(tj)∈Rd\hat{M}(t_j) \in \mathbb{R}^d is the student model's predicted sentence embedding for target sentence tjt_j.
    • (⋅)2(\cdot)^2 denotes the squared Euclidean norm (or element-wise squared error summed across the embedding dimension dd).
  3. Knowl 3 — Semantic Textual Similarity Performance of Distilled Multilingual Sentence Embeddings

    data/table

    Multilingual sentence embeddings evaluated on the STS 2017 Semantic Textual Similarity benchmark demonstrate state-of-the-art performance when trained via multilingual knowledge distillation into XLM-RoBERTa (XLM-R). Evaluated using cosine similarity scored against human judgments via Spearman rank correlation (ρ×100\rho \times 100):

    Model EN-EN ES-ES AR-AR Monolingual Avg.
    mBERT mean 54.4 56.7 50.9 54.0
    XLM-R mean 50.7 51.8 25.7 42.7
    mBERT-nli-stsb 80.2 83.9 65.3 76.5
    XLM-R-nli-stsb 78.2 83.1 64.4 75.3
    Knowledge Distillation
    mBERT ←\leftarrow SBERT-nli-stsb 82.5 83.0 78.8 81.4
    DistilmBERT ←\leftarrow SBERT-nli-stsb 82.1 84.0 77.7 81.2
    XLM-R ←\leftarrow SBERT-nli-stsb 82.5 83.5 79.9 82.0
    XLM-R ←\leftarrow SBERT-paraphrases 88.8 86.3 79.6 84.6
    Other Systems
    LASER 77.6 79.7 68.9 75.4
    mUSE 86.4 86.9 76.4 83.2
    LaBSE 79.4 80.8 69.1 76.4

    For cross-lingual STS across 7 language pairs (EN-AR, EN-DE, EN-TR, EN-ES, EN-FR, EN-IT, EN-NL):

    • XLM-R mean achieves an average of 17.8.
    • XLM-R-nli-stsb achieves an average of 55.6.
    • LASER achieves an average of 67.0.
    • LaBSE achieves an average of 73.5.
    • mUSE achieves an average of 81.1.
    • Distilled XLM-R ←\leftarrow SBERT-nli-stsb achieves an average of 77.9.
    • Distilled XLM-R ←\leftarrow SBERT-paraphrases achieves the best performance with an average of 83.7 (82.3 EN-AR, 84.0 EN-DE, 80.9 EN-TR, 83.1 EN-ES, 84.9 EN-FR, 86.3 EN-IT, 84.5 EN-NL).
  4. Knowl 4 — Language Bias in Multilingual Sentence Embedding Models

    empirical result

    Language bias is the tendency of a multilingual embedding model to map sentences of the same language closer together in vector space simply because they share a language, rather than because of shared semantic content. This bias negatively impacts tasks involving mixed-language sentence pools.

    To quantify language bias, models are evaluated on individual language pairs of the STS 2017 dataset versus a single joined pool combining all 10 language pairs. If no language bias exists, the Spearman correlation ρ\rho on the joined set matches the mean expected score of the individual pairs. A performance drop indicates that within-language sentence similarities are systematically scored higher than cross-language pairs:

    Model Expected Score Actual Score Difference
    LASER 69.5 68.6 -0.92
    mUSE 81.7 81.6 -0.19
    LaBSE 74.4 73.1 -1.29
    XLM-R ←\leftarrow SBERT-paraphrases 84.0 83.9 -0.11

    Both LASER (−0.92-0.92) and LaBSE (−1.29-1.29) suffer from statistically significant language bias (p<0.001p < 0.001), visible as discrete per-language spatial clustering in principal component projections. In contrast, multilingual knowledge distillation (−0.11-0.11) and mUSE (−0.19-0.19) exhibit negligible, statistically insignificant language bias.

  5. Knowl 5 — Evaluation on BUCC Bitext Mining and Semantic Similarity Trade-Off

    data/table

    On the BUCC bitext retrieval task, parallel sentence pairs are extracted between English and four target languages (German, French, Russian, Chinese) from corpora containing 150k--1.2M sentences with 2--3% true parallel pairs using margin-based cosine scoring. Performance is measured in F1F_1 score:

    Model DE-EN FR-EN RU-EN ZH-EN Avg.
    mBERT mean 44.1 47.2 38.0 37.4 41.7
    XLM-R mean 5.2 6.6 22.1 12.4 11.6
    mBERT-nli-stsb 38.9 39.5 26.4 30.2 33.7
    XLM-R-nli-stsb 44.0 51.0 51.5 44.0 47.6
    Knowledge Distillation
    XLM-R ←\leftarrow SBERT-nli-stsb 86.8 84.4 86.3 85.1 85.7
    XLM-R ←\leftarrow SBERT-paraphrase 90.8 87.1 88.6 87.8 88.6
    Other Systems
    mUSE 88.5 86.3 89.1 86.9 87.7
    LASER 95.4 92.4 92.3 91.7 93.0
    LaBSE 95.9 92.5 92.4 93.0 93.5

    Translation-oriented models (LASER and LaBSE) outperform semantic-similarity models (SBERT distillation and mUSE) on strict bitext extraction because distillation maps semantically similar non-exact translations close together in vector space, which BUCC labels as non-parallel.

  6. Knowl 6 — Low-Resource and Unseen Language Alignment on Tatoeba

    empirical result

    Multilingual knowledge distillation was tested on low-resource language pairs from the Tatoeba similarity search benchmark (measuring cosine retrieval accuracy for up to 1,000 English-aligned pairs in both directions) using JW300 training data. Models were evaluated on four languages with limited parallel data:

    • Georgian (KA, 296k sentence pairs)
    • Swahili (SW, 173k sentence pairs)
    • Tagalog (TL, 36k sentence pairs)
    • Tatar (TT, 119k sentence pairs)

    Notably, Tagalog and Tatar were not part of the 100 languages seen during XLM-RoBERTa's pre-training (meaning no specialized vocabulary or masked language model tuning was present for them).

    Model / Direction KA SW TL TT
    LASER
    en →\rightarrow xx 39.7 54.4 52.6 28.0
    xx →\rightarrow en 32.2 60.8 48.5 34.3
    XLM-R ←\leftarrow SBERT-nli-stsb
    en →\rightarrow xx 73.1 85.4 86.2 54.5
    xx →\rightarrow en 71.7 86.7 84.0 52.3

    Distillation improves accuracy by over 30 to 40 percentage points over LASER on low-resource pairs and generalizes effectively even to languages absent from the student model's original pre-training.

  7. Knowl 7 — Cross-Lingual Knowledge Distillation Versus Target-Language Direct Fine-Tuning

    empirical result

    Cross-lingual knowledge distillation was compared to direct supervised fine-tuning in a target language using the Korean KorNLI and KorSTS datasets. Models were evaluated on the Korean STS benchmark test set (measuring Spearman rank correlation ρ×100\rho \times 100):

    Model KO-KO
    LASER 68.44
    mUSE 76.32
    Trained on KorNLI KorSTS
    Korean RoBERTa-base 80.29
    Korean RoBERTa-large 80.49
    XLM-R 79.19
    XLM-R-large 81.84
    Multilingual Knowledge Distillation
    XLM-R ←\leftarrow SBERT-nli-stsb 81.47
    XLM-R-large ←\leftarrow SBERT-large-nli-stsb 83.00

    Distilling knowledge from an English teacher into XLM-R outperforms models fine-tuned directly on translated Korean task data. This indicates that multilingual knowledge distillation eliminates the need for target-language task datasets while simultaneously creating a joint aligned vector space across English and Korean.

  8. Knowl 8 — Impact of Training Dataset Size and Linguistic Similarity on Cross-Lingual Alignment

    empirical result

    Evaluating multilingual knowledge distillation across varied training corpora on STS 2017 demonstrates differing data requirements for linguistically similar vs. dissimilar language pairs:

    1. Structurally similar pairs with shared alphabets (English-German, EN-DE):

      • Models require very small datasets to align: training XLM-R on 1,000 TED2020 sentence pairs achieves an EN-DE Spearman ρ\rho of 71.5; 25,000 pairs reach 80.0 (near the full 483,000-pair score of 80.4).
      • Text domain has minimal impact: scores across GlobalVoices (78.1), Tatoeba (79.5), JW300 (80.0), and Europarl (78.7) remain consistent.
    2. Linguistically dissimilar pairs with distinct scripts (English-Arabic, EN-AR):

      • Models require substantially more data: 1,000 pairs yield ρ=48.4\rho = 48.4, 25,000 pairs yield 70.2, and full TED2020 (774,000 pairs) yields 78.0.
      • Text domain strongly influences performance: Tatoeba (27k pairs) achieves ρ=76.7\rho = 76.7, whereas the out-of-domain UNPC corpus (8 million pairs) achieves only ρ=66.1\rho = 66.1, demonstrating that domain alignment outweighs raw parallel data volume.

    Bilingual models outperform 10-language models by 2.2 points for EN-DE and 1.2 points for EN-AR, reflecting the curse of multilinguality where fixed model capacity is split across languages.

  9. Knowl 9 — Margin-Based Scoring Function for Bitext Mining

    equation

    To identify parallel sentence pairs between monolingual corpora while mitigating hubness and scale distortions in cosine similarity spaces, a margin-based scoring function is employed:

    score(x,y)=cos⁡(x,y)12k∑z∈NNk(x)cos⁡(x,z)+12k∑z∈NNk(y)cos⁡(y,z)\text{score}(x, y) = \frac{\cos(x, y)}{\displaystyle \frac{1}{2k} \sum_{z \in \text{NN}_k(x)} \cos(x, z) + \frac{1}{2k} \sum_{z \in \text{NN}_k(y)} \cos(y, z)}

    where:

    • x∈Rdx \in \mathbb{R}^d and y∈Rdy \in \mathbb{R}^d are the sentence embeddings of candidate sentences in the source and target corpora, respectively.
    • cos⁡(u,v)=u⋅v∥u∥∥v∥\cos(u, v) = \frac{u \cdot v}{\|u\| \|v\|} denotes the cosine similarity between two vector embeddings.
    • NNk(x)\text{NN}_k(x) is the set of the kk nearest neighbors of sentence embedding xx within the target language corpus.
    • NNk(y)\text{NN}_k(y) is the set of the kk nearest neighbors of sentence embedding yy within the source language corpus.
    • kk is the neighborhood size hyperparameter.
    • The denominator represents the average cosine similarity of xx and yy to their respective kk nearest cross-lingual neighbors.
  10. Knowl 10 — Unlabeled Parallel Pairs in the BUCC Bitext Mining Benchmark

    limitation

    The BUCC bitext mining benchmark exhibits systematic false-negative annotation noise. The dataset was constructed by pairing parallel sentences from News Commentary with Wikipedia sentences that were presumed to be non-parallel without exhaustive human verification.

    A manual evaluation of 60 false-positive German-English sentence pairs (20 top scoring false-positive predictions each from SBERT distillation, mUSE, and LASER) revealed that 57 out of 60 pairs (95%) were in fact genuine, high-quality translation pairs present in the Wikipedia split.

    As a consequence, F1F_1 scores on BUCC penalize systems that successfully retrieve unannotated parallel sentences from Wikipedia, meaning published precision and F1F_1 figures underestimate true bitext retrieval capability.

Coverage note — None was omitted; all key contributions, mathematical objectives, experimental datasets, benchmark results, and limitations are fully covered.

References

  1. 1.Űeljko Agić and Ivan Vulić. 2019. JW300: A Wide-Coverage Parallel Corpus for Low-Resource Languages. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3204–3210, Florence, Italy. Association for Computational Linguistics.
  2. 2.Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2018. Generalizing and improving bilingual word embedding mappings with a multi-step framework of linear transformations. AAAI Conference on Artificial Intelligence.
  3. 3.Mikel Artetxe and Holger Schwenk. 2019a. Margin-based Parallel Corpus Mining with Multilingual Sentence Embeddings. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3197–3203, Florence, Italy. Association for Computational Linguistics.
  4. 4.Mikel Artetxe and Holger Schwenk. 2019b. Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and Beyond. Transactions of the Association for Computational Linguistics, 7(0):597–610.
  5. 5.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal. Association for Computational Linguistics.
  6. 6.Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017. SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation. In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), pages 1–14, Vancouver, Canada. Association for Computational Linguistics.
  7. 7.Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. 2018. Universal Sentence Encoder. arXiv preprint arXiv:1803.11175.
  8. 8.Muthu Chidambaram, Yinfei Yang, Daniel Cer, Steve Yuan, Yunhsuan Sung, Brian Strope, and Ray Kurzweil. 2019. Learning Cross-Lingual Sentence Representations via a Multi-task Dual-Encoder Model. In Proceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019), pages 250–259, Florence, Italy. Association for Computational Linguistics.
  9. 9.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Unsupervised Cross-lingual Representation Learning at Scale. arXiv preprint arXiv:1911.02116.
  10. 10.Alexis Conneau, Douwe Kiela, Holger Schwenk, Loïc Barrault, and Antoine Bordes. 2017a. Supervised Learning of Universal Sentence Representations from Natural Language Inference Data. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 670–680, Copenhagen, Denmark. Association for Computational Linguistics.
  11. 11.Alexis Conneau, Guillaume Lample, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2017b. Word Translation Without Parallel Data. arXiv preprint arXiv:1710.04087.
  12. 12.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating Cross-lingual Sentence Representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2475–2485, Brussels, Belgium. Association for Computational Linguistics.
  13. 13.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv preprint arXiv:1810.04805.
  14. 14.Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. 2020. Language-agnostic BERT Sentence Embedding. arXiv preprint arXiv:2007.01852.
  15. 15.Mandy Guo, Qinlan Shen, Yinfei Yang, Heming Ge, Daniel Cer, Gustavo Hernandez Abrego, Keith Stevens, Noah Constant, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. 2018. Effective Parallel Corpus Mining using Bilingual Sentence Embeddings. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 165–176, Brussels, Belgium. Association for Computational Linguistics.
  16. 16.Jiyeon Ham, Yo Joong Choe, Kyubyong Park, Ilji Choi, and Hyungjoon Soh. 2020. KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language Understanding. arXiv preprint arXiv:2004.03289.
  17. 17.Ryan Kiros, Yukun Zhu, Ruslan R Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Skip-Thought Vectors. In C. Cortes, N. D. Lawrence, D. D. Lee, M. Sugiyama, and R. Garnett, editors, Advances in Neural Information Processing Systems 28, pages 3294–3302. Curran Associates, Inc.
  18. 18.Philipp Koehn. 2005. Europarl: A Parallel Corpus for Statistical Machine Translation. In Conference Proceedings: the tenth Machine Translation Summit, pages 79–86, Phuket, Thailand. AAMT, AAMT.
  19. 19.Guillaume Lample, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2018. Unsupervised Machine Translation Using Monolingual Corpora Only. In International Conference on Learning Representations.
  20. 20.Pierre Lison and Jörg Tiedemann. 2016. OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016), Paris, France. European Language Resources Association (ELRA).
  21. 21.Nils Reimers, Philip Beyer, and Iryna Gurevych. 2016. Task-Oriented Intrinsic Evaluation of Semantic Textual Similarity. In Proceedings of the 26th International Conference on Computational Linguistics (COLING), pages 87–96.
  22. 22.Nils Reimers and Iryna Gurevych. 2017. Reporting score distributions makes a difference: Performance study of LSTM-networks for sequence tagging. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 338–348, Copenhagen, Denmark. Association for Computational Linguistics.
  23. 23.Nils Reimers and Iryna Gurevych. 2018. Why Comparing Single Performance Scores Does Not Allow to Draw Conclusions About Machine Learning Approaches. arXiv preprint arXiv:1803.09578, abs/1803.09578.
  24. 24.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China. Association for Computational Linguistics.
  25. 25.Uma Roy, Noah Constant, Rami Al-Rfou, Aditya Barua, Aaron Phillips, and Yinfei Yang. 2020. LAReQA: Language-agnostic answer retrieval from a multilingual pool. arXiv preprint arXiv:2004.05484.
  26. 26.Sebastian Ruder. 2017. A survey of cross-lingual embedding models. arXiv preprint arXiv:1706.04902, abs/1706.04902.
  27. 27.Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzmán. 2019. WikiMatrix: Mining 135M Parallel Sentences in 1620 Language Pairs from Wikipedia. arXiv preprint arXiv:11907.05791.
  28. 28.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014. Sequence to Sequence Learning with Neural Networks. In Z. Ghahramani, M. Welling, C. Cortes, N. D. Lawrence, and K. Q. Weinberger, editors, Advances in Neural Information Processing Systems 27, pages 3104–3112. Curran Associates, Inc.
  29. 29.Jörg Tiedemann. 2012. Parallel Data, Tools and Interfaces in OPUS. In Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC’12), Istanbul, Turkey. European Language Resources Association (ELRA).
  30. 30.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122. Association for Computational Linguistics.
  31. 31.Yinfei Yang, Daniel Cer, Amin Ahmad, Mandy Guo, Jax Law, Noah Constant, Gustavo Hernández Ábrego, Steve Yuan, Chris Tar, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. 2019. Multilingual Universal Sentence Encoder for Semantic Retrieval. arXiv preprint arXiv:1907.04307, abs/1907.04307.
  32. 32.Yinfei Yang, Steve Yuan, Daniel Cer, Sheng-Yi Kong, Noah Constant, Petr Pilar, Heming Ge, Yun-hsuan Sung, Brian Strope, and Ray Kurzweil. 2018. Learning Semantic Textual Similarity from Conversations. In Proceedings of The Third Workshop on Representation Learning for NLP, pages 164–174, Melbourne, Australia. Association for Computational Linguistics.
  33. 33.Micha Ziemski, Marcin Junczys-Dowmunt, and Bruno Pouliquen. 2016. The United Nations Parallel Corpus v1.0. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016), Paris, France. European Language Resources Association (ELRA).
  34. 34.Pierre Zweigenbaum, Serge Sharoff, and Reinhar Rapp. 2018. Overview of the Third BUCC Shared Task: Spotting Parallel Sentences in Comparable Corpora. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Paris, France. European Language Resources Association (ELRA).
  35. 35.Pierre Zweigenbaum, Serge Sharoff, and Reinhard Rapp. 2017. Overview of the Second BUCC Shared Task: Spotting Parallel Sentences in Comparable Corpora. In Proceedings of the 10th Workshop on Building and Using Comparable Corpora, pages 60–67, Vancouver, Canada. Association for Computational Linguistics.

Citation

MLA
Reimers, N., and I. Gurevych. “Making Monolingual Sentence Embeddings Multilingual Using Knowledge Distillation”. arXiv, 2020, http://arxiv.org/abs/2004.09813v2.
APA
Reimers, N., & Gurevych, I. (2020). Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation. arXiv. http://arxiv.org/abs/2004.09813v2
Chicago
Reimers, N., and I. Gurevych. 2020. “Making Monolingual Sentence Embeddings Multilingual Using Knowledge Distillation”. arXiv. http://arxiv.org/abs/2004.09813v2.
Harvard
Reimers, N. and Gurevych, I. (2020) “Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2004.09813v2.
Vancouver
1. Reimers N, Gurevych I (2020) Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation. arXiv

BibTeX

@article{reimers2020making,
  title = {Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation},
  author = {Reimers, Nils and Gurevych, Iryna},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2004.09813v2},
  eprint = {2004.09813}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/