Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich Languages

Yuanchi ZhangYile WangZijun LiuShuo WangXiaolong WangPeng LiMaosong SunYang Liu

article2024ACL55 citations

Proposes SDRRL, a self-distillation framework that transfers knowledge from an LLM's own high-resource language outputs to low-resource languages, improving multilingual comprehension and generation while preserving source-language performance.

Listen

Modern large language models are predominantly pre-trained on text from a few high-resource languages such as English, leaving their performance in mid- and low-resource languages substantially behind. The common practice of translating training datasets into lower-resource languages often introduces significant translation noise and can even degrade the model's core capabilities in its primary language. To address this disparity, the article evaluates a new fine-tuning framework called Self-Distillation from Resource-Rich Languages (SDRRL), demonstrating how an AI model can transfer its own strong internal competence from high-resource languages into lower-resource languages.

The evaluated approach constructs training pairs using the model's own high-quality responses generated in a resource-rich language, translates these pairs across various source-target language combinations, and applies word-level code-switching for linguistic diversity. In addition, the framework incorporates a small set of clean external parallel text to serve as a regularization mechanism, countering machine translation noise without requiring massive new datasets. The researchers tested this method across 14 target languages using standard open-source models, including LLaMA-2 and SeaLLM, measuring performance on standardized language understanding, summarization, translation, and question-answering benchmarks.

The findings show that this self-distillation approach consistently outperforms standard fine-tuning and translation baselines across all evaluated benchmarks. Multilingual understanding accuracy improved over baselines, while bidirectional translation scores increased substantially—gaining up to roughly 6 BLEU points. Generation quality and robustness saw marked improvements, with noticeable reductions in translation errors, hallucinations, and off-target language responses where models mistakenly reply in the wrong language. Crucially, the method achieved these gains while fully preserving, and in some cases slightly improving, the model's original performance in English, a benchmark where conventional methods typically experience performance drops.

These results demonstrate that organizations can significantly enhance multilingual AI capabilities without the costly and complex requirement of massive multilingual pre-training or flawless human translation datasets. By leveraging internal knowledge transfer, teams can deploy higher-quality regional language services with lower compute and data acquisition costs. However, decision-makers should be aware that aligning lower-resource outputs to high-resource language representations carries the risk of transferring cultural norms and biases from the source language. Future work should pilot this technique across mixed multi-language settings and explore self-translation mechanisms to further reduce pipeline dependencies.

arXiv: 2402.12204THUNLP-MT/SDRRL

No sufficiently relevant recommendations were found.

Cover for Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich Languages

Abstract

While large language models (LLMs) have been pre-trained on multilingual corpora, their performance still lags behind in most languages compared to a few resource-rich languages. One common approach to mitigate this issue is to translate training data from resource-rich languages into other languages and then continue training. However, using the data obtained solely relying on translation while ignoring the original capabilities of LLMs across languages is not always effective, which we show will limit the performance of cross-lingual knowledge transfer. In this work, we propose SDRRL, a method based on Self-Distillation from Resource-Rich Languages that effectively improve multilingual performance by leveraging the internal capabilities of LLMs on resource-rich languages. We evaluate on different LLMs (LLaMA-2 and SeaLLM) and source languages (English and French) across various comprehension and generation tasks, experimental results demonstrate that SDRRL can significantly enhance multilingual capabilities while minimizing the impact on original performance in resource-rich languages.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 SFT and Translate-then-SFT Paradigm
  • 3.2 Self-Distillation from Resource-Rich Languages (SDRRL)
  • 3.2.1 Transfer Set Construction
  • 3.2.2 Transfer Set Translation
  • 3.2.3 Applying Code-Switching
  • 3.2.4 Incorporating External Parallel Corpus
  • 3.2.5 Training Objective
  • 4 Experiments
  • 4.1 Setup
  • 4.2 Main Results
  • 4.3 Ablation Study
  • 4.4 Visualization of Representation Space for Source and Target Langauges
  • 4.5 SeaLLM as Different Backbone Model
  • 4.6 Further Analysis
  • 5 Conclusion and Future Work
  • References
  • A Implementation Details
  • B More Detailed Results on SeaLLM
  • C Experiments with Non-English Language as the Source Language
  • D Case Study
  • E Off-Target Issue Analysis
  • F Potential Risks of Our Method

Knowls

  1. Knowl 1 — Self-distillation transfers a model’s resource-rich-language responses into multilingual instruction training

    model/method

    SDRRL improves a large language model’s target-language performance by using the model’s own responses in a resource-rich source language as additional supervision. For each instruction, the training data can include its original answer and the model-generated answer, together with translations that vary the language of the question and answer. SDRRL also adds a small amount of external parallel-corpus data, so that supervision is not wholly dependent on machine-translated synthetic text. The resulting fine-tuning objective combines six kinds of data: source-language questions and answers, mixed-language question-answer pairs, target-language questions and answers, machine-translation instructions, and target-language sentence-completion instructions. The method is intended both to transfer source-language capabilities to target languages and to preserve performance in the source language.

  2. Knowl 2 — SDRRL improves multilingual performance across comprehension and generation tasks

    empirical result

    In the main experiments, LLaMA-2-7B used English as its source language and was evaluated on Czech, Danish, Ukrainian, Bulgarian, Finnish, Hungarian, Norwegian, Indonesian, Japanese, Korean, Portuguese, Slovenian, Vietnamese, and Polish. The Stanford Alpaca instruction dataset supplied 52,002 English examples; NLLB-200-3.3B supplied translations. Fine-tuning ran for four epochs with early stopping. Across target-language averages, SDRRL scored 50.58 on BELEBELE, compared with 49.09 for the strongest baseline, CIT; on FLORES it scored 30.64 BLEU for target-to-English and 29.16 BLEU for English-to-target, compared with the strongest baseline scores of 24.83 and 24.68, respectively. Its FLORES COMET averages were 85.06 and 84.39 for those directions, versus strongest baseline scores of 81.65 and 80.53. On XL-SUM, SDRRL’s target-language ROUGE-1 and ROUGE-L were 21.69 and 16.18, compared with the best baseline scores of 21.14 and 15.87; on MKQA, its target-language average was 39.38 versus 38.09. SDRRL also retained or improved English performance: its BELEBELE English average was 66.29 versus 65.39 for SFT, and its English averages on XL-SUM and MKQA were higher than those of the baselines. These results support gains across comprehension, translation, summarization, and question answering, rather than gains confined to one task type.

  3. Knowl 3 — The transfer set balances ground-truth and model-generated answers

    model/method

    Let D={(xi,yi)}i=1ND=\{(x_i,y_i)\}_{i=1}^{N} be an instruction dataset in a resource-rich language, where xix_i is an instruction and yiy_i its ground-truth answer. A language model MθM_\theta generates an answer y^i=Mθ(xi)\hat{y}_i=M_\theta(x_i) for each instruction, producing G={(xi,y^i)}i=1NG=\{(x_i,\hat{y}_i)\}_{i=1}^{N}. SDRRL constructs its synthetic transfer set by sampling equally from the original examples and the model-generated examples: Dsynth=Sample⁡(D)∪Sample⁡(G)D_{\mathrm{synth}}=\operatorname{Sample}(D)\cup\operatorname{Sample}(G). Thus, the transfer set includes both reference answers and the model’s own source-language responses as supervision. The selected questions and answers are translated into the target language with a machine-translation system; translations with a WMT22-CometKiwi-DA quality score below 0.80.8 are rejected.

  4. Knowl 4 — Four source–target language pairings provide linguistically varied supervision

    definition

    For a question-answer pair, let the source language be the resource-rich language and the target language be the language being enhanced. SDRRL constructs four pair types: source-language question with source-language answer (LL); target-language question with source-language answer (TL); source-language question with target-language answer (LT); and target-language question with target-language answer (TT). The answer can be either the ground-truth response or the model-generated source-language response, with target-language answers formed by translation. These semantically corresponding but linguistically varied pairs provide sentence-level cross-language supervision; the TL and LT forms also expose the model to questions and answers in different languages. Translation quality is screened using a reference-free CometKiwi-DA score threshold of 0.80.8.

  5. Knowl 5 — Question-only code-switching adds token-level language variation

    model/method

    SDRRL applies dictionary-based code-switching to the question component of examples in each of its four language-pairing types. For each question token that has a bilingual-dictionary equivalent, the token is replaced by its equivalent with probability p=0.15p=0.15; otherwise it remains unchanged. Tokens without a dictionary entry are not replaced. Answers are left unchanged, whether they are in the source or target language. This adds token-level language variation to the sentence-level variation from translation without directly altering the answer text used for generation supervision.

  6. Knowl 6 — A small external parallel corpus supplies natural target-language supervision

    model/method

    To reduce the effect of errors in synthetic machine translations, SDRRL supplements its transfer set with a small parallel corpus P={(si,ti)}i=1LP=\{(s_i,t_i)\}_{i=1}^{L}, where sis_i is a source-language sentence and tit_i is its target-language counterpart. From this corpus it constructs machine-translation instruction data in both translation directions, and target-language sentence-completion instructions by randomly splitting target-language sentences into a prefix and completion. The machine-translation examples are denoted DmtD_{\mathrm{mt}} and the completion examples DcompD_{\mathrm{comp}}. In the LLaMA-2 experiments, 1,000 Opus100 parallel examples were sampled for each target language. The authors treat these examples as a small source of natural-language supervision and as a regularizer against deterioration from noisy synthetic translations, not as sufficient standalone fine-tuning data.

  7. Knowl 7 — The SDRRL objective equally averages the six data-subset losses

    equation

    Let U={LL,TL,LT,TT,mt,comp}\mathcal{U}=\{\mathrm{LL},\mathrm{TL},\mathrm{LT},\mathrm{TT},\mathrm{mt},\mathrm{comp}\} index the four language-pairing subsets and the machine-translation and sentence-completion subsets. For each d∈Ud\in\mathcal{U}, let DdD_d be that subset’s training examples and let ℓCE(Mθ(x),y)\ell_{\mathrm{CE}}(M_\theta(x),y) be the token-level cross-entropy for model MθM_\theta predicting answer sequence yy from input xx. SDRRL sums the mean loss within each subset, giving each subset equal weight regardless of its number of examples:

    LSDRRL=∑d∈U1∣Dd∣∑(x,y)∈DdℓCE(Mθ(x),y).\displaystyle \mathcal{L}_{\mathrm{SDRRL}}=\sum_{d\in\mathcal{U}}\frac{1}{|D_d|}\sum_{(x,y)\in D_d}\ell_{\mathrm{CE}}(M_\theta(x),y).

    Here ∣Dd∣|D_d| is the number of input-answer examples in subset dd; xx and yy are text sequences, and θ\theta denotes the language-model parameters.

  8. Knowl 8 — Ablations show that each SDRRL component contributes to performance

    empirical result

    The LLaMA-2 ablation reports target-language and English averages for natural-language understanding (NLU) and generation (NLG), in that order. The full method scored 50.58 and 66.29 on NLU, and 28.24 and 31.69 on NLG. Removing the mixed-language TL and LT subsets reduced these scores to 49.56, 65.93, 26.15, and 30.55. Replacing the synthetic transfer set with original data, thereby removing model-generated responses, gave 48.59, 65.10, 25.16, and 30.10. Removing the external translation and completion data gave 50.41, 66.01, 26.61, and 30.19; removing code-switching gave 50.37, 65.94, 27.13, and 30.69. Training only on the external translation and completion data scored 41.25, 61.61, 17.89, and 22.28. The ablations indicate that generated responses and cross-language pairings are important supervision sources, while the external data and code-switching add smaller benefits; the external parallel data alone does not account for SDRRL’s gains.

  9. Knowl 9 — SDRRL also works with SeaLLM and with French as the source language

    empirical result

    On SeaLLM-7B, the reported target-language average across BELEBELE, XL-SUM, FLORES, and MKQA was 33.01 for SDRRL, compared with 29.01 for SFT. The corresponding English average was 36.68 for SDRRL and 35.89 for SFT. In a separate LLaMA-2 experiment using French rather than English as the source language and Indonesian, Japanese, and Korean as targets, SDRRL’s gains over SFT were +2.94 average points on NLU and +1.77 on NLG; with English as source, the reported gains were +6.29 and +5.31. The French-source gains were positive but smaller, which the authors relate to LLaMA-2’s weaker French performance and lower effectiveness of the external translation system for that source language. Additional LLaMA-2-13B experiments also reported SDRRL outperforming the compared baselines.

  10. Knowl 10 — Representation and generation analyses are consistent with cross-lingual alignment

    empirical result

    For a representation-space analysis, the authors represented each LLaMA-2 instruction using the last hidden state of its final token (a 4,096-dimensional vector) and visualized the vectors in two dimensions using t-SNE. The plotted representations of semantically equivalent source- and target-language instructions were closer after SDRRL fine-tuning, consistent with the intended cross-lingual alignment. A separate off-target analysis evaluated LLaMA-2 responses on Dolly and its multilingual extension, Bactrian-X; the plotted occurrence rates declined over SDRRL training epochs across the examined languages. The paper’s examples also report fewer off-target responses and improvements in fluency and grammaticality after fine-tuning. These are supporting analyses and examples, rather than controlled measurements establishing that every individual generation error is eliminated.

  11. Knowl 11 — The method depends on external translation and may transfer source-language cultural norms

    limitation

    SDRRL relies on an external machine-translation system to create target-language examples, and its experiments include a small amount of machine-translated parallel data; translation noise therefore remains a potential constraint. The experiments primarily transfer from one source language at a time, leaving the effects of using multiple source and target languages together unexplored. The method does not modify the language-model architecture, which the authors note may limit its effectiveness for extremely low-resource languages where vocabulary expansion or continued pretraining could help. The authors also identify a potential cultural risk: aligning responses in other languages toward a high-resource language could cause those responses to reflect the high-resource language’s cultural practices or social norms.

Coverage note — The full per-language score breakdowns and individual case-study texts are omitted because the aggregate results and the main qualitative findings capture their contribution; the reported larger-model result is included without reproducing its full score tables.

References

  1. 1.Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. 2023. The belebele benchmark: a parallel reading comprehension dataset in 122 language variants.
  2. 2.Linzheng Chai, Jian Yang, Tao Sun, Hongcheng Guo, Jiaheng Liu, Bing Wang, Xiannian Liang, Jiaqi Bai, Tongliang Li, Qiyao Peng, and Zhoujun Li. 2024. xcot: Cross-lingual instruction tuning for cross-lingual chain-of-thought reasoning.
  3. 3.Nuo Chen, Zinan Zheng, Ning Wu, Ming Gong, Yangqiu Song, Dongmei Zhang, and Jia Li. 2023a. Breaking language barriers in multilingual mathematical reasoning: Insights and observations.
  4. 4.Yen-Chun Chen, Zhe Gan, Yu Cheng, Jingzhou Liu, and Jingjing Liu. 2020. Distilling knowledge learned in bert for text generation.
  5. 5.Zhihong Chen, Shuo Yan, Juhao Liang, Feng Jiang, Xiangbo Wu, Fei Yu, Guiming Hardy Chen, Junying Chen, Hongbo Zhang, Li Jianquan, Wan Xiang, and Benyou Wang. 2023b. MultilingualSIFT: Multilingual Supervised Instruction Fine-tuning.
  6. 6.Zewen Chi, Li Dong, Bo Zheng, Shaohan Huang, Xian-Ling Mao, Heyan Huang, and Furu Wei. 2021. Improving pretrained cross-lingual language models via self-labeled word alignment.
  7. 7.Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023. Free dolly: Introducing the world’s first truly open instruction-tuned llm.
  8. 8.Marta R Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al. 2022. No language left behind: Scaling human-centered machine translation. arXiv preprint arXiv:2207.04672.
  9. 9.Julen Etxaniz, Gorka Azkune, Aitor Soroa, Oier Lopez de Lacalle, and Mikel Artetxe. 2023. Do multilingual language models think better in english?
  10. 10.Gemini Team Google, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805.
  11. 11.Mitchell A. Gordon and Kevin Duh. 2019. Explaining sequence-level knowledge distillation as data-augmentation for neural machine translation.
  12. 12.Jianping Gou, Baosheng Yu, Stephen J Maybank, and Dacheng Tao. 2021. Knowledge distillation: A survey. International Journal of Computer Vision, 129:1789–1819.
  13. 13.Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. 2022. The Flores-101 evaluation benchmark for low-resource and multilingual machine translation. Transactions of the Association for Computational Linguistics, 10:522–538.
  14. 14.Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021. XL-sum: Large-scale multilingual abstractive summarization for 44 languages. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4693–4703, Online. Association for Computational Linguistics.
  15. 15.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the knowledge in a neural network.
  16. 16.Hiyouga. 2023. Llama factory. https://github.com/hiyouga/LLaMA-Factory.
  17. 17.Haoyang Huang, Tianyi Tang, Dongdong Zhang, Wayne Xin Zhao, Ting Song, Yan Xia, and Furu Wei. 2023. Not all languages are created equal in llms: Improving multilingual capability by cross-lingual-thought prompting.
  18. 18.Lianzhe Huang, Shuming Ma, Dongdong Zhang, Furu Wei, and Houfeng Wang. 2022. Zero-shot cross-lingual transfer of prompt-based tuning with a unified multilingual prompt.
  19. 19.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825.
  20. 20.Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024. Mixtral of experts.
  21. 21.Tannon Kew, Florian Schottmann, and Rico Sennrich. 2023. Turning english-centric llms into polyglots: How much multilinguality is needed?
  22. 22.Yoon Kim and Alexander M. Rush. 2016. Sequence-level knowledge distillation.
  23. 23.Guillaume Lample and Alexis Conneau. 2019. Cross-lingual language model pretraining.
  24. 24.Chong Li, Shaonan Wang, Jiajun Zhang, and Chengqing Zong. 2023a. Align after pre-train: Improving multilingual generative models with cross-lingual alignment.
  25. 25.Haonan Li, Fajri Koto, Minghao Wu, Alham Fikri Aji, and Timothy Baldwin. 2023b. Bactrian-x: Multilingual replicable instruction-following models with low-rank adaptation.
  26. 26.Junyi Li, Tianyi Tang, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2022. Pretrained language models for text generation: A survey.
  27. 27.Zongxi Li, Xianming Li, Yuzhang Liu, Haoran Xie, Jing Li, Fu lee Wang, Qing Li, and Xiaoqin Zhong. 2023c. Label supervised llama finetuning.
  28. 28.Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, and Xian Li. 2022. Few-shot learning with multilingual language models.
  29. 29.Zehui Lin, Xiao Pan, Mingxuan Wang, Xipeng Qiu, Jiangtao Feng, Hao Zhou, and Lei Li. 2021. Pre-training multilingual neural machine translation by leveraging alignment information.
  30. 30.Shayne Longpre, Yi Lu, and Joachim Daiber. 2020. Mkqa: A linguistically diverse benchmark for multilingual open domain question answering.
  31. 31.Yukun Ma, Trung Hieu Nguyen, and Bin Ma. 2022. Cpt: Cross-modal prefix-tuning for speech-to-text translation. In ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 6217–6221.
  32. 32.Zhuoyuan Mao and Yen Yu. 2024a. Tuning llms with contrastive alignment instructions for machine translation in unseen, low-resource languages.
  33. 33.Zhuoyuan Mao and Yen Yu. 2024b. Tuning llms with contrastive alignment instructions for machine translation in unseen, low-resource languages.
  34. 34.Xuan-Phi Nguyen, Wenxuan Zhang, Xin Li, Mahani Aljunied, Qingyu Tan, Liying Cheng, Guanzheng Chen, Yue Deng, Sen Yang, Chaoqun Liu, Hang Zhang, and Lidong Bing. 2023. Seallms – large language models for southeast asia.
  35. 35.Joel Niklaus, Veton Matoshi, Pooja Rani, Andrea Galassi, Matthias Stürmer, and Ilias Chalkidis. 2023. Lextreme: A multi-lingual and multi-task benchmark for the legal domain. In Findings of the Association for Computational Linguistics: EMNLP 2023. Association for Computational Linguistics.
  36. 36.Jeroen Ooms. 2024. cld3: Google’s Compact Language Detector 3. R package version 1.6.0.
  37. 37.OpenAI. 2022. ChatGPT. https://openai.com/chatgpt.
  38. 38.OpenAI. 2023. GPT-4 technical report.
  39. 39.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  40. 40.Saurabh Pahune and Manoj Chandrasekharan. 2023. Several categories of large language models (llms): A short survey. International Journal for Research in Applied Science and Engineering Technology, 11(7):615–633.
  41. 41.Minh Pham, Minsu Cho, Ameya Joshi, and Chinmay Hegde. 2022. Revisiting self-distillation.
  42. 42.Matt Post. 2018. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191, Brussels, Belgium. Association for Computational Linguistics.
  43. 43.Libo Qin, Qiguang Chen, Fuxuan Wei, Shijue Huang, and Wanxiang Che. 2023. Cross-lingual prompting: Improving zero-shot chain-of-thought reasoning across languages.
  44. 44.Leonardo Ranaldi and Giulia Pucci. 2023. Does the english matter? elicit cross-lingual abilities of large language models. In Proceedings of the 3rd Workshop on Multi-lingual Representation Learning (MRL), pages 173–183.
  45. 45.Leonardo Ranaldi and Fabio Massimo Zanzotto. 2023. Empowering multi-step reasoning across languages via tree-of-thoughts.
  46. 46.Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD ’20, page 3505–3506, New York, NY, USA. Association for Computing Machinery.
  47. 47.Ricardo Rei, José G. C. de Souza, Duarte Alves, Chrysoula Zerva, Ana C Farinha, Taisiya Glushkova, Alon Lavie, Luisa Coheur, and André F. T. Martins. 2022a. COMET-22: Unbabel-IST 2022 submission for the metrics shared task. In Proceedings of the Seventh Conference on Machine Translation (WMT), pages 578–585, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computational Linguistics.
  48. 48.Ricardo Rei, Marcos Treviso, Nuno M. Guerreiro, Chrysoula Zerva, Ana C Farinha, Christine Maroti, José G. C. de Souza, Taisiya Glushkova, Duarte Alves, Luisa Coheur, Alon Lavie, and André F. T. Martins. 2022b. CometKiwi: IST-unbabel 2022 submission for the quality estimation shared task. In Proceedings of the Seventh Conference on Machine Translation (WMT), pages 634–645, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computational Linguistics.
  49. 49.Iman Saberi, Fatemeh Fard, and Fuxiang Chen. 2024. Advfusion: Multilingual adapter-based knowledge transfer for code summarization.
  50. 50.Tal Schuster, Ori Ram, Regina Barzilay, and Amir Globerson. 2019a. Cross-lingual alignment of contextual word embeddings, with applications to zero-shot dependency parsing.
  51. 51.Tal Schuster, Ori Ram, Regina Barzilay, and Amir Globerson. 2019b. Cross-lingual alignment of contextual word embeddings, with applications to zero-shot dependency parsing.
  52. 52.Uri Shaham, Jonathan Herzig, Roee Aharoni, Idan Szpektor, Reut Tsarfaty, and Matan Eyal. 2024. Multilingual instruction tuning with just a pinch of multilinguality.
  53. 53.Haipeng Sun, Rui Wang, Kehai Chen, Masao Utiyama, Eiichiro Sumita, and Tiejun Zhao. 2020. Knowledge distillation for multilingual unsupervised neural machine translation.
  54. 54.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  55. 55.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a. Llama: Open and efficient foundation language models.
  56. 56.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023b. Llama 2: Open foundation and fine-tuned chat models.
  57. 57.Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-sne. Journal of machine learning research, 9(11).
  58. 58.Guan Wang, Sijie Cheng, Xianyuan Zhan, Xiangang Li, Sen Song, and Yang Liu. 2023a. Openchat: Advancing open-source language models with mixed-quality data.
  59. 59.Guan Wang, Sijie Cheng, Xianyuan Zhan, Xiangang Li, Sen Song, and Yang Liu. 2023b. Openchat: Advancing open-source language models with mixed-quality data. arXiv preprint arXiv:2309.11235.
  60. 60.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  61. 61.Andrea Wen-Yi and David Mimno. 2023. Hyperpolyglot llms: Cross-lingual interpretability in token embeddings. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics.
  62. 62.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  63. 63.BigScience Workshop, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model.
  64. 64.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mt5: A massively multilingual pre-trained text-to-text transformer.
  65. 65.Chenglin Yang, Lingxi Xie, Chi Su, and Alan L. Yuille. 2019. Snapshot distillation: Teacher-student optimization in one generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  66. 66.Wen Yang, Chong Li, Jiajun Zhang, and Chengqing Zong. 2023. Bigtranslate: Augmenting large language models with multilingual translation capability over 100 languages.
  67. 67.Dongkeun Yoon, Joel Jang, Sungdong Kim, Seungone Kim, Sheikh Shafayat, and Minjoon Seo. 2024. Langbridge: Multilingual reasoning without multilingual supervision.
  68. 68.Xianfeng Zeng, Yijin Liu, Ernan Li, Qiu Ran, Fandong Meng, Peng Li, Jinan Xu, and Jie Zhou. 2021. WeChat neural machine translation systems for WMT21. In Proceedings of the Sixth Conference on Machine Translation, pages 243–254, Online. Association for Computational Linguistics.
  69. 69.Biao Zhang, Philip Williams, Ivan Titov, and Rico Sennrich. 2020. Improving massively multilingual neural machine translation and zero-shot translation.
  70. 70.L. Zhang, C. Bao, and K. Ma. 2022a. Self-distillation: Towards efficient and compact neural networks. IEEE Transactions on Pattern Analysis; Machine Intelligence, 44(08):4388–4403.
  71. 71.Linfeng Zhang, Chenglong Bao, and Kaisheng Ma. 2022b. Self-distillation: Towards efficient and compact neural networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(8):4388–4403.
  72. 72.Linfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen, Chenglong Bao, and Kaisheng Ma. 2019. Be your own teacher: Improve the performance of convolutional neural networks via self distillation.
  73. 73.Yuanchi Zhang, Peng Li, Maosong Sun, and Yang Liu. 2023. Continual knowledge distillation for neural machine translation.
  74. 74.Yuanchi Zhang and Yang Liu. 2021. Directquote: A dataset for direct quotation extraction and attribution in news articles. arXiv preprint arXiv:2110.07827.
  75. 75.Jun Zhao, Zhihao Zhang, Luhui Gao, Qi Zhang, Tao Gui, and Xuanjing Huang. 2024. Llama beyond english: An empirical study on language capability transfer.
  76. 76.Wenhao Zhu, Shujian Huang, Fei Yuan, Shuaijie She, Jiajun Chen, and Alexandra Birch. 2024. Question translation training for better multilingual reasoning.

Citation

MLA
Zhang, Y., et al. “Enhancing Multilingual Capabilities of Large Language Models Through Self-Distillation from Resource-Rich Languages”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 11189–204, https://doi.org/10.18653/v1/2024.acl-long.603.
APA
Zhang, Y., Wang, Y., Liu, Z., Wang, S., Wang, X., Li, P., (孙茂松), M. S., & Liu, Y. (2024). Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich Languages. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 11189–11204. https://doi.org/10.18653/v1/2024.acl-long.603
Chicago
Zhang, Y., Y. Wang, Z. Liu, et al. 2024. “Enhancing Multilingual Capabilities of Large Language Models Through Self-Distillation from Resource-Rich Languages”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 11189–204. https://doi.org/10.18653/v1/2024.acl-long.603.
Harvard
Zhang, Y. et al. (2024) “Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich Languages”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 11189–11204. Available at: https://doi.org/10.18653/v1/2024.acl-long.603.
Vancouver
1. Zhang Y, Wang Y, Liu Z, Wang S, Wang X, Li P, (孙茂松) MS, Liu Y (2024) Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich Languages. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 11189–11204

BibTeX

@inproceedings{zhang-etal-2024-enhancing-multilingual,
    title = "Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich Languages",
    author = "Zhang, Yuanchi  and
      Wang, Yile  and
      Liu, Zijun  and
      Wang, Shuo  and
      Wang, Xiaolong  and
      Li, Peng  and
      Sun, Maosong  and
      Liu, Yang",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.603/",
    doi = "10.18653/v1/2024.acl-long.603",
    pages = "11189--11204"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/