Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich Languages
Yuanchi ZhangYile WangZijun LiuShuo WangXiaolong WangPeng LiMaosong SunYang Liu
Proposes SDRRL, a self-distillation framework that transfers knowledge from an LLM's own high-resource language outputs to low-resource languages, improving multilingual comprehension and generation while preserving source-language performance.
Modern large language models are predominantly pre-trained on text from a few high-resource languages such as English, leaving their performance in mid- and low-resource languages substantially behind. The common practice of translating training datasets into lower-resource languages often introduces significant translation noise and can even degrade the model's core capabilities in its primary language. To address this disparity, the article evaluates a new fine-tuning framework called Self-Distillation from Resource-Rich Languages (SDRRL), demonstrating how an AI model can transfer its own strong internal competence from high-resource languages into lower-resource languages.
The evaluated approach constructs training pairs using the model's own high-quality responses generated in a resource-rich language, translates these pairs across various source-target language combinations, and applies word-level code-switching for linguistic diversity. In addition, the framework incorporates a small set of clean external parallel text to serve as a regularization mechanism, countering machine translation noise without requiring massive new datasets. The researchers tested this method across 14 target languages using standard open-source models, including LLaMA-2 and SeaLLM, measuring performance on standardized language understanding, summarization, translation, and question-answering benchmarks.
The findings show that this self-distillation approach consistently outperforms standard fine-tuning and translation baselines across all evaluated benchmarks. Multilingual understanding accuracy improved over baselines, while bidirectional translation scores increased substantially—gaining up to roughly 6 BLEU points. Generation quality and robustness saw marked improvements, with noticeable reductions in translation errors, hallucinations, and off-target language responses where models mistakenly reply in the wrong language. Crucially, the method achieved these gains while fully preserving, and in some cases slightly improving, the model's original performance in English, a benchmark where conventional methods typically experience performance drops.
These results demonstrate that organizations can significantly enhance multilingual AI capabilities without the costly and complex requirement of massive multilingual pre-training or flawless human translation datasets. By leveraging internal knowledge transfer, teams can deploy higher-quality regional language services with lower compute and data acquisition costs. However, decision-makers should be aware that aligning lower-resource outputs to high-resource language representations carries the risk of transferring cultural norms and biases from the source language. Future work should pilot this technique across mixed multi-language settings and explore self-translation mechanisms to further reduce pipeline dependencies.
- Paper: Large Language Models Can Self-Improve, Jiaxin Huang et al. (2023). Its self-training framework establishes how models can generate and learn from their own outputs, the central premise that SDRRL adapts for cross-language transfer.
- Paper: Self-Instruct: Aligning Language Models with Self-Generated Instructions, Yizhong Wang et al. (2023). Its self-generated instruction data provides a useful foundation for understanding how model-produced examples can support fine-tuning when human-created data is scarce.
No sufficiently relevant recommendations were found.
