PLUG: Leveraging Pivot Language in Cross-Lingual Instruction Tuning
Zhihan ZhangDong-Ho LeeYuwei FangWenhao YuMengzhao JiaMeng JiangFrancesco Barbieri
Proposes a cross-lingual instruction tuning framework that processes queries and drafts intermediate responses in a high-resource pivot language before outputting in the target language, boosting lower-resource language performance by an average of 29%.
Large language models excel at understanding and executing complex human instructions in high-resource languages like English. However, expanding these capabilities into lower-resource languages remains a major bottleneck due to the significant imbalance in the pre-training data of foundational models. Standard approaches that train models to respond directly in the target language frequently produce subpar, unfaithful, or incomplete answers, limiting the global usability and deployment of generative artificial intelligence.
The article introduces and evaluates Pivot Language Guided Generation, a method designed to improve multilingual instruction tuning by leveraging a high-resource pivot language. The objective is to demonstrate that processing a non-English user instruction through an intermediate representation in a high-resource language enables large language models to deliver higher-quality, more reliable target-language responses.
The evaluated approach mirrors human second-language learning strategies. When given an instruction in a target language, the model is trained in a single pass to translate or comprehend the instruction in a pivot language (primarily English), draft an intermediate response in that pivot language, and finally generate the response in the target language. To rigorously evaluate open-ended performance, the researchers created a professionally translated benchmark across Chinese, Korean, Italian, and Spanish. They evaluated 13-billion-parameter foundation and instruction-tuned models across open-ended generation, factual truthfulness benchmarks, and mathematical reasoning tasks, using both automated model-based judges and blind human assessments.
The analysis reveals several decisive findings. First, the pivot-guided approach improved instruction-following quality by an average of 29% across target languages compared to standard monolingual response training, yielding a 32% net gain on the English-centric model and 28% on the multilingual model. Second, the performance benefits were greatest in lower-resource languages, reaching an average win-rate gain of 46% in Korean and 31% in Italian. Third, the method substantially increased data efficiency: models fine-tuned on just 2,000 pivot-guided examples outperformed baseline models trained on up to 96,000 standard examples. Fourth, the approach proved versatile beyond English, as other competent languages like Spanish functioned effectively as pivots for lower-resource targets. Finally, the method increased factual truthfulness by up to 39.9% and enhanced mathematical reasoning accuracy without degrading original performance in the pivot language.
These findings indicate that organizations can achieve superior multilingual performance with vastly smaller training datasets and without the complexity of separate external machine translation pipelines. By restructuring training prompts to allow models to "think" in their strongest language, enterprises can reduce development costs, shorten fine-tuning cycles, and mitigate factual inaccuracies and hallucinations in non-English customer-facing applications.
Organizations developing multilingual applications should adopt pivot-guided generation workflows when fine-tuning models on underrepresented languages. When choosing pivot languages, practitioners should align the choice with the base model's pre-training strengths and language family similarities. However, engineering teams must weigh the operational trade-off of higher generation latency, as producing intermediate pivot tokens increases total sequence length during inference.
Confidence in these findings is supported by high agreement between automated metrics and independent human evaluators. Nevertheless, stakeholders should note that the evaluation was confined to four target languages and 13-billion-parameter models. Furthermore, for extremely long input prompts, producing the intermediate pivot text may strain context window limits, requiring practitioners to consider modified variations that generate only the pivot response before producing target outputs.
- Paper: Finetuned Language Models Are Zero-Shot Learners, Jason Wei et al. (2022). FLAN establishes instruction tuning as a way to teach pretrained models to generalize across tasks, the foundation PLUG adapts to multilingual instruction following.
No sufficiently relevant recommendations were found.
