PLUG: Leveraging Pivot Language in Cross-Lingual Instruction Tuning

Zhihan ZhangDong-Ho LeeYuwei FangWenhao YuMengzhao JiaMeng JiangFrancesco Barbieri

article2024ACL48 citations

Proposes a cross-lingual instruction tuning framework that processes queries and drafts intermediate responses in a high-resource pivot language before outputting in the target language, boosting lower-resource language performance by an average of 29%.

Listen

Large language models excel at understanding and executing complex human instructions in high-resource languages like English. However, expanding these capabilities into lower-resource languages remains a major bottleneck due to the significant imbalance in the pre-training data of foundational models. Standard approaches that train models to respond directly in the target language frequently produce subpar, unfaithful, or incomplete answers, limiting the global usability and deployment of generative artificial intelligence.

The article introduces and evaluates Pivot Language Guided Generation, a method designed to improve multilingual instruction tuning by leveraging a high-resource pivot language. The objective is to demonstrate that processing a non-English user instruction through an intermediate representation in a high-resource language enables large language models to deliver higher-quality, more reliable target-language responses.

The evaluated approach mirrors human second-language learning strategies. When given an instruction in a target language, the model is trained in a single pass to translate or comprehend the instruction in a pivot language (primarily English), draft an intermediate response in that pivot language, and finally generate the response in the target language. To rigorously evaluate open-ended performance, the researchers created a professionally translated benchmark across Chinese, Korean, Italian, and Spanish. They evaluated 13-billion-parameter foundation and instruction-tuned models across open-ended generation, factual truthfulness benchmarks, and mathematical reasoning tasks, using both automated model-based judges and blind human assessments.

The analysis reveals several decisive findings. First, the pivot-guided approach improved instruction-following quality by an average of 29% across target languages compared to standard monolingual response training, yielding a 32% net gain on the English-centric model and 28% on the multilingual model. Second, the performance benefits were greatest in lower-resource languages, reaching an average win-rate gain of 46% in Korean and 31% in Italian. Third, the method substantially increased data efficiency: models fine-tuned on just 2,000 pivot-guided examples outperformed baseline models trained on up to 96,000 standard examples. Fourth, the approach proved versatile beyond English, as other competent languages like Spanish functioned effectively as pivots for lower-resource targets. Finally, the method increased factual truthfulness by up to 39.9% and enhanced mathematical reasoning accuracy without degrading original performance in the pivot language.

These findings indicate that organizations can achieve superior multilingual performance with vastly smaller training datasets and without the complexity of separate external machine translation pipelines. By restructuring training prompts to allow models to "think" in their strongest language, enterprises can reduce development costs, shorten fine-tuning cycles, and mitigate factual inaccuracies and hallucinations in non-English customer-facing applications.

Organizations developing multilingual applications should adopt pivot-guided generation workflows when fine-tuning models on underrepresented languages. When choosing pivot languages, practitioners should align the choice with the base model's pre-training strengths and language family similarities. However, engineering teams must weigh the operational trade-off of higher generation latency, as producing intermediate pivot tokens increases total sequence length during inference.

Confidence in these findings is supported by high agreement between automated metrics and independent human evaluators. Nevertheless, stakeholders should note that the evaluation was confined to four target languages and 13-billion-parameter models. Furthermore, for extremely long input prompts, producing the intermediate pivot text may strain context window limits, requiring practitioners to consider modified variations that generate only the pivot response before producing target outputs.

  • Paper: Finetuned Language Models Are Zero-Shot Learners, Jason Wei et al. (2022). FLAN establishes instruction tuning as a way to teach pretrained models to generalize across tasks, the foundation PLUG adapts to multilingual instruction following.

No sufficiently relevant recommendations were found.

Cover for PLUG: Leveraging Pivot Language in Cross-Lingual Instruction Tuning

Abstract

Instruction tuning has remarkably advanced large language models (LLMs) in understanding and responding to diverse human instructions. Despite the success in high-resource languages, its application in lower-resource ones faces challenges due to the imbalanced foundational abilities of LLMs across different languages, stemming from the uneven language distribution in their pre-training data. To tackle this issue, we propose pivot language guided generation (PLUG), an approach that utilizes a high-resource language, primarily English, as the pivot to enhance instruction tuning in lower-resource languages. It trains the model to first process instructions in the pivot language, and then produce responses in the target language. To evaluate our approach, we introduce a benchmark, X-AlpacaEval, of instructions in 4 languages (Chinese, Korean, Italian, and Spanish), each annotated by professional translators. Our approach demonstrates a significant improvement in the instruction-following abilities of LLMs by 29% on average, compared to directly responding in the target language alone. Further experiments validate the versatility of our approach by employing alternative pivot languages beyond English to assist languages where LLMs exhibit lower proficiency.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Instruction Tuning
  • 2.2 Multilingual LLMs
  • 3 Pivot Language Guided Generation
  • 4 Evaluation Settings
  • 4.1 Benchmarks
  • 4.2 Model Settings
  • 4.3 Methods to Compare
  • 5 Results
  • 5.1 Open-Ended Instructions
  • 5.2 Study of Pivot Languages
  • 5.3 Ablation Study
  • 5.4 Data Efficiency of PLUG
  • 5.5 Comparison against Translation-Based Approaches
  • 5.6 Truthfulness & Reasoning
  • 6 Conclusion
  • References
  • A Training Prompts
  • B Evaluation Settings and Prompts
  • B.1 X-AlpacaEval
  • B.2 TruthfulQA
  • B.3 SVAMP
  • C Additional Experiments
  • C.1 Response Quality in Pivot Language
  • C.2 Inter-Annotator Agreement
  • C.3 TruthfulQA Results on PolyLM
  • C.4 Inference Efficiency
  • D Training Settings
  • E Case Study

Knowls

  1. Knowl 1 — PLUG generates a pivot-language interpretation and response before the target-language answer

    model/method

    Pivot language guided generation (PLUG) trains a language model to answer a target-language instruction using an explicit high-proficiency pivot-language intermediate. For a target-language instruction xtx^t, its pivot-language translation is xpx^p, the corresponding pivot-language answer is ypy^p, and the final target-language answer is yty^t. Instead of training only to produce yty^t from xtx^t, PLUG trains the model, in one autoregressive generation, to produce the concatenated sequence [xp;yp;yt][x^p; y^p; y^t] conditioned on xtx^t. The components are marked with language or response indicator tokens, which also make it possible to extract yty^t as the user-facing answer. The intended effect is to use the model’s stronger comprehension and response capabilities in the pivot language to guide its target-language answer.

  2. Knowl 2 — X-AlpacaEval provides professionally translated parallel instruction tests

    definition

    X-AlpacaEval is a multilingual extension of AlpacaEval for comparing open-ended instruction-following responses. It starts from 805 English instructions drawn from five test sets and provides professionally translated parallel instructions in Chinese, Korean, Italian, and Spanish; the translators were native speakers of the respective target languages. Model responses are evaluated in pairwise comparisons for how well they follow the instruction, using GPT-4 as the judge in the main evaluation. The study also used human pairwise judgments from two native speakers per example on a random sample of 80 instructions.

  3. Knowl 3 — PLUG improves open-ended instruction following over monolingual response training

    empirical result

    On X-AlpacaEval, PLUG improved response quality over monolingual response training across Chinese, Korean, Italian, and Spanish for LLaMA-2-13B, PolyLM-13B, and PolyLM-Instruct-13B. The paper reports average gains in win-loss differential of 32% for LLaMA-2, 28% for PolyLM, and 26% for PolyLM-Instruct. These are averages of the pairwise win percentage minus loss percentage, not changes in a standalone accuracy score. Gains were especially large for languages less represented in pre-training: averaging across the tested models, the reported advantage over monolingual response training was 46% for Korean and 31% for Italian. PLUG also outperformed code-switching training by reported differentials of 21% for LLaMA-2 and 19% for PolyLM. Human judgments on 80 instructions favored PLUG over monolingual training in every tested language: LLaMA-2 differentials were +32.5% Chinese, +47.5% Korean, +15.0% Italian, and +22.5% Spanish; PolyLM differentials were +18.8%, +53.8%, +8.8%, and +10.0%, respectively. Human–GPT-4 agreement was 80.6% when ties were excluded, compared with 78.0% agreement between human annotators.

  4. Knowl 4 — The evaluation compares three 13B models under shared training conditions

    experimental setup

    Experiments used LLaMA-2-13B, PolyLM-13B, and PolyLM-Instruct-13B, with English primarily serving as the pivot and Chinese, Korean, Italian, and Spanish as target languages. Instruction tuning used the 52,000-example GPT4-Alpaca dataset; ChatGPT translated its original English examples into the other languages. All models were trained for four epochs in bfloat16 with batch size 64, a peak learning rate of 5×10−65\times10^{-6}, warmup over the first 3% of steps, and linear decay afterward; inference used greedy decoding. Comparisons included pivot-only training, monolingual response training on pivot and target data, code-switching training that paired pivot instructions with target answers and vice versa, auxiliary instruction-style translation tasks, and PLUG. The same hyperparameters were used across methods.

  5. Knowl 5 — Other languages can serve as pivots, with effectiveness depending on the model and language pair

    data/table

    On PolyLM, the study compared each pivot language with the target-language monolingual response baseline, excluding pivot-language instruction-response data from all training sets so that the comparisons used the same baseline. The values below are win-loss differentials in percent, ordered by target Chinese, Korean, Italian, and Spanish; a dash marks the pivot’s own language.

    English pivot: +21.6, +54.4, +35.9, +30.3. Chinese pivot: —, +36.6, +3.1, −8.7. Korean pivot: −42.2, —, −39.4, −42.1. Italian pivot: −5.7, +36.5, —, +2.9. Spanish pivot: +4.1, +41.9, +17.5, —.

    English was the strongest pivot overall for PolyLM, whose pre-training data contained the most English. However, non-English pivots also helped some targets: Spanish was more effective than Chinese for Italian, and the paper reports an average 42% improvement for Korean when using the tested pivots. Korean, the least represented tested language in PolyLM’s pre-training data, was generally ineffective as a pivot for the other target languages.

  6. Knowl 6 — Ablations show contributions from both the pivot instruction and pivot response

    empirical result

    The PLUG-PRO ablation removes the pivot-language interpretation and generates a pivot response followed by a target response. For final target responses, PLUG’s win-loss differentials over PLUG-PRO, in Chinese, Korean, Italian, and Spanish order, were: PolyLM +8.1, +14.8, +11.7, +3.4; PolyLM-Instruct +3.9, +4.8, +4.7, −1.2; and LLaMA-2 −0.9, +2.4, +2.6, +4.8. In a separate comparison of the pivot responses within the bilingual outputs, the same PLUG-over-PLUG-PRO differentials were PolyLM +9.4, +9.2, +14.0, +2.2; PolyLM-Instruct +5.7, +4.2, +4.5, +0.9; and LLaMA-2 −0.6, 0.0, +8.1, +3.4. This indicates that interpreting the instruction in the pivot language can improve the intermediate response as well as the final target response. The pivot response itself also mattered: PLUG-PRO outperformed monolingual response training in every model-language comparison, with differentials ranging from +6.1% to +44.8%. All reported values are win-loss differentials.

  7. Knowl 7 — Small PLUG datasets can outperform much larger monolingual datasets

    empirical result

    In a LLaMA-2 evaluation on 200 randomly sampled X-AlpacaEval instructions, models trained with 2,000 PLUG examples outperformed monolingual-response baselines trained with substantially more data, including the full 96,000-example monolingual set. Increasing the PLUG training set from 2,000 to 96,000 examples further improved performance, indicating that PLUG benefited from additional data as well as performing strongly in the low-data comparison.

  8. Knowl 8 — PLUG generally exceeds external translation pipelines

    empirical result

    On PolyLM, PLUG was compared with round-trip translation, which translates the target instruction into the pivot language, generates a pivot-language answer, and translates that answer back. The reported PLUG-versus-round-trip win-loss differentials, ordered Chinese, Korean, Italian, Spanish, were +28.2, +59.3, +37.0, +38.3 when PolyLM-Instruct was the translator, and +34.4, −0.6, +14.9, +10.7 with NLLB-3.3B. The Korean result against NLLB was the exception to PLUG’s usual advantage. PLUG also exceeded direct translation of the pivot response extracted from a PLUG output: differentials were +43.6, +76.0, +68.8, +80.0 with PolyLM-Instruct and +65.5, +40.7, +39.9, +38.9 with NLLB. These results support the paper’s finding that generating the target answer depends on the preceding instruction and pivot-response context, rather than being equivalent to translating a pivot answer. Values are win-loss differentials.

  9. Knowl 9 — PLUG improves factual truthfulness and mathematical reasoning in target languages

    empirical result

    For LLaMA-2, zero-shot PLUG tuning improved scores over monolingual response training on TruthfulQA and SVAMP in all four target languages. TruthfulQA reports the percentage of answers judged both truthful and informative: monolingual versus PLUG scores were 46.4 versus 61.1 in Chinese, 43.9 versus 61.4 in Korean, 52.8 versus 57.9 in Italian, and 53.0 versus 57.9 in Spanish. On SVAMP, measured by answer accuracy, the corresponding scores were 38.5 versus 41.4, 39.1 versus 42.9, 43.0 versus 44.3, and 38.1 versus 42.7. The questions were translated into the target languages using GPT-4. The paper also reports that PolyLM improved on TruthfulQA in all four target languages, but does not report a PolyLM SVAMP result because its English accuracy was only about 10%.

  10. Knowl 10 — Long instructions and extra generated tokens are limitations of PLUG

    limitation

    The paper identifies very long instructions as a limitation because generating a lengthy pivot-language interpretation can be inefficient or exceed sequence-length constraints. PLUG-PRO, which omits that interpretation, is suggested as a possible workaround, with some loss of performance. PLUG also produces longer inference outputs because it generates intermediate pivot-language text: for Chinese, Korean, Italian, and Spanish, the reported total output-token counts were 691, 957, 638, and 647, compared with 496, 858, 380, and 360 for monolingual responses—reported increases of 39%, 11%, 68%, and 80%, respectively. The experiments covered only four target languages, which the authors attribute to the cost of GPT-4 evaluations and recruiting human evaluators.

Coverage note — Qualitative response case studies and detailed evaluator prompts are omitted because they illustrate the findings or specify operational evaluation details rather than adding separate generalizable contributions.

References

  1. 1.Kabir Ahuja, Rishav Hada, Millicent Ochieng, Prachi Jain, Harshita Diddee, Samuel Maina, Tanuja Ganu, Sameer Segal, Maxamed Axmed, Kalika Bali, and Sunayana Sitaram. 2023. MEGA: multilingual evaluation of generative AI. ArXiv preprint, 2303.12528.
  2. 2.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark, Sam McCandlish, Chris Olah, Benjamin Mann, and Jared Kaplan. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. ArXiv preprint, 2204.05862.
  3. 3.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020.
  4. 4.Lichang Chen, Shiyang Li, Jun Yan, Hai Wang, Kalpa Gunaratna, Vikas Yadav, Zheng Tang, Vijay Srinivasan, Tianyi Zhou, Heng Huang, and Hongxia Jin. 2023a. Alpagasus: Training A better alpaca with fewer data. ArXiv preprint, 2307.08701.
  5. 5.Nuo Chen, Zinan Zheng, Ning Wu, Linjun Shou, Ming Gong, Yangqiu Song, Dongmei Zhang, and Jia Li. 2023b. Breaking language barriers in multilingual mathematical reasoning: Insights and observations. ArXiv preprint, 2310.20246.
  6. 6.Pinzhen Chen, Shaoxiong Ji, Nikolay Bogoychev, Barry Haddow, and Kenneth Heafield. 2023c. Monolingual or multilingual instruction tuning: Which makes a better alpaca. ArXiv preprint, 2309.08958.
  7. 7.Zhihong Chen, Feng Jiang, Junying Chen, Tiannan Wang, Fei Yu, Guiming Chen, Hongbo Zhang, Juhao Liang, Chen Zhang, Zhiyi Zhang, Jianquan Li, Xiang Wan, Benyou Wang, and Haizhou Li. 2023d. Phoenix: Democratizing chatgpt across languages. ArXiv preprint, 2304.10453.
  8. 8.Joseph Cheung. 2023. Guanaco - generative universal assistant for natural-language adaptive context-aware omnilingual outputs. Blog Post, https://guanaco-model.github.io/.
  9. 9.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. Blog Post, https://lmsys.org/blog/2023-03-30-vicuna.
  10. 10.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2023. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research.
  11. 11.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Y. Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022. Scaling instruction-finetuned language models. ArXiv preprint, 2210.11416.
  12. 12.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020.
  13. 13.Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loïc Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. 2022. No language left behind: Scaling human-centered machine translation. ArXiv preprint, 2207.04672.
  14. 14.Shahul Es, Jithin James, Luis Espinosa Anke, and Steven Schockaert. 2023. RAGAS: automated evaluation of retrieval augmented generation. ArXiv preprint, 2309.15217.
  15. 15.Xinyang Geng, Arnav Gudibande, Hao Liu, Eric Wallace, Pieter Abbeel, Sergey Levine, and Dawn Song. 2023. Koala: A dialogue model for academic research. Blog Post, https://bair.berkeley.edu/blog/2023/04/03/koala.
  16. 16.Wenxiang Jiao, Jen-tse Huang, Wenxuan Wang, Xing Wang, Shuming Shi, and Zhaopeng Tu. 2023. Parrot: Translating during chat using large language models. ArXiv preprint, 2304.02426.
  17. 17.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. In Advances in Neural Information Processing Systems, NeurIPS 2022.
  18. 18.Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and Alexander Mattick. 2023. Openassistant conversations - democratizing large language model alignment. ArXiv preprint, 2304.07327.
  19. 19.Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. AlpacaEval: An automatic evaluator of instruction-following models. GitHub repository, https://github.com/tatsu-lab/alpaca_eval.
  20. 20.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022a. Truthfulqa: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, ACL 2022.
  21. 21.Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona T. Diab, Veselin Stoyanov, and Xian Li. 2022b. Few-shot learning with multilingual generative language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022.
  22. 22.Yuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, and Qi Zhang. 2023. Calibrating llm-based evaluator. ArXiv preprint, 2309.13308.
  23. 23.Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar, Sahaj Agarwal, Hamid Palangi, and Ahmed Hassan Awadallah. 2023. Orca: Progressive learning from complex explanation traces of GPT-4. ArXiv preprint, 2306.02707.
  24. 24.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, NeurIPS 2022.
  25. 25.Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. Are NLP models really able to solve simple math word problems? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021.
  26. 26.Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. 2023. Instruction tuning with GPT-4. ArXiv preprint, 2304.03277.
  27. 27.Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. Zero: memory optimizations toward training trillion parameter models. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC 2020.
  28. 28.Leonardo Ranaldi, Giulia Pucci, and Andre Freitas. 2023. Empowering cross-lingual abilities of instruction-tuned large language models by translation-following demonstrations. ArXiv preprint, 2308.14186.
  29. 29.Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In KDD ’20: The 26th ACM SIGKDD Conference on Knowledge Discovery and Data Mining.
  30. 30.Sebastian Ruder, Noah Constant, Jan A. Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, and Melvin Johnson. 2021. XTREME-R: towards more challenging and nuanced multilingual evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021.
  31. 31.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal V. Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Févry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M. Rush. 2022. Multitask prompted training enables zero-shot task generalization. In The Tenth International Conference on Learning Representations, ICLR 2022. OpenReview.net.
  32. 32.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatunji Ruwase, Rachel Bawden, Stas Bekman, Angelina McMillan-Major, Iz Beltagy, Huu Nguyen, Lucile Saulnier, Samson Tan, Pedro Ortiz Suarez, Victor Sanh, Hugo Laurençon, Yacine Jernite, Julien Launay, Margaret Mitchell, Colin Raffel, Aaron Gokaslan, Adi Simhi, Aitor Soroa, Alham Fikri Aji, Amit Alfassy, Anna Rogers, Ariel Kreisberg Nitzav, Canwen Xu, Chenghao Mou, Chris Emezue, Christopher Klamm, Colin Leong, Daniel van Strien, David Ifeoluwa Adelani, and et al. 2022. BLOOM: A 176b-parameter open-access multilingual language model. ArXiv preprint, 2211.05100.
  33. 33.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. Blog Post, https://crfm.stanford.edu/2023/03/13/alpaca.html.
  34. 34.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a. Llama: Open and efficient foundation language models. ArXiv preprint, 2302.13971.
  35. 35.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton-Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurélien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023b. Llama 2: Open foundation and fine-tuned chat models. ArXiv preprint, 2307.09288.
  36. 36.Peiyi Wang, Lei Li, Liang Chen, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023a. Large language models are not fair evaluators. ArXiv preprint, 2305.17926.
  37. 37.Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, David Wadden, Kelsey MacMillan, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi. 2023b. How far can camels go? exploring the state of instruction tuning on open resources. ArXiv preprint, 2306.04751.
  38. 38.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023c. Self-instruct: Aligning language models with self-generated instructions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, ACL 2023.
  39. 39.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022. Finetuned language models are zero-shot learners. In The Tenth International Conference on Learning Representations, ICLR 2022. OpenReview.net.
  40. 40.Xiangpeng Wei, Haoran Wei, Huan Lin, Tianhao Li, Pei Zhang, Xingzhang Ren, Mei Li, Yu Wan, Zhiwei Cao, Binbin Xie, Tianxiang Hu, Shangjie Li, Binyuan Hui, Bowen Yu, Dayiheng Liu, Baosong Yang, Fei Huang, and Jun Xie. 2023. Polylm: An open source polyglot large language model. ArXiv preprint, 2307.06018.
  41. 41.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023a. Wizardlm: Empowering large language models to follow complex instructions. ArXiv preprint, 2304.12244.
  42. 42.Canwen Xu, Daya Guo, Nan Duan, and Julian J. McAuley. 2023b. Baize: An open-source chat model with parameter-efficient tuning on self-chat data. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023.
  43. 43.Haoran Xu, Young Jin Kim, Amr Sharaf, and Hany Hassan Awadalla. 2023c. A paradigm shift in machine translation: Boosting translation performance of large language models. ArXiv preprint, 2309.11674.
  44. 44.Shaolei Zhang, Qingkai Fang, Zhuocheng Zhang, Zhengrui Ma, Yan Zhou, Langlin Huang, Mengyu Bu, Shangtong Gui, Yunji Chen, Xilin Chen, and Yang Feng. 2023a. Bayling: Bridging cross-lingual alignment and instruction following through interactive translation for large language models. ArXiv preprint, 2306.10968.
  45. 45.Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, and Guoyin Wang. 2023b. Instruction tuning for large language models: A survey. ArXiv preprint, 2308.10792.
  46. 46.Wenxuan Zhang, Sharifah Mahani Aljunied, Chang Gao, Yew Ken Chia, and Lidong Bing. 2023c. M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models. ArXiv preprint, 2306.05179.
  47. 47.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. ArXiv preprint, 2306.05685.
  48. 48.Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. 2023. LIMA: less is more for alignment. ArXiv preprint, 2305.11206.
  49. 49.Wenhao Zhu, Yunzhe Lv, Qingxiu Dong, Fei Yuan, Jingjing Xu, Shujian Huang, Lingpeng Kong, Jiajun Chen, and Lei Li. 2023. Extrapolating large language models to non-english by aligning languages. ArXiv preprint, 2308.04948.

Citation

MLA
Zhang, Z., et al. “PLUG: Leveraging Pivot Language in Cross-Lingual Instruction Tuning”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 7025–46, https://doi.org/10.18653/v1/2024.acl-long.379.
APA
Zhang, Z., Lee, D.-H., Fang, Y., Yu, W., Jia, M., Jiang, M., & Barbieri, F. (2024). PLUG: Leveraging Pivot Language in Cross-Lingual Instruction Tuning. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7025–7046. https://doi.org/10.18653/v1/2024.acl-long.379
Chicago
Zhang, Z., D.-H. Lee, Y. Fang, et al. 2024. “PLUG: Leveraging Pivot Language in Cross-Lingual Instruction Tuning”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7025–46. https://doi.org/10.18653/v1/2024.acl-long.379.
Harvard
Zhang, Z. et al. (2024) “PLUG: Leveraging Pivot Language in Cross-Lingual Instruction Tuning”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7025–7046. Available at: https://doi.org/10.18653/v1/2024.acl-long.379.
Vancouver
1. Zhang Z, Lee D-H, Fang Y, Yu W, Jia M, Jiang M, Barbieri F (2024) PLUG: Leveraging Pivot Language in Cross-Lingual Instruction Tuning. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7025–7046

BibTeX

@inproceedings{zhang-etal-2024-plug,
    title = "{PLUG}: Leveraging Pivot Language in Cross-Lingual Instruction Tuning",
    author = "Zhang, Zhihan  and
      Lee, Dong-Ho  and
      Fang, Yuwei  and
      Yu, Wenhao  and
      Jia, Mengzhao  and
      Jiang, Meng  and
      Barbieri, Francesco",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.379/",
    doi = "10.18653/v1/2024.acl-long.379",
    pages = "7025--7046"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/