Improving In-context Learning of Multilingual Generative Language Models with Cross-lingual Alignment

Chong LiShaonan WangJiajun ZhangChengqing Zong

article2024NAACL58 citations

Proposes an efficient post-training alignment framework combining multilingual contrastive learning and cross-lingual instruction following to bridge the performance gap between high- and low-resource languages using fewer than one million parallel samples.

Listen

Modern multilingual artificial intelligence models exhibit notable performance disparities across languages, frequently favoring high-resource languages such as English over less-represented ones. This imbalance stems from uneven training data and internally isolated language representations, which impede the model's ability to transfer learned knowledge across languages. Given the prohibitive computational cost of retraining massive foundation models from scratch, there is an urgent operational need for lightweight, post-pretraining alignment techniques that bridge this cross-lingual gap.

The article demonstrates and evaluates Align aFter Pre-training (AFP), a framework designed to enhance the cross-lingual in-context learning capabilities of generative language models using limited parallel translation data. The core objective is to align internal sentence representations and generation outputs across languages, thereby reducing performance disparities and enabling efficient cross-lingual knowledge transfer.

To achieve this, the authors implemented a two-part approach evaluated across multiple model architectures (XGLM, BLOOM, and Llama) spanning sizes from 560 million to 7.5 billion parameters. The first component, Multilingual Contrastive Learning, pulls the internal representations of translation pairs closer together within the model's early transformer layers. The second component, Cross-lingual Instruction Following, trains the model to respond in a specified target language given a prompt in a source language. The experiments utilized fewer than one million parallel translation samples—totaling approximately 20 million tokens, or less than 0.05 per mille of standard pretraining volume—and tested performance across standardized understanding, reasoning, and translation benchmarks covering up to 52 languages.

The findings show consistent and significant performance gains across all evaluated benchmarks. In bilingual English-Chinese testing, the framework improved overall accuracy by an average of 3.31%, with understanding tasks improving by 4.28% and reasoning by 2.67%. When scaled across 52 languages using English as a pivot language, the method yielded an average performance boost of 2.6% while narrowing performance variance across high-, medium-, and low-resource languages. Furthermore, the framework notably improved zero-shot translation quality and generalized effectively to unseen languages; for example, BLOOM's performance on non-pretrained languages such as Thai and Turkish improved by nearly 4%. An ablation analysis confirmed that combining internal contrastive learning with cross-lingual instruction following outperforms standard instruction tuning, whereas using internal contrastive learning alone degraded generation capabilities.

These results demonstrate that language models do not require costly, massive multilingual pretraining to achieve robust cross-lingual equity. Deploying targeted representation alignment after pretraining provides an efficient path to improve non-English capabilities, drastically lowering computational expenses and shortening deployment timelines. For organizations deploying global language solutions, the findings offer a practical mechanism to mitigate regional quality disparities and reduce risks associated with language bias in production systems.

Decision-makers should consider adopting this cross-lingual alignment framework as a standard post-processing step for generative models deployed across international markets. Practitioners should use English as a pivot language for broad multi-language alignment and configure contrastive learning at early transformer layers, while balancing cross-lingual and monolingual instruction prompts. For future work, technical teams should explore unsupervised alignment methods to expand coverage to dialects lacking parallel translation data, as well as investigate extending this framework to multimodal settings.

Confidence in these findings is supported by consistent empirical improvements across varied model families and standard benchmarks. However, several operational constraints apply. The framework currently relies on high-quality parallel labeled datasets, introducing risks of translation error propagation. Additionally, because the study evaluated models up to 7.5 billion parameters, additional validation is recommended when scaling to very large foundation models. Finally, using English as a central pivot language risks propagating English-centric cultural biases into target language outputs, which requires ongoing monitoring during deployment.

No sufficiently relevant recommendations were found.

Cover for Improving In-context Learning of Multilingual Generative Language Models with Cross-lingual Alignment

Abstract

Multilingual generative models obtain remarkable cross-lingual in-context learning capabilities through pre-training on large-scale corpora. However, they still exhibit a performance bias toward high-resource languages and learn isolated distributions of multilingual sentence representations, which may hinder knowledge transfer across languages. To bridge this gap, we propose a simple yet effective cross-lingual alignment framework exploiting pairs of translation sentences. It aligns the internal sentence representations across different languages via multilingual contrastive learning and aligns outputs by following cross-lingual instructions in the target language. Experimental results show that even with less than 0.1 ‰ of pre-training tokens, our alignment framework significantly boosts the cross-lingual abilities of generative language models and mitigates the performance gap. Further analyses reveal that it results in a better internal multilingual representation distribution of multilingual models.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Multilingual Generative Language Model
  • 2.2 Multilingual Instruction Tuning
  • 2.3 Contrastive Learning in Natural Langauge Processing
  • 3 Method
  • 3.1 Multilingual Contrastive Learning
  • 3.2 Cross-lingual Instruction Following
  • 4 Experiments
  • 4.1 Experiments Settings
  • 4.2 Bilingual Results and Analyses
  • 4.2.1 AFP Brings Better Bilingual Representations
  • 4.2.2 Multilingual Contrastive Learning on Bottom Layer Performs Better
  • 4.2.3 Cross-lingual Instruction Following or Multilingual Instruction Tuning?
  • 4.3 Multilingual Results and Analyses
  • 4.3.1 English as a pivot language or Pairwise Alignment?
  • 4.3.2 Combination with Other Cross-lingual Methods
  • 4.4 Extended to Alignment in 52 Languages
  • 4.5 Ablation Study
  • 5 Conclusion and Future Work
  • References
  • A Hyperparameters
  • B Additional Results
  • B.1 Pooling Methods
  • B.2 Weight of Cross-lingual Instruction Following
  • B.3 Distribution of multilingual representations
  • B.4 Performance on multilingual datasets
  • C Task Descriptions and Prompt Templates
  • D Additional Information about Language Code

Knowls

  1. Knowl 1 — AFP aligns internal representations and generated outputs

    model/method

    Align aFter Pre-training (AFP) fine-tunes a pretrained generative language model with two complementary objectives: Multilingual Contrastive Learning (MCL), which brings together internal representations of translation pairs, and Cross-lingual Instruction Following (CIF), which trains the model to answer in a target language even when its instruction is in another language. The combined objective is LAFP(θ)=LMCL(θ)+αLCIF(θ)L_{\mathrm{AFP}}(\theta)=L_{\mathrm{MCL}}(\theta)+\alpha L_{\mathrm{CIF}}(\theta), where θ\theta denotes model parameters and α≥0\alpha\geq 0 weights the instruction-following loss. The schematic on page 4 depicts these as internal-representation and output-alignment paths through the same generative model.

  2. Knowl 2 — MCL uses translation pairs as positives for decoder representations

    model/method

    For a parallel sentence pair (si,si+)(s_i,s_i^+), where sis_i is a source-language sentence and si+s_i^+ its translation, MCL extracts token representations from layer ll of a generative model f(⋅;θ)f(\cdot;\theta) and pools them into sentence vectors: hi=g(fl(si;θ))h_i=g(f_l(s_i;\theta)) and hi+=g(fl(si+;θ))h_i^+=g(f_l(s_i^+;\theta)). Here gg is a pooling operation, such as mean or max pooling. The contrastive objective is LMCL(θ)=E(si,si+)∼D[−log⁡exp⁡(sim(hi,hi+)/τ)∑jexp⁡(sim(hi,hj)/τ)]L_{\mathrm{MCL}}(\theta)=\mathbb{E}_{(s_i,s_i^+)\sim D}\left[-\log\frac{\exp(\mathrm{sim}(h_i,h_i^+)/\tau)}{\sum_j\exp(\mathrm{sim}(h_i,h_j)/\tau)}\right], where DD is the parallel training set, hjh_j is the representation of another sentence in the same mini-batch, sim\mathrm{sim} is cosine similarity, and τ\tau is a temperature. The translated sentence is the positive and other batch sentences provide negatives. The authors selected the first transformer layer after the embedding layer for alignment based on development-set performance.

  3. Knowl 3 — CIF trains responses in a language different from the prompt

    model/method

    Given an instruction context ciac_i^a and response riar_i^a in source language aa, CIF translates the response into a target language bb using a machine translation system ta→bt^{a\to b}, and appends a prompt pbp^b asking for an answer in language bb to the original context: cia→b=cia+pbc_i^{a\to b}=c_i^a+p^b and rib=ta→b(ria)r_i^b=t^{a\to b}(r_i^a). The model is trained by next-token prediction of the target-language response, with loss LCIF(θ)=E(cia,ria)[−∑jlog⁡P(rijb∣cia→b,ri,<jb;θ)]L_{\mathrm{CIF}}(\theta)=\mathbb{E}_{(c_i^a,r_i^a)}\left[-\sum_j\log P(r_{ij}^b\mid c_i^{a\to b},r_{i,<j}^b;\theta)\right], where rijbr_{ij}^b is response token jj and ri,<jbr_{i,<j}^b denotes preceding response tokens. With probability psrcp_{\mathrm{src}}, the target language is set equal to the source language; setting psrc=1p_{\mathrm{src}}=1 reduces CIF to same-language multilingual instruction tuning. The experiments found that an intermediate value, rather than either endpoint, worked best and used psrc=0.5p_{\mathrm{src}}=0.5.

  4. Knowl 4 — AFP uses modest parallel data and post-training fine-tuning

    experimental setup

    The experiments use Bactrian-X instruction data, translated from Alpaca into 52 languages and containing 67,000 samples per language, together with 100,000 selected OPUS-100 translation pairs. The English–Chinese setup uses 167,000 parallel samples in total, corresponding to about 20 million tokens—approximately 0.05‰ of BLOOM pre-training tokens. Models include XGLM (564M, 1.7B, and 7.5B parameters), BLOOM (560M, 1.7B, and 7.1B), and English-pretrained Llama 7B. Training uses AdamW with learning rate 10−510^{-5}, β1=0.9\beta_1=0.9, and β2=0.999\beta_2=0.999; MCL temperature τ=0.05\tau=0.05; batch size 128; and 10,000 steps. Mean or last-token pooling is selected using development performance. The CIF weight was tested over α∈{1,1.5,2}\alpha\in\{1,1.5,2\}; α=1.5\alpha=1.5 performed best for XGLM564M. Training used mixed precision and ZeRO on eight A100 80GB GPUs.

  5. Knowl 5 — Bilingual AFP improves in-context performance across model families

    empirical result

    On English–Chinese in-context evaluations covering XNLI and PAWS-X (understanding) and XCOPA, XStoryCloze, and XWinograd (reasoning), AFP improves the aggregate performance of the evaluated generative models using 167,000 parallel samples. The paper reports an average improvement of 3.31%, with 4.28% on the first two tasks and 2.67% on the reasoning tasks. For example, XGLM7.5B rises from an average score of 61.2 to 64.1, exceeding the reported 61.0 average for GPT-3 6.7B under the same evaluation template. BLOOM7.1B rises from 59.3 to 63.0, above BLOOMZ7.1B’s 61.1 despite BLOOMZ being fine-tuned on 78 million multilingual instructions. English-pretrained Llama7B also improves, from 62.9 to 65.5 average, including gains on Chinese.

  6. Knowl 6 — English-pivot alignment improves five-language results

    empirical result

    For alignment across English, Chinese, Thai, Turkish, and Swahili, the authors use English as a pivot: parallel pairs are drawn from English–other-language corpora, aligning each language to English. On the combined XNLI and XCOPA evaluation, AFP raises the reported average score from 46.1 to 50.7 for XGLM564M, 53.7 to 57.1 for XGLM7.5B, 44.9 to 48.5 for BLOOM560M, 47.2 to 50.6 for BLOOM1.7B, and 48.6 to 52.3 for BLOOM7.1B. For the alignment-policy comparison, English-pivot versus pairwise alignment gives average scores of 50.7 versus 50.3 for XGLM564M and 48.5 versus 48.0 for BLOOM560M. Pairwise alignment scores slightly higher on Swahili for these two models (48.4 versus 48.0 and 46.8 versus 46.6, respectively), although its overall average is lower.

  7. Knowl 7 — AFP extends to 52 languages and reduces cross-language performance spread

    empirical result

    Using English as the pivot, AFP is extended to all 52 languages in Bactrian-X. Across five multilingual in-context tasks, the reported average rises from 55.3 to 57.8 for XGLM7.5B and from 52.0 to 54.7 for BLOOM7.1B. The reported across-language dispersion also decreases: from ±8.5 to ±8.0 for XGLM7.5B and from ±8.2 to ±8.0 for BLOOM7.1B. The paper reports a 2.6% average improvement across the evaluated models and tasks, and a 2.8% improvement for BLOOM7.1B on languages unseen during its pre-training.

  8. Knowl 8 — AFP substantially raises zero-shot bilingual translation scores

    empirical result

    On the FLORES-101 development test set, AFP improves XGLM’s bilingual English–Chinese translation scores measured by COMET, especially in the zero-shot setting. Averaging the two translation directions and both model sizes, the reported zero-shot score increases from 27.3 to 61.2. For XGLM564M, the average rises from 26.0 to 59.3; for XGLM7.5B, it rises from 28.6 to 63.1. AFP also improves the reported scores in the one-, five-, and ten-shot settings.

  9. Knowl 9 — Representation alignment improves with AFP, but MCL alone hurts task performance

    empirical result

    The paper quantifies representation alignment as ℓalign=E(x,x+)∼Dpos∥f(x)−f(x+)∥2\ell_{\mathrm{align}}=\mathbb{E}_{(x,x^+)\sim D_{\mathrm{pos}}}\|f(x)-f(x^+)\|^2, where DposD_{\mathrm{pos}} contains translation pairs and ff is a sentence representation. It quantifies uniformity as ℓuniform=log⁡Ex,y∼i.i.d.Dexp⁡(−2∥f(x)−f(y)∥2)\ell_{\mathrm{uniform}}=\log\mathbb{E}_{x,y\overset{\mathrm{i.i.d.}}{\sim}D}\exp(-2\|f(x)-f(y)\|^2), where DD is the representation distribution; lower values indicate better alignment and uniformity. During the first 5,000 training steps, AFP reduces both metrics, whereas bilingual pre-training improves uniformity but not alignment. The t-SNE visualizations on pages 1, 5, and 16 likewise show language-separated sentence-representation clusters before AFP and more intermingled distributions after alignment. In an XGLM564M ablation averaged over five English–Chinese tasks, the 0-/3-/5-shot scores are 52.94/51.71/52.03 without extra training, 50.23/48.66/48.60 with MCL alone, 54.25/53.54/52.93 with same-language instruction tuning, 55.31/54.02/53.58 with CIF, and 55.97/55.50/56.15 with AFP. Thus, MCL alone worsens in-context task performance in this test, while combining it with CIF yields the best scores.

  10. Knowl 10 — AFP depends on parallel supervision and has scale and translation-error limits

    limitation

    AFP requires labeled parallel training data, so it cannot directly align a language for which suitable parallel samples are unavailable. Because of computational constraints, the experiments cover generative models with at most 7.5 billion parameters; performance at larger scales is not established by this work. CIF also depends on machine-translated responses, so translation errors can propagate into its training targets. The English-pivot policy may transfer English cultural biases or offensive response patterns into other languages.

Coverage note — Secondary pooling-method and loss-weight sweeps, the small additional gain from combining AFP with semantic-aligned demonstrations, and supplementary per-language benchmark scores are omitted because they do not add a core method or a primary result.

References

  1. 1.Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023. Palm 2 technical report. arXiv preprint arXiv:2305.10403.
  2. 2.Akari Asai, Sneha Kudugunta, Xinyan Velocity Yu, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, and Hannaneh Hajishirzi. 2023. Buffet: Benchmarking large language models for few-shot cross-lingual transfer. arXiv preprint arXiv:2305.14857.
  3. 3.Steven Cao, Nikita Kitaev, and Dan Klein. 2020. Multilingual alignment of contextual word representations. In International Conference on Learning Representations.
  4. 4.Zewen Chi, Li Dong, Furu Wei, Nan Yang, Saksham Singhal, Wenhui Wang, Xia Song, Xian-Ling Mao, Heyan Huang, and Ming Zhou. 2021. InfoXLM: An information-theoretic framework for cross-lingual language model pre-training. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3576–3588, Online. Association for Computational Linguistics.
  5. 5.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2475–2485, Brussels, Belgium. Association for Computational Linguistics.
  6. 6.Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023. Free dolly: Introducing the world’s first truly open instruction-tuned llm. databricks.
  7. 7.Yiming Cui, Ziqing Yang, and Xin Yao. 2023. Efficient and effective text encoding for chinese llama and alpaca. arXiv preprint arXiv:2304.08177.
  8. 8.Hongchao Fang, Sicheng Wang, Meng Zhou, Jiayuan Ding, and Pengtao Xie. 2020. Cert: Contrastive self-supervised learning for language understanding. arXiv preprint arXiv:2005.12766.
  9. 9.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. SimCSE: Simple contrastive learning of sentence embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6894–6910, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  10. 10.Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. 2022. The Flores-101 evaluation benchmark for low-resource and multilingual machine translation. Transactions of the Association for Computational Linguistics, 10:522–538.
  11. 11.Hao He, Qian Wang, Zhipeng Yu, Yang Zhao, Jiajun Zhang, and Chengqing Zong. 2021. Synchronous interactive decoding for multilingual neural machine translation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 12981–12988.
  12. 12.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  13. 13.Haonan Li, Fajri Koto, Minghao Wu, Alham Fikri Aji, and Timothy Baldwin. 2023. Bactrian-x: A multilingual replicable instruction-following model with low-rank adaptation. arXiv preprint arXiv:2305.15011.
  14. 14.Victor Weixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung, and James Y Zou. 2022. Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning. In Advances in Neural Information Processing Systems, volume 35, pages 17612–17625. Curran Associates, Inc.
  15. 15.Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, and Xian Li. 2022. Few-shot learning with multilingual generative language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 9019–9052, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  16. 16.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual Denoising Pre-training for Neural Machine Translation. Transactions of the Association for Computational Linguistics, 8:726–742.
  17. 17.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In International Conference on Learning Representations.
  18. 18.Jinliang Lu, Yu Lu, and Jiajun Zhang. 2023. Take a closer look at multilinguality! improve multilingual pre-training using monolingual corpora only. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 2891–2907, Singapore. Association for Computational Linguistics.
  19. 19.Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al. 2018. Mixed precision training. In International Conference on Learning Representations.
  20. 20.Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. 2023. Crosslingual generalization through multitask finetuning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15991–16111, Toronto, Canada. Association for Computational Linguistics.
  21. 21.Jianmo Ni, Gustavo Hernandez Abrego, Noah Constant, Ji Ma, Keith Hall, Daniel Cer, and Yinfei Yang. 2022. Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models. In Findings of the Association for Computational Linguistics: ACL 2022, pages 1864–1874, Dublin, Ireland. Association for Computational Linguistics.
  22. 22.OpenAI. 2022. Introducing chatgpt. OpenAI blog.
  23. 23.OpenAI. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  24. 24.Lin Pan, Chung-Wei Hang, Haode Qi, Abhishek Shah, Saloni Potdar, and Mo Yu. 2021a. Multilingual BERT post-pretraining alignment. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 210–219, Online. Association for Computational Linguistics.
  25. 25.Lin Pan, Chung-Wei Hang, Avirup Sil, and Saloni Potdar. 2022. Improved text classification via contrastive adversarial training. Proceedings of the AAAI Conference on Artificial Intelligence, 36(10):11130–11138.
  26. 26.Xiao Pan, Mingxuan Wang, Liwei Wu, and Lei Li. 2021b. Contrastive learning for many-to-many multilingual neural machine translation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 244–258, Online. Association for Computational Linguistics.
  27. 27.Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulic, and Anna Korhonen. 2020. XCOPA: A multilingual dataset for causal commonsense reasoning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2362–2376, Online. Association for Computational Linguistics.
  28. 28.Libo Qin, Qiguang Chen, Tianbao Xie, Qixin Li, Jian-Guang Lou, Wanxiang Che, and Min-Yen Kan. 2022. GL-CLeF: A global–local contrastive learning framework for cross-lingual spoken language understanding. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2677–2686, Dublin, Ireland. Association for Computational Linguistics.
  29. 29.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 8748–8763. PMLR.
  30. 30.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551.
  31. 31.Leonardo Ranaldi, Giulia Pucci, and Andre Freitas. 2023. Empowering cross-lingual abilities of instruction-tuned large language models by translation-following demonstrations. arXiv preprint arXiv:2308.14186.
  32. 32.Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD ’20, page 3505–3506, New York, NY, USA. Association for Computing Machinery.
  33. 33.Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. COMET: A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685–2702, Online. Association for Computational Linguistics.
  34. 34.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China. Association for Computational Linguistics.
  35. 35.Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon. 2011. Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In 2011 AAAI Spring Symposium Series.
  36. 36.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  37. 37.Tom Sherborne, Tom Hosking, and Mirella Lapata. 2023. Optimal transport posterior alignment for cross-lingual semantic parsing. Transactions of the Association for Computational Linguistics, 11:1432–1450.
  38. 38.Saleh Soltan, Shankar Ananthakrishnan, Jack Fitzgerald, Rahul Gupta, Wael Hamza, Haidar Khan, Charith Peris, Stephen Rawls, Andy Rosenbaum, Anna Rumshisky, et al. 2022. Alexatm 20b: Few-shot learning using a large-scale multilingual seq2seq model. arXiv preprint arXiv:2208.01448.
  39. 39.Alex Tamkin, Miles Brundage, Jack Clark, and Deep Ganguli. 2021. Understanding the capabilities, limitations, and societal impact of large language models. arXiv preprint arXiv:2102.02503.
  40. 40.Eshaan Tanwar, Subhabrata Dutta, Manish Borthakur, and Tanmoy Chakraborty. 2023. Multilingual LLMs are better cross-lingual in-context learners with alignment. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6292–6307, Toronto, Canada. Association for Computational Linguistics.
  41. 41.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  42. 42.Alexey Tikhonov and Max Ryabinin. 2021. It’s All in the Heads: Using Attention Heads as a Baseline for Cross-Lingual Transfer in Commonsense Reasoning. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3534–3546, Online. Association for Computational Linguistics.
  43. 43.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023a. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  44. 44.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023b. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  45. 45.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  46. 46.Liang Wang, Wei Zhao, and Jingming Liu. 2021. Aligning cross-lingual sentence representations with dual momentum contrast. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3807–3815, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  47. 47.Qian Wang, Jiajun Zhang, and Chengqing Zong. 2022. Synchronous inference for multilingual neural machine translation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:1827–1839.
  48. 48.Tongzhou Wang and Phillip Isola. 2020. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 9929–9939. PMLR.
  49. 49.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. Self-instruct: Aligning language models with self-generated instructions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13484–13508, Toronto, Canada. Association for Computational Linguistics.
  50. 50.Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le. 2022. Finetuned language models are zero-shot learners. In International Conference on Learning Representations.
  51. 51.Xiangpeng Wei, Haoran Wei, Huan Lin, Tianhao Li, Pei Zhang, Xingzhang Ren, Mei Li, Yu Wan, Zhiwei Cao, Binbin Xie, et al. 2023. PolyLM: An open source polyglot large language model. arXiv preprint arXiv:2307.06018.
  52. 52.Xiangpeng Wei, Rongxiang Weng, Yue Hu, Luxi Xing, Heng Yu, and Weihua Luo. 2021. On learning universal representations across languages. In International Conference on Learning Representations.
  53. 53.Zhuofeng Wu, Sinong Wang, Jiatao Gu, Madian Khabsa, Fei Sun, and Hao Ma. 2020. Clear: Contrastive learning for sentence representation. arXiv preprint arXiv:2012.15466.
  54. 54.Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. 2021. VideoCLIP: Contrastive pre-training for zero-shot video-text understanding. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6787–6800, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  55. 55.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online. Association for Computational Linguistics.
  56. 56.Nan Yang, Furu Wei, Binxing Jiao, Daxing Jiang, and Linjun Yang. 2021. xMoCo: Cross momentum contrastive learning for open-domain question answering. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6120–6129, Online. Association for Computational Linguistics.
  57. 57.Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. 2019. PAWS-X: A cross-lingual adversarial dataset for paraphrase identification. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3687–3692, Hong Kong, China. Association for Computational Linguistics.
  58. 58.Biao Zhang, Philip Williams, Ivan Titov, and Rico Sennrich. 2020. Improving massively multilingual neural machine translation and zero-shot translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1628–1639, Online. Association for Computational Linguistics.
  59. 59.Shaolei Zhang, Qingkai Fang, Zhuocheng Zhang, Zhengrui Ma, Yan Zhou, Langlin Huang, Mengyu Bu, Shangtong Gui, Yunji Chen, Xilin Chen, et al. 2023. Bayling: Bridging cross-lingual alignment and instruction following through interactive translation for large language models. arXiv preprint arXiv:2306.10968.
  60. 60.Yang Zhao, Jiajun Zhang, and Chengqing Zong. 2023. Transformer: A general framework from machine translation to others. Machine Intelligence Research, 20(4):514–538.
  61. 61.Wenhao Zhu, Yunzhe Lv, Qingxiu Dong, Fei Yuan, Jingjing Xu, Shujian Huang, Lingpeng Kong, Jiajun Chen, and Lei Li. 2023. Extrapolating large language models to non-english by aligning languages. arXiv preprint arXiv:2308.04948.

Citation

MLA
Li, C., et al. “Improving In-context Learning of Multilingual Generative Language Models with Cross-lingual Alignment”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 8058–76, https://doi.org/10.18653/v1/2024.naacl-long.445.
APA
Li, C., Wang, S., Zhang, J., & Zong, C. (2024). Improving In-context Learning of Multilingual Generative Language Models with Cross-lingual Alignment. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 8058–8076. https://doi.org/10.18653/v1/2024.naacl-long.445
Chicago
Li, C., S. Wang, J. Zhang, and C. Zong. 2024. “Improving In-context Learning of Multilingual Generative Language Models with Cross-lingual Alignment”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 8058–76. https://doi.org/10.18653/v1/2024.naacl-long.445.
Harvard
Li, C. et al. (2024) “Improving In-context Learning of Multilingual Generative Language Models with Cross-lingual Alignment”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 8058–8076. Available at: https://doi.org/10.18653/v1/2024.naacl-long.445.
Vancouver
1. Li C, Wang S, Zhang J, Zong C (2024) Improving In-context Learning of Multilingual Generative Language Models with Cross-lingual Alignment. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 8058–8076

BibTeX

@inproceedings{li-etal-2024-improving-context,
    title = "Improving In-context Learning of Multilingual Generative Language Models with Cross-lingual Alignment",
    author = "Li, Chong  and
      Wang, Shaonan  and
      Zhang, Jiajun  and
      Zong, Chengqing",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.445/",
    doi = "10.18653/v1/2024.naacl-long.445",
    pages = "8058--8076"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/