Improving In-context Learning of Multilingual Generative Language Models with Cross-lingual Alignment
Chong LiShaonan WangJiajun ZhangChengqing Zong
Proposes an efficient post-training alignment framework combining multilingual contrastive learning and cross-lingual instruction following to bridge the performance gap between high- and low-resource languages using fewer than one million parallel samples.
Modern multilingual artificial intelligence models exhibit notable performance disparities across languages, frequently favoring high-resource languages such as English over less-represented ones. This imbalance stems from uneven training data and internally isolated language representations, which impede the model's ability to transfer learned knowledge across languages. Given the prohibitive computational cost of retraining massive foundation models from scratch, there is an urgent operational need for lightweight, post-pretraining alignment techniques that bridge this cross-lingual gap.
The article demonstrates and evaluates Align aFter Pre-training (AFP), a framework designed to enhance the cross-lingual in-context learning capabilities of generative language models using limited parallel translation data. The core objective is to align internal sentence representations and generation outputs across languages, thereby reducing performance disparities and enabling efficient cross-lingual knowledge transfer.
To achieve this, the authors implemented a two-part approach evaluated across multiple model architectures (XGLM, BLOOM, and Llama) spanning sizes from 560 million to 7.5 billion parameters. The first component, Multilingual Contrastive Learning, pulls the internal representations of translation pairs closer together within the model's early transformer layers. The second component, Cross-lingual Instruction Following, trains the model to respond in a specified target language given a prompt in a source language. The experiments utilized fewer than one million parallel translation samples—totaling approximately 20 million tokens, or less than 0.05 per mille of standard pretraining volume—and tested performance across standardized understanding, reasoning, and translation benchmarks covering up to 52 languages.
The findings show consistent and significant performance gains across all evaluated benchmarks. In bilingual English-Chinese testing, the framework improved overall accuracy by an average of 3.31%, with understanding tasks improving by 4.28% and reasoning by 2.67%. When scaled across 52 languages using English as a pivot language, the method yielded an average performance boost of 2.6% while narrowing performance variance across high-, medium-, and low-resource languages. Furthermore, the framework notably improved zero-shot translation quality and generalized effectively to unseen languages; for example, BLOOM's performance on non-pretrained languages such as Thai and Turkish improved by nearly 4%. An ablation analysis confirmed that combining internal contrastive learning with cross-lingual instruction following outperforms standard instruction tuning, whereas using internal contrastive learning alone degraded generation capabilities.
These results demonstrate that language models do not require costly, massive multilingual pretraining to achieve robust cross-lingual equity. Deploying targeted representation alignment after pretraining provides an efficient path to improve non-English capabilities, drastically lowering computational expenses and shortening deployment timelines. For organizations deploying global language solutions, the findings offer a practical mechanism to mitigate regional quality disparities and reduce risks associated with language bias in production systems.
Decision-makers should consider adopting this cross-lingual alignment framework as a standard post-processing step for generative models deployed across international markets. Practitioners should use English as a pivot language for broad multi-language alignment and configure contrastive learning at early transformer layers, while balancing cross-lingual and monolingual instruction prompts. For future work, technical teams should explore unsupervised alignment methods to expand coverage to dialects lacking parallel translation data, as well as investigate extending this framework to multimodal settings.
Confidence in these findings is supported by consistent empirical improvements across varied model families and standard benchmarks. However, several operational constraints apply. The framework currently relies on high-quality parallel labeled datasets, introducing risks of translation error propagation. Additionally, because the study evaluated models up to 7.5 billion parameters, additional validation is recommended when scaling to very large foundation models. Finally, using English as a central pivot language risks propagating English-centric cultural biases into target language outputs, which requires ongoing monitoring during deployment.
- Paper: Multilingual LLMs are Better Cross-lingual In-context Learners with Alignment, Eshaan Tanwar et al. (2023). This earlier framework establishes how semantic and task alignment can improve cross-lingual in-context learning, directly preparing readers for the source’s alignment-based approach.
No sufficiently relevant recommendations were found.
