Multilingual LLMs are Better Cross-lingual In-context Learners with Alignment
Eshaan TanwarSubhabrata DuttaManish BorthakurTanmoy Chakraborty
Proposes X-InSTA, a prompt construction strategy that combines semantic coherence and task-based alignment between source and target languages to improve cross-lingual in-context learning across diverse low-resource text classification tasks.
Large language models can perform downstream tasks through in-context learning by providing a few labeled examples directly in the prompt without modifying the underlying model weights. While this method significantly lowers annotation costs and inference compute overhead, extending it across languages presents severe hurdles. Historically, cross-lingual prompting has relied on randomly selecting demonstrations from a high-resource language to classify target text in a low-resource language. This arbitrary sampling causes prompt incoherence and abrupt semantic shifts, leading to suboptimal classification accuracy.
The article aims to evaluate the weaknesses of standard cross-lingual prompting and introduces a novel framework called Cross-lingual In-context Source-Target Alignment to improve text classification across language boundaries. The proposed method demonstrates how structured alignment between source and target representations enables multilingual models to generalize more effectively across diverse languages.
To address prompt discordance, the approach introduces two complementary mechanisms. First, semantic alignment dynamically selects labeled demonstrations from the source language that are most similar in meaning to the target input using multilingual sentence embeddings. Second, task alignment inserts an explicit bridging statement into the prompt to define the target language and map source labels directly into target-language terms. The authors tested this framework across 44 cross-lingual language pairings spanning three standard benchmark tasks—multilingual product review classification, cross-lingual sentiment analysis, and hate-speech detection—primarily utilizing a 7.5-billion-parameter multilingual model.
The findings show that combining semantic and task alignment markedly boosts model accuracy. Across all evaluated tasks, the integrated framework outperformed standard random prompt selection with an average relative gain of approximately 18% in performance metrics, achieving improvements of up to 22% to 23% on review and sentiment benchmarks. Semantic alignment alone produced consistent gains across nearly all language pairs, while task alignment confirmed that informing the model about target label mappings is critical for accurate inference. However, exceptions emerged: target languages such as German in review tasks and English in hate-speech benchmarks exhibited near-random performance under certain configurations, and Mandarin performed better when label spaces were kept uniform rather than translated.
These results demonstrate that organizations can significantly enhance cross-lingual automated text processing without the expense of fine-tuning models or acquiring extensive labeled data in every target language. By structuring prompts with semantic relevance and explicit label definitions, teams can lower computing costs, decrease deployment timelines, and improve multi-market service capabilities. However, practitioners must note that alignment effectiveness varies across specific language pairs, and model performance remains vulnerable to cultural nuances, such as subtle sarcasm or localized hate speech cues.
Stakeholders looking to implement cross-lingual classification should adopt semantic and task alignment prompt templates over random sampling. Before broad operational rollout, organizations should conduct targeted validation on their specific source-target language pairs to identify languages that may require alternate label strategies, such as uniform label spaces. Future work should focus on automating dynamic bridge generation and integrating cultural context into prompt pipelines.
The analysis carries high confidence for the evaluated benchmark tasks and language pairs on mid-scale multilingual models. Nevertheless, confidence should remain cautious when scaling to larger, unverified commercial systems, excessively long input contexts exceeding prompt limits, or sensitive content moderation domains where cultural differences can cause misclassifications.
- Paper: Cross-lingual Language Model Pretraining, Guillaume Lample et al. (2019). Its multilingual pretraining methods establish how cross-lingual representations arise, the foundation this paper uses to align source and target examples.
- Paper: Making Pre-trained Language Models Better Few-shot Learners, Tianyu Gao et al. (2021). Its semantic selection of few-shot demonstrations introduces the retrieval approach this paper adapts to cross-lingual prompts.
- Paper: Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering, Zhiyong Wu et al. (2023). Its query-specific semantic retrieval framework provides a direct precursor to this paper’s alignment-based selection of demonstrations.
No sufficiently relevant recommendations were found.
