Multilingual Large Language Models Are Not (Yet) Code-Switchers
Ruochen ZhangSamuel CahyawijayaJan Christian Blaise CruzGenta Indra WinataAlham Fikri Aji
Reveals through systematic benchmarking across four distinct NLP tasks that prompted multilingual large language models consistently underperform substantially smaller fine-tuned models on code-switched text, demonstrating that standard multilingual pretraining fails to confer proficiency in mixed-language communication.
In multilingual communities globally, speakers frequently switch between two or more languages within a single conversation, a practice known as code-switching. Although multilingual large language models are increasingly used for natural language processing, their training focuses primarily on monolingual datasets. As a result, it remains unclear whether these models can effectively interpret and generate mixed-language speech and text, posing a risk of exclusion and degraded system performance for global user bases.
The article aims to evaluate the code-switching proficiency of existing multilingual large language models across multiple tasks and languages. It demonstrates whether general multilingual pretraining inherently equips these systems to process mixed-language inputs or if task-specific customization remains necessary.
To conduct this evaluation, the researchers benchmarked open-source and proprietary models—including BLOOMZ, mT0, XGLM, and ChatGPT—across four tasks: sentiment analysis, machine translation, summarization, and word-level language identification. The study covered several language pairs, such as Spanish-English, Hindi-English, Tamil-English, Malayalam-English, and Arabic dialects. The researchers compared the performance of large models using zero-shot and few-shot prompting techniques directly against much smaller, fine-tuned models such as XLM-RoBERTa, mBERT, and mBART.
The findings show that smaller, fine-tuned models consistently outperform large language models evaluated through prompting. In machine translation from Hindi-English to English, fine-tuned models scored between 25 and 32 BLEU points, while the largest prompted open-source model achieved under 20 BLEU. In word-level language identification, fine-tuned models achieved accuracy around 70 to 86 percent Macro F1, whereas open-source prompted models struggled significantly, rarely exceeding 20 percent due to formatting failures. While scaling up model size provided modest improvements, the positive effect of scale was noticeably weaker for mixed-language data than for monolingual data. Furthermore, adding few-shot examples often failed to help and sometimes degraded performance, as models frequently defaulted to monolingual English rather than generating code-switched text. ChatGPT performed competitively with fine-tuned models, but its proprietary nature prevents full analysis of its underlying data and mechanisms.
These results imply that multilingual large language models do not automatically understand or generate code-switched text simply because they were trained on individual constituent languages. Relying solely on general prompting for mixed-language applications introduces severe operational risks, including poor translation, failed compliance with structured outputs, and alienation of multilingual users. For cost-effective and accurate deployment, fine-tuning smaller architectures remains the superior approach.
The article recommends that organizations working with multilingual user bases continue utilizing targeted fine-tuning for code-switched applications. To improve foundation models long-term, developers should integrate mixed-language pairs into pretraining corpora, utilize data augmentation, and adopt code-switching optimization objectives during model alignment.
Readers should note that the study evaluated a limited set of language pairs and tasks due to the scarcity of high-quality, annotated code-switching data. While confidence in the benchmarked task results is high, evaluating emerging model architectures and a broader array of language combinations remains an area for further analysis.
- Paper: How Multilingual is Multilingual BERT?, Telmo Pires et al. (2019). Its evidence that multilingual BERT can transfer to code-switched Hindi-English provides an important earlier baseline for understanding this paper’s broader evaluation of code-switching in prompted language models.
- Paper: mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer, Linting Xue et al. (2020). This paper introduces mT5, one of the multilingual text-to-text foundations relevant to the source’s comparisons of prompted generative models with fine-tuned systems.
- Paper: Multilingual Denoising Pre-training for Neural Machine Translation, Yinhan Liu et al. (2020). Its introduction of mBART as a multilingual denoising-pretrained model helps explain the fine-tuned translation baseline used to assess code-switched inputs.
- Paper: Multitask Prompted Training Enables Zero-Shot Task Generalization, Victor Sanh et al. (2021). Its account of T0’s multitask prompt training clarifies the prompted-model approach whose limits the source tests on mixed-language tasks.
- Paper: MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks, Sanchit Ahuja et al. (2024). This later, much broader multilingual benchmark extends the source’s warning about uneven model performance by testing models across more languages, modalities, and task types.
- Paper: BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual Transfer, Akari Asai et al. (2024). This later benchmark directly develops the source’s comparison of prompting and fine-tuning into a systematic study of few-shot cross-lingual transfer across many languages and tasks.
- Paper: LLMs Are Few-Shot In-Context Low-Resource Language Learners, Samuel Cahyawijaya et al. (2024). This later study follows up on the source’s finding that prompting often fails by testing alignment strategies designed to make cross-lingual in-context learning more effective.
