Multilingual Large Language Models Are Not (Yet) Code-Switchers

Ruochen ZhangSamuel CahyawijayaJan Christian Blaise CruzGenta Indra WinataAlham Fikri Aji

article2023EMNLP103 citations

Reveals through systematic benchmarking across four distinct NLP tasks that prompted multilingual large language models consistently underperform substantially smaller fine-tuned models on code-switched text, demonstrating that standard multilingual pretraining fails to confer proficiency in mixed-language communication.

Listen

In multilingual communities globally, speakers frequently switch between two or more languages within a single conversation, a practice known as code-switching. Although multilingual large language models are increasingly used for natural language processing, their training focuses primarily on monolingual datasets. As a result, it remains unclear whether these models can effectively interpret and generate mixed-language speech and text, posing a risk of exclusion and degraded system performance for global user bases.

The article aims to evaluate the code-switching proficiency of existing multilingual large language models across multiple tasks and languages. It demonstrates whether general multilingual pretraining inherently equips these systems to process mixed-language inputs or if task-specific customization remains necessary.

To conduct this evaluation, the researchers benchmarked open-source and proprietary models—including BLOOMZ, mT0, XGLM, and ChatGPT—across four tasks: sentiment analysis, machine translation, summarization, and word-level language identification. The study covered several language pairs, such as Spanish-English, Hindi-English, Tamil-English, Malayalam-English, and Arabic dialects. The researchers compared the performance of large models using zero-shot and few-shot prompting techniques directly against much smaller, fine-tuned models such as XLM-RoBERTa, mBERT, and mBART.

The findings show that smaller, fine-tuned models consistently outperform large language models evaluated through prompting. In machine translation from Hindi-English to English, fine-tuned models scored between 25 and 32 BLEU points, while the largest prompted open-source model achieved under 20 BLEU. In word-level language identification, fine-tuned models achieved accuracy around 70 to 86 percent Macro F1, whereas open-source prompted models struggled significantly, rarely exceeding 20 percent due to formatting failures. While scaling up model size provided modest improvements, the positive effect of scale was noticeably weaker for mixed-language data than for monolingual data. Furthermore, adding few-shot examples often failed to help and sometimes degraded performance, as models frequently defaulted to monolingual English rather than generating code-switched text. ChatGPT performed competitively with fine-tuned models, but its proprietary nature prevents full analysis of its underlying data and mechanisms.

These results imply that multilingual large language models do not automatically understand or generate code-switched text simply because they were trained on individual constituent languages. Relying solely on general prompting for mixed-language applications introduces severe operational risks, including poor translation, failed compliance with structured outputs, and alienation of multilingual users. For cost-effective and accurate deployment, fine-tuning smaller architectures remains the superior approach.

The article recommends that organizations working with multilingual user bases continue utilizing targeted fine-tuning for code-switched applications. To improve foundation models long-term, developers should integrate mixed-language pairs into pretraining corpora, utilize data augmentation, and adopt code-switching optimization objectives during model alignment.

Readers should note that the study evaluated a limited set of language pairs and tasks due to the scarcity of high-quality, annotated code-switching data. While confidence in the benchmarked task results is high, evaluating emerging model architectures and a broader array of language combinations remains an area for further analysis.

arXiv: 2305.14235
Cover for Multilingual Large Language Models Are Not (Yet) Code-Switchers

Abstract

Multilingual Large Language Models (LLMs) have recently shown great capabilities in a wide range of tasks, exhibiting state-of-the-art performance through zero-shot or few-shot prompting methods. While there have been extensive studies on their abilities in monolingual tasks, the investigation of their potential in the context of code-switching (CSW), the practice of alternating languages within an utterance, remains relatively uncharted. In this paper, we provide a comprehensive empirical analysis of various multilingual LLMs, benchmarking their performance across four tasks: sentiment analysis, machine translation, summarization and word-level language identification. Our results indicate that despite multilingual LLMs exhibiting promising outcomes in certain tasks using zero or few-shot prompting, they still underperform in comparison to fine-tuned models of much smaller scales. We argue that current “multilingualism" in LLMs does not inherently imply proficiency with code-switching texts, calling for future research to bridge this discrepancy.

Table of Contents

  • 1 Introduction
  • 2 Experimental Setup
  • 2.1 Datasets
  • 2.2 Models
  • 3 Results and Discussion
  • 3.1 Sentiment Analysis Results
  • 3.2 Machine Translation Results
  • 3.3 Summarization Results
  • 3.4 Language Identification Results
  • 3.5 Few-Shot Results
  • 3.6 Benchmarking ChatGPT
  • 4 Implications for Future LLMs
  • 5 Related Work
  • 6 Conclusion
  • Limitations
  • Ethical Considerations
  • Acknowledgements
  • References
  • A Fine-tuning Model Setup
  • B Fine-tuning Model Results
  • C Prompt Templates
  • Sentiment Analysis
  • Machine Translation
  • Summarization
  • Word-level LID
  • D Detailed Results

Knowls

  1. Knowl 1 — Benchmark design compares prompting with task-specific fine-tuning

    experimental setup

    The study evaluates multilingual language models on code-switched sentiment analysis, machine translation, summarization, and word-level language identification. Prompted generative models include BLOOMZ (560M–7.1B parameters), mT0 (300M–13B), and XGLM (564M–7.5B); GPT-3.5 Turbo is also evaluated. Fine-tuned comparisons use XLM-R, mBERT, and mDeBERTa for sentiment analysis and language identification; mT0, M2M100, and mBART-50 for translation; and mT0 and standard mBART-50 for summarization. The prompting experiments use five prompt templates and 0-, 1-, 3-, and 5-shot settings; the primary reported results use zero-shot prompting for sentiment analysis, translation, and summarization, and five-shot prompting for language identification. GPT-3.5 Turbo uses one-shot prompting for language identification because of budget limits. Sentiment labels are selected by comparing the model’s next-token probabilities; the other tasks use text generation. Language-identification generations are required to tag each token in a specified bracket format and are parsed with a dynamic-programming method. Fine-tuning uses all available training examples. The reported maximum training schedules are 15 epochs for sentiment analysis, 3 for language identification, 10 for translation, and 3 for summarization; Adafactor is used, with learning rates of 2e-5 for sentiment analysis and 5e-5 for translation.

  2. Knowl 2 — The benchmark covers five code-switching language pairs and four task datasets

    data/table

    The evaluation covers Spanish–English, Malayalam–English, Tamil–English, Hinglish–English, English–Hindi, and Modern Standard Arabic–Egyptian Arabic, with four task categories. Sentiment analysis uses positive/negative examples from Sentimix and MixSentiment; examples with other labels are filtered out. Its train/validation/test sizes are 8,831/8,831/1,342 for Spanish–English, 2,541/275/703 for Malayalam–English, and 9,075/1,022/2,499 for Tamil–English. Hinglish–English translation uses MixMT with 8,060/942/960 sentence pairs. Hinglish-to-English summarization uses Gupshup with 5,831 training pairs and 500 each for validation and testing. Word-level language identification uses LinCE subsets: English–Hindi has 4,832 training and 744 validation examples, while MSA–Egyptian Arabic has 8,464 training and 1,116 validation examples. The language-identification test set is not publicly available, so results for that task are reported on validation data. These datasets define the language and task scope of the benchmark rather than a broad sample of all code-switching communities.

  3. Knowl 3 — Fine-tuned models generally outperform prompted public LLMs, while GPT-3.5 is competitive

    empirical result

    Across the tested datasets, the best fine-tuned models generally exceed prompted public multilingual LLMs, including much larger models. GPT-3.5 Turbo narrows this gap but does not surpass the best fine-tuned result in the reported comparisons. For sentiment analysis, the best fine-tuned Macro F1 scores are 79.94 on Malayalam–English, 78.81 on Spanish–English, and 68.28 on Tamil–English, all from 560M XLM-R; GPT-3.5 Turbo obtains 65.92, 75.64, and 63.15, respectively. For Hinglish-to-English translation, fine-tuned mT0 with prompting reaches 31.88 BLEU, compared with 19.28 for the best prompted public model and 27.64 for GPT-3.5 Turbo. For English-to-Hinglish, the corresponding scores are 13.90, 2.86, and 4.32 BLEU. For Hinglish-to-English summarization, fine-tuned mT0 reaches 40.12 ROUGE-L, versus 30.67 for the best prompted public model and 25.07 for GPT-3.5 Turbo. In language identification, 560M XLM-R reaches F1 scores of 86.65 on English–Hindi and 79.79 on MSA–Egyptian Arabic. GPT-3.5 Turbo obtains 80.19 and 71.41 with one in-context example, whereas the best public prompted-model results in the five-shot setting are 19.61 and 20.51, respectively. These comparisons use different prompting conditions for GPT-3.5 Turbo and public models on language identification.

  4. Knowl 4 — Scaling gains depend on task, architecture, and whether the input is code-switched

    empirical result

    Increasing model size often improves prompted performance, but the gains are neither uniform across tasks nor sufficient to close the fine-tuning gap. The paper’s regression slopes (β) and intercepts (α) for performance against model scaling are as follows. For machine translation, code-switched versus monolingual β/α are mT0 1.057/6.403 versus 1.626/6.075, and BLOOMZ 0.192/2.373 versus 0.824/6.240. For summarization they are mT0 0.712/5.471 versus 0.738/9.228, BLOOMZ 0.312/3.507 versus 0.644/8.637, and XGLM 0.029/0.444 versus 0.012/0.883. The translation slopes show a stronger scaling effect for monolingual Hindi-to-English than for Hinglish-to-English, particularly for BLOOMZ. Summarization slopes are similar for mT0 across input types, while the matched English-input summaries score consistently higher than summaries from Hinglish inputs. These results support task- and model-dependent scaling trends, not a universal scaling benefit for code-switching.

  5. Knowl 5 — Prompted sentiment analysis is unstable across model sizes

    empirical result

    On the Spanish–English, Malayalam–English, and Tamil–English sentiment datasets, fine-tuned models occupy the strongest performance range despite being much smaller than the largest prompted models. Scaling BLOOMZ and XGLM produces only small gains, while mT0 scores fluctuate around 50 Macro F1 rather than improving consistently with size. The majority-class baseline has an average F1 of 46 across the three datasets. The authors interpret mT0’s near-baseline and unstable results as evidence that it struggles to infer sentiment reliably from these code-switched texts; the observed size increases do not yield a dependable improvement.

  6. Knowl 6 — Models struggle to generate code-switched translation outputs

    empirical result

    Translation performance depends strongly on whether the code-switched text is the input or the required output. For Hinglish-to-English translation, encoder–decoder mT0 improves visibly with model size, whereas decoder-only BLOOMZ and XGLM show little improvement and remain weak overall. When the requested output is English-to-Hinglish, both prompted and fine-tuned systems perform substantially worse than on Hinglish-to-English. Prompted mT0 often returns the original English sentence instead of translating it into Hinglish, including in few-shot settings that provide Hinglish examples. This indicates that the models’ ability to translate from code-switched input does not imply an ability to produce code-switched target text.

  7. Knowl 7 — English-source summarization is consistently easier than Hinglish-source summarization

    empirical result

    On the Gupshup summarization task, fine-tuning outperforms zero-shot prompting, and mT0 is the strongest prompted model family with performance that generally rises with model size. The study also compares parallel evaluation examples with Hinglish versus English input, keeping the target summaries aligned. Models perform consistently better when summarizing the English input, even though the scaling impact is broadly similar for the two input conditions. The result isolates an input-language disadvantage for code-switched summarization rather than attributing the entire difference to a change in evaluation examples.

  8. Knowl 8 — Word-level language identification exposes instruction-format failures

    empirical result

    In zero-shot language identification, all tested prompted multilingual LLMs score zero on both the English–Hindi and MSA–Egyptian Arabic datasets. The models often fail to follow the required word-by-word output format, which represents each token and its language tag as a bracketed pair. Adding in-context examples improves public models somewhat, and their results tend to be best at five shots, but format errors remain common: generations may combine several words in one pair, omit words, or fail to reproduce the input tokens exactly. Fine-tuned token classifiers substantially outperform these prompted systems, showing that general multilingual instruction-following does not ensure reliable sequence tagging in this format.

  9. Knowl 9 — In-context examples affect code-switched tasks differently

    empirical result

    Increasing the number of demonstrations from zero to one tends to reduce prompted performance on sentiment analysis and summarization, while additional examples produce negligible changes in machine translation. For English-to-Hinglish translation, models may continue copying the English input rather than generating the requested mixed-language sentence, even after seeing Hinglish examples. Language identification behaves differently: additional demonstrations help models learn the unfamiliar token-tag output convention, with most models performing best at five shots. Even there, larger models frequently fail to preserve every token or adhere to the required format. Thus, few-shot prompting is not a general remedy for code-switching weaknesses; its effect depends on the task and its output requirements.

  10. Knowl 10 — The evidence is limited in language, task, and model coverage

    limitation

    The results cover only the language pairs and four tasks represented in the available datasets; the paper notes that scarcity and collection difficulty for high-quality code-switched data limit broader evaluation. Model sizes and the number of in-context examples are also constrained by local compute resources, and several newer model families are outside the evaluation. The findings therefore do not establish generality across all code-switching languages, tasks, or larger models. As possible future directions—not tested outcomes—the paper proposes improving code-switching representation in multilingual pretraining or instruction tuning, including pair-aware sampling, and adapting objectives such as token- or span-level language identification.

Coverage note — The broader normative discussion of inclusive language technology is omitted because it does not add a separate empirical or methodological finding; proposed technical directions are included with the scope limitations.

References

  1. 1.Muhammad Farid Adilazuarda, Samuel Cahyawijaya, Genta Indra Winata, Pascale Fung, and Ayu Purwarianti. 2022. Indorobusta: Towards robustness against diverse code-mixed indonesian local languages. In Proceedings of the First Workshop on Scaling Up Multilingual Evaluation, pages 25–34.
  2. 2.Gustavo Aguilar, Sudipta Kar, and Thamar Solorio. 2020. Lince: A centralized benchmark for linguistic code-switching evaluation. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 1803–1813.
  3. 3.Roee Aharoni, Melvin Johnson, and Orhan Firat. 2019. Massively multilingual neural machine translation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3874–3884, Minneapolis, Minnesota. Association for Computational Linguistics.
  4. 4.Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, and Pascale Fung. 2023. A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity.
  5. 5.John Baugh. 2005. Linguistic profiling. In Black linguistics, pages 167–180. Routledge.
  6. 6.Ondřej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, Radu Soricut, Lucia Specia, and Aleš Tamchyna. 2014. Findings of the 2014 workshop on statistical machine translation. In Proceedings of the Ninth Workshop on Statistical Machine Translation, pages 12–58, Baltimore, Maryland, USA. Association for Computational Linguistics.
  7. 7.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  8. 8.Samuel Cahyawijaya, Holy Lovenia, Tiezheng Yu, Willy Chung, and Pascale Fung. 2023. Instruct-align: Teaching novel languages with to llms through alignment-based cross-lingual instruction.
  9. 9.Samuel Cahyawijaya, Genta Indra Winata, Bryan Wilie, Karissa Vincentio, Xiaohong Li, Adhiguna Kuncoro, Sebastian Ruder, Zhi Yuan Lim, Syafri Bahar, Masayu Khodra, Ayu Purwarianti, and Pascale Fung. 2021. IndoNLG: Benchmark and resources for evaluating Indonesian natural language generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8875–8898, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  10. 10.Bharathi Raja Chakravarthi, Navya Jose, Shardul Suryawanshi, Elizabeth Sherly, and John P McCrae. 2020a. A sentiment analysis dataset for code-mixed malayalam-english. arXiv preprint arXiv:2006.00210.
  11. 11.Bharathi Raja Chakravarthi, Vigneshwaran Muralidaran, Ruba Priyadharshini, and John Philip McCrae. 2020b. Corpus creation for sentiment analysis in code-mixed Tamil-English text. In Proceedings of the 1st Joint Workshop on Spoken Language Technologies for Under-resourced languages (SLTU) and Collaboration and Computing for Under-Resourced Languages (CCURL), pages 202–210, Marseille, France. European Language Resources association.
  12. 12.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  13. 13.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Édouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451.
  14. 14.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186.
  15. 15.Kaustubh Dhole, Varun Gangal, Sebastian Gehrmann, Aadesh Gupta, Zhenhao Li, Saad Mahamood, Abinaya Mahadiran, Simon Mille, Ashish Shrivastava, Samson Tan, Tongshang Wu, Jascha Sohl-Dickstein, Jinho Choi, Eduard Hovy, Ondřej Dušek, Sebastian Ruder, Sajant Anand, Nagender Aneja, Rabin Banjade, Lisa Barthe, Hanna Behnke, Ian Berlot-Attwell, Connor Boyle, Caroline Brun, Marco Antonio Sobrevilla Cabezudo, Samuel Cahyawijaya, Emile Chapuis, Wanxiang Che, Mukund Choudhary, Christian Clauss, Pierre Colombo, Filip Cornell, Gautier Dagan, Mayukh Das, Tanay Dixit, Thomas Dopierre, Paul-Alexis Dray, Suchitra Dubey, Tatiana Ekeinhor, Marco Di Giovanni, Tanya Goyal, Rishabh Gupta, Louanes Hamla, Sang Han, Fabrice Harel-Canada, Antoine Honoré, Ishan Jindal, Przemysław Joniak, Denis Kleyko, Venelin Kovatchev, Kalpesh Krishna, Ashutosh Kumar, Stefan Langer, Seungjae Ryan Lee, Corey James Levinson, Hualou Liang, Kaizhao Liang, Zhexiong Liu, Andrey Lukyanenko, Vukosi Marivate, Gerard De Melo, Simon Meoni, Maxine Meyer, Afnan Mir, Nafise Sadat Moosavi, Niklas Meunnighoff, Timothy Sum Hon Mun, Kenton Murray, Marcin Namysl, Maria Obedkova, Priti Oli, Nivranshu Pasricha, Jan Pfister, Richard Plant, Vinay Prabhu, Vasile Pais, Libo Qin, Shahab Raji, Pawan Kumar Rajpoot, Vikas Raunak, Roy Rinberg, Nicholas Roberts, Juan Diego Rodriguez, Claude Roux, Vasconcellos Samus, Ananya Sai, Robin Schmidt, Thomas Scialom, Tshephisho Sefara, Saqib Shamsi, Xudong Shen, Yiwen Shi, Haoyue Shi, Anna Shvets, Nick Siegel, Damien Sileo, Jamie Simon, Chandan Singh, Roman Sitelew, Priyank Soni, Taylor Sorensen, William Soto, Aman Srivastava, Aditya Srivatsa, Tony Sun, Mukund Varma, A Tabassum, Fiona Tan, Ryan Teehan, Mo Tiwari, Marie Tolkiehn, Athena Wang, Zijian Wang, Zijie Wang, Gloria Wang, Fuxuan Wei, Bryan Wilie, Genta Indra Winata, Xinyu Wu, Witold Wydmanski, Tianbao Xie, Usama Yaseen, Michael Yee, Jing Zhang, and Yue Zhang. 2023. Nl-augmenter: A framework for task-sensitive natural language augmentatio. Northern European Journal of Language Technology, 9(1).
  16. 16.Mark Dingemanse and Andreas Liesenfeld. 2022. From text to talk: Harnessing conversational corpora for humane and diversity-aware language technology. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5614–5633, Dublin, Ireland. Association for Computational Linguistics.
  17. 17.A. Seza Dogruöz, Sunayana Sitaram, Barbara E. Bullock, and Almeida Jacqueline Toribio. 2021. A survey of code-switching: Linguistic and social perspectives for language technologies. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1654–1666, Online. Association for Computational Linguistics.
  18. 18.Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, Naman Goyal, Tom Birch, Vitaliy Liptchinsky, Sergey Edunov, Edouard Grave, Michael Auli, and Armand Joulin. 2020. Beyond english-centric multilingual machine translation.
  19. 19.Patrick Fernandes, Behrooz Ghorbani, Xavier Garcia, Markus Freitag, and Orhan Firat. 2023. Scaling laws for multilingual neural machine translation. arXiv preprint arXiv:2302.09650.
  20. 20.Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer. 2019. SAMSum corpus: A human-annotated dialogue dataset for abstractive summarization. In Proceedings of the 2nd Workshop on New Frontiers in Summarization, pages 70–79, Hong Kong, China. Association for Computational Linguistics.
  21. 21.Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2021. Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing.
  22. 22.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361.
  23. 23.Simran Khanuja, Sandipan Dandapat, Anirudh Srinivasan, Sunayana Sitaram, and Monojit Choudhury. 2020. Gluecos: An evaluation benchmark for code-switched nlp. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3575–3585.
  24. 24.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. In ICML 2022 Workshop on Knowledge Retrieval and Language Models.
  25. 25.Guillaume Lample and Alexis Conneau. 2019. Cross-lingual language model pretraining. Advances in Neural Information Processing Systems (NeurIPS).
  26. 26.Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro Von Werra, Chenghao Mou, Eduardo González Ponferrada, Huu Nguyen, et al. 2022. The bigscience roots corpus: A 1.6 tb composite multilingual dataset. Advances in Neural Information Processing Systems, 35:31809–31826.
  27. 27.Ke Li, Jinyu Li, Guoli Ye, Rui Zhao, and Yifan Gong. 2019. Towards code-switching asr for end-to-end ctc models. In ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 6076–6080.
  28. 28.Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, et al. 2021. Few-shot learning with multilingual language models. arXiv preprint arXiv:2112.10668.
  29. 29.Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2021. What makes good in-context examples for gpt-3? arXiv preprint arXiv:2101.06804.
  30. 30.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual denoising pretraining for neural machine translation. Transactions of the Association for Computational Linguistics, 8:726–742.
  31. 31.Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V Le, Barret Zoph, Jason Wei, et al. 2023. The flan collection: Designing data and methods for effective instruction tuning. arXiv preprint arXiv:2301.13688.
  32. 32.Laiba Mehnaz, Debanjan Mahata, Rakesh Gosangi, Uma Sushmitha Gunturi, Riya Jain, Gauri Gupta, Amardeep Kumar, Isabelle G. Lee, Anish Acharya, and Rajiv Ratn Shah. 2021. GupShup: Summarizing open-domain code-switched conversations. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6177–6192, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  33. 33.Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, et al. 2022. Crosslingual generalization through multitask finetuning. arXiv preprint arXiv:2211.01786.
  34. 34.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155.
  35. 35.Giovanni Paolini, Ben Athiwaratkun, Jason Krone, Jie Ma, Alessandro Achille, Rishita Anubhai, Cicero Nogueira dos Santos, Bing Xiang, and Stefano Soatto. 2021. Structured prediction as translation between augmented natural languages. arXiv preprint arXiv:2101.05779.
  36. 36.Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. The RefinedWeb dataset for Falcon LLM: outperforming curated corpora with web data, and web data only. arXiv preprint arXiv:2306.01116.
  37. 37.S. Poplack. 2001. Code switching: Linguistic. In International Encyclopedia of the Social & Behavioral Sciences, pages 2062–2065. Elsevier.
  38. 38.Shana Poplack. 1980. Sometimes i’ll start a sentence in spanish y termino en espanol: toward a typology of code-switching1.
  39. 39.Vineel Pratap, Andros Tjandra, Bowen Shi, Paden Tomasello, Arun Babu, Sayani Kundu, Ali Elkahky, Zhaoheng Ni, Apoorv Vyas, Maryam Fazel-Zarandi, et al. 2023. Scaling speech technology to 1,000+ languages. arXiv preprint arXiv:2305.13516.
  40. 40.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv e-prints.
  41. 41.Together Computer SambaNova Systems. 2023. BLOOMChat: a New Open Multilingual Chat LLM.
  42. 42.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Stella Biderman, Leo Gao, Tali Bers, Thomas Wolf, and Alexander M. Rush. 2021. Multitask prompted training enables zero-shot task generalization.
  43. 43.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  44. 44.Noam Shazeer and Mitchell Stern. 2018. Adafactor: Adaptive learning rates with sublinear memory cost. In International Conference on Machine Learning, pages 4596–4604. PMLR.
  45. 45.Vivek Srivastava and Mayank Singh. 2022. Overview and results of MixMT shared-task at WMT 2022. In Proceedings of the Seventh Conference on Machine Translation (WMT), pages 806–811, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computational Linguistics.
  46. 46.Samson Tan and Shafiq Joty. 2021. Code-mixing on sesame street: Dawn of the adversarial polyglots. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3596–3616.
  47. 47.Yuqing Tang, Chau Tran, Xian Li, Peng-Jen Chen, Naman Goyal, Vishrav Chaudhary, Jiatao Gu, and Angela Fan. 2020. Multilingual translation with extensible multilingual pretraining and finetuning. arXiv preprint arXiv:2008.00401.
  48. 48.Yuqing Tang, Chau Tran, Xian Li, Peng-Jen Chen, Naman Goyal, Vishrav Chaudhary, Jiatao Gu, and Angela Fan. 2021. Multilingual translation from denoising pre-training. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3450–3466, Online. Association for Computational Linguistics.
  49. 49.Eshaan Tanwar, Manish Borthakur, Subhabrata Dutta, and Tanmoy Chakraborty. 2023. Multilingual llms are better cross-lingual in-context learners with alignment. arXiv preprint arXiv:2305.05940.
  50. 50.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. 2022. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239.
  51. 51.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  52. 52.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2021. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652.
  53. 53.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems.
  54. 54.Genta Winata, Shijie Wu, Mayank Kulkarni, Thamar Solorio, and Daniel Preotiuc-Pietro. 2022a. Cross-lingual few-shot learning on unseen languages. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 777–791, Online only. Association for Computational Linguistics.
  55. 55.Genta Indra Winata, Alham Fikri Aji, Zheng-Xin Yong, and Thamar Solorio. 2022b. The decades progress on code-switching research in nlp: A systematic survey on trends and challenges. arXiv preprint arXiv:2212.09660.
  56. 56.Genta Indra Winata, Samuel Cahyawijaya, Zhaojiang Lin, Zihan Liu, Peng Xu, and Pascale Fung. 2020. Meta-transfer learning for code-switched speech recognition. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3770–3776, Online. Association for Computational Linguistics.
  57. 57.Genta Indra Winata, Samuel Cahyawijaya, Zihan Liu, Zhaojiang Lin, Andrea Madotto, and Pascale Fung. 2021a. Are multilingual models effective in code-switching? In Proceedings of the Fifth Workshop on Computational Approaches to Linguistic Code-Switching, pages 142–153.
  58. 58.Genta Indra Winata, Andrea Madotto, Zhaojiang Lin, Rosanne Liu, Jason Yosinski, and Pascale Fung. 2021b. Language models are few-shot multilingual learners. In Proceedings of the 1st Workshop on Multilingual Representation Learning, pages 1–15.
  59. 59.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online. Association for Computational Linguistics.
  60. 60.Zheng-Xin Yong, Ruochen Zhang, Jessica Zosa Forde, Skyler Wang, Samuel Cahyawijaya, Holy Lovenia, Genta Indra Winata, Lintang Sutawika, Jan Christian Blaise Cruz, Long Phan, et al. 2023. Prompting multilingual large language models to generate code-mixed texts: The case of south east asian languages. arXiv preprint arXiv:2303.13592.
  61. 61.Ruochen Zhang and Carsten Eickhoff. 2023. Crocosum: A benchmark dataset for cross-lingual code-switched summarization. arXiv preprint arXiv:2303.04092.

Citation

MLA
Zhang, R., et al. “Multilingual Large Language Models Are Not (Yet) Code-Switchers”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 12567–82, https://doi.org/10.18653/v1/2023.emnlp-main.774.
APA
Zhang, R., Cahyawijaya, S., Cruz, J. C. B., Winata, G. I., & Aji, A. F. (2023). Multilingual Large Language Models Are Not (Yet) Code-Switchers. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 12567–12582. https://doi.org/10.18653/v1/2023.emnlp-main.774
Chicago
Zhang, R., S. Cahyawijaya, J. C. B. Cruz, G. I. Winata, and A. F. Aji. 2023. “Multilingual Large Language Models Are Not (Yet) Code-Switchers”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 12567–82. https://doi.org/10.18653/v1/2023.emnlp-main.774.
Harvard
Zhang, R. et al. (2023) “Multilingual Large Language Models Are Not (Yet) Code-Switchers”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 12567–12582. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.774.
Vancouver
1. Zhang R, Cahyawijaya S, Cruz JCB, Winata GI, Aji AF (2023) Multilingual Large Language Models Are Not (Yet) Code-Switchers. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 12567–12582

BibTeX

@inproceedings{zhang-etal-2023-multilingual,
    title = "Multilingual Large Language Models Are Not (Yet) Code-Switchers",
    author = "Zhang, Ruochen  and
      Cahyawijaya, Samuel  and
      Cruz, Jan Christian Blaise  and
      Winata, Genta  and
      Aji, Alham Fikri",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.774/",
    doi = "10.18653/v1/2023.emnlp-main.774",
    pages = "12567--12582"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/