Multilingual LLMs are Better Cross-lingual In-context Learners with Alignment

Eshaan TanwarSubhabrata DuttaManish BorthakurTanmoy Chakraborty

article2023ACL103 citationsOutstanding Paper Award

Proposes X-InSTA, a prompt construction strategy that combines semantic coherence and task-based alignment between source and target languages to improve cross-lingual in-context learning across diverse low-resource text classification tasks.

Listen

Large language models can perform downstream tasks through in-context learning by providing a few labeled examples directly in the prompt without modifying the underlying model weights. While this method significantly lowers annotation costs and inference compute overhead, extending it across languages presents severe hurdles. Historically, cross-lingual prompting has relied on randomly selecting demonstrations from a high-resource language to classify target text in a low-resource language. This arbitrary sampling causes prompt incoherence and abrupt semantic shifts, leading to suboptimal classification accuracy.

The article aims to evaluate the weaknesses of standard cross-lingual prompting and introduces a novel framework called Cross-lingual In-context Source-Target Alignment to improve text classification across language boundaries. The proposed method demonstrates how structured alignment between source and target representations enables multilingual models to generalize more effectively across diverse languages.

To address prompt discordance, the approach introduces two complementary mechanisms. First, semantic alignment dynamically selects labeled demonstrations from the source language that are most similar in meaning to the target input using multilingual sentence embeddings. Second, task alignment inserts an explicit bridging statement into the prompt to define the target language and map source labels directly into target-language terms. The authors tested this framework across 44 cross-lingual language pairings spanning three standard benchmark tasks—multilingual product review classification, cross-lingual sentiment analysis, and hate-speech detection—primarily utilizing a 7.5-billion-parameter multilingual model.

The findings show that combining semantic and task alignment markedly boosts model accuracy. Across all evaluated tasks, the integrated framework outperformed standard random prompt selection with an average relative gain of approximately 18% in performance metrics, achieving improvements of up to 22% to 23% on review and sentiment benchmarks. Semantic alignment alone produced consistent gains across nearly all language pairs, while task alignment confirmed that informing the model about target label mappings is critical for accurate inference. However, exceptions emerged: target languages such as German in review tasks and English in hate-speech benchmarks exhibited near-random performance under certain configurations, and Mandarin performed better when label spaces were kept uniform rather than translated.

These results demonstrate that organizations can significantly enhance cross-lingual automated text processing without the expense of fine-tuning models or acquiring extensive labeled data in every target language. By structuring prompts with semantic relevance and explicit label definitions, teams can lower computing costs, decrease deployment timelines, and improve multi-market service capabilities. However, practitioners must note that alignment effectiveness varies across specific language pairs, and model performance remains vulnerable to cultural nuances, such as subtle sarcasm or localized hate speech cues.

Stakeholders looking to implement cross-lingual classification should adopt semantic and task alignment prompt templates over random sampling. Before broad operational rollout, organizations should conduct targeted validation on their specific source-target language pairs to identify languages that may require alternate label strategies, such as uniform label spaces. Future work should focus on automating dynamic bridge generation and integrating cultural context into prompt pipelines.

The analysis carries high confidence for the evaluated benchmark tasks and language pairs on mid-scale multilingual models. Nevertheless, confidence should remain cautious when scaling to larger, unverified commercial systems, excessively long input contexts exceeding prompt limits, or sensitive content moderation domains where cultural differences can cause misclassifications.

arXiv: 2305.05940EshaanT/X-InSTA

No sufficiently relevant recommendations were found.

Cover for Multilingual LLMs are Better Cross-lingual In-context Learners with Alignment

Abstract

In-context learning (ICL) unfolds as large language models become capable of inferring test labels conditioned on a few labeled samples without any gradient update. ICL-enabled large language models provide a promising step forward toward bypassing recurrent annotation costs in a low-resource setting. Yet, only a handful of past studies have explored ICL in a cross-lingual setting, in which the need for transferring label-knowledge from a high-resource language to a low-resource one is immensely crucial. To bridge the gap, we provide the first in-depth analysis of ICL for cross-lingual text classification. We find that the prevalent mode of selecting random input-label pairs to construct the prompt-context is severely limited in the case of cross-lingual ICL, primarily due to the lack of alignment in the input as well as the output spaces. To mitigate this, we propose a novel prompt construction strategy – Cross-lingual In-context Source-Target Alignment (X-InSTA). With an injected coherence in the semantics of the input examples and a task-based alignment across the source and target languages, X-InSTA is able to outperform random prompt selection by a large margin across three different tasks using 44 different cross-lingual pairs.

Table of Contents

  • 1 Introduction
  • 2 Prompting Techniques
  • 2.1 Preliminaries
  • 2.2 Semantic Alignment
  • 2.3 Task-based Alignment
  • 2.4 X-InSTA
  • 3 Results and Analysis
  • 3.1 Comparing Alignment Techniques
  • 3.2 Why does Task Alignment Work?
  • 3.3 Role of semantic alignment
  • 3.4 Automated aligner generation
  • 3.5 Error Analysis
  • 4 Related Works
  • 5 Conclusion
  • Limitations
  • Ethics statement
  • References
  • A Dataset Details
  • B Model Variants
  • C Hyperparameters
  • D Miscellaneous
  • D.1 Language Code
  • D.2 Prompt Examples
  • ACL 2023 Responsible NLP Checklist

Knowls

  1. Knowl 1 — X-InSTA combines semantic example retrieval with cross-language label guidance

    model/method

    Cross-lingual In-context Source-Target Alignment (X-InSTA) constructs a prompt for classifying an unlabeled target-language input using labeled examples from a source language. For each target input, it encodes the target input and source examples with a multilingual sentence encoder, ranks source examples by cosine similarity, and selects the top kk examples with their source-language labels. It then appends a manually written task aligner that states the target language and explains the mapping between source- and target-language labels. The multilingual LLM receives this concatenated context followed by the target input and predicts the most probable label from the target-language label set. The paper uses k=4k=4 in its experiments. The method is intended to make demonstrations semantically relevant while also explicitly communicating how the task labels transfer across languages.

  2. Knowl 2 — X-InSTA improves reported macro-F1 across three cross-lingual classification datasets

    empirical result

    The authors report relative macro-F1 improvements for X-InSTA over random prompting of 23% on multilingual Amazon Reviews (MARC), 22% on cross-language sentiment classification (CLS), and 14% on HatEval. The score vectors below give macro-F1 averaged over source languages for each target language; target-language order is stated for each dataset. On MARC, target order German, English, Spanish, French, Japanese, Mandarin: random prompting scores are [0.345, 0.633, 0.731, 0.557, 0.499, 0.462]; semantic alignment [0.375, 0.713, 0.757, 0.681, 0.610, 0.557]; task alignment [0.338, 0.722, 0.830, 0.758, 0.730, 0.335]; and X-InSTA [0.350, 0.795, 0.865, 0.822, 0.805, 0.337]. On CLS, target order German, English, French, Japanese: random [0.524, 0.602, 0.495, 0.631]; semantic [0.531, 0.621, 0.543, 0.697]; task [0.490, 0.686, 0.711, 0.776]; and X-InSTA [0.483, 0.715, 0.757, 0.803]. On HatEval, target order Spanish, English: random [0.435, 0.274]; semantic [0.493, 0.284]; task [0.499, 0.269]; and X-InSTA [0.542, 0.269]. Thus, X-InSTA is strongest on most target-language averages, but not uniformly: German and Mandarin on MARC and German on CLS remain weak, and English HatEval does not improve over random prompting.

  3. Knowl 3 — Semantic alignment retrieves source demonstrations similar to each target input

    model/method

    Semantic alignment selects labeled source-language demonstrations separately for each unlabeled target-language input. A multilingual sentence encoder produces embeddings for the target input and the source examples; cosine similarity between the target embedding and each source embedding determines the ranking. The method takes the top kk source examples, retains their labels, concatenates them into the prompt, and appends the target input for prediction in the target label set. In the paper’s experiments, k=4k=4. The procedure is based on the hypothesis that multilingual embeddings expose enough cross-language semantic similarity for relevant source demonstrations to help the LLM classify the target input.

  4. Knowl 4 — Task alignment states the target-language label mapping in the prompt

    model/method

    Task-based alignment adds a manually written task aligner to a prompt containing randomly selected labeled source-language examples. The aligner is written in the source language and identifies the target language and the target-language forms of the task labels—for example, stating that Spanish “malo” means “bad” and “bueno” means “good.” The source examples, their labels, and the aligner are concatenated before the target input; the LLM then predicts among the target-language labels. Unlike semantic alignment, this method does not retrieve demonstrations according to the individual target input. Its purpose is to explicitly connect the source and target label spaces.

  5. Knowl 5 — Cross-lingual in-context classification transfers from labeled source data to target inputs

    definition

    In the paper’s cross-lingual classification setup, a source language has a labeled dataset of input-label pairs, while a target language has inputs to classify. A prompt context is formed by concatenating selected source-language demonstrations, each consisting of an input and its label. Given a target-language test input and that context, the multilingual LLM selects the target-language label with the highest conditional probability. The source and target label sets are natural-language label expressions connected by a one-to-one translation mapping. The model is used without task-specific gradient updates.

  6. Knowl 6 — Evaluation covers three datasets, 44 language-pair setups, and a fixed few-shot protocol

    experimental setup

    Experiments evaluate macro-F1 on 44 cross-lingual setups across three classification datasets. MARC covers German, English, Spanish, French, Japanese, and Mandarin; each language has 200,000 training reviews available for demonstrations and a test set of 40,000 reviews. CLS covers German, English, French, and Japanese, with 2,000 training and 2,000 test sentences per language. HatEval covers English and Spanish hate-speech classification, with 10,000 English and 5,000 Spanish training posts and test sets of 3,000 English and 1,600 Spanish posts. The main experiments use XGLM-7.5B, selected after comparison with XGLM-1.7B and BLOOM-7.1B using random prompting. The reported settings are k=4k=4, maximum input length 1,024 tokens, batch size 4, an NVIDIA A100 GPU, and seeds 32, 5, 232, 100, and 42.

  7. Knowl 7 — Label-space information explains much of the task-aligner effect

    empirical result

    On MARC, the authors compare random prompting with four changes to task alignment: making the target label space identical to the source label space, providing only target-language information, using an aligner for an unrelated third language, or using the correct task aligner. The following are average macro-F1 scores across source-target pairs, in target order German, English, Spanish, French, Japanese, Mandarin. Random prompting: [0.345, 0.633, 0.731, 0.557, 0.499, 0.462]. Uniform label space: [0.441, 0.570, 0.493, 0.414, 0.483, 0.594]. Language information only: [0.346, 0.645, 0.733, 0.575, 0.543, 0.508]. Third-language aligner: [0.345, 0.687, 0.755, 0.673, 0.601, 0.423]. Incorrect task aligner: [0.338, 0.665, 0.787, 0.647, 0.544, 0.339]. Correct task alignment: [0.338, 0.722, 0.830, 0.758, 0.730, 0.335]. Correct task alignment is best for English, Spanish, French, and Japanese; uniform label space is best for German and Mandarin. The authors interpret the generally small gain from language-only information, compared with label-mapping prompts, as evidence that label information is particularly important. They also suggest the German and Mandarin results reflect difficulty aligning those target label spaces.

  8. Knowl 8 — Dissimilar demonstrations generally underperform semantic retrieval

    empirical result

    The paper tests semantic alignment on CLS against a variant that selects the source sentences least similar to the target input. Scores are macro-F1, averaged over target languages for each source language; source-language order is German, English, French, Japanese. Random prompting gives [0.524, 0.602, 0.495, 0.631], non-semantic selection gives [0.531, 0.561, 0.453, 0.515], and semantic selection gives [0.531, 0.621, 0.543, 0.697]. The non-semantic variant falls below random prompting for English, French, and Japanese, while German is an exception. The authors report an average 8% fall for non-semantic selection relative to random prompting and a 10% gain for semantic alignment.

  9. Knowl 9 — An mT5-generated bridge improves on random prompts but does not replace task alignment

    empirical result

    The authors test an automated alternative to manually writing task aligners. For each prompt, they concatenate randomly selected source-language demonstrations, a mask token, and the target input; mT5 generates a span in the masked position, which is appended to the context before the classification LLM predicts a target-language label. This experiment uses English as the source language. The macro-F1 vectors, in target order MARC German, Spanish, French, Japanese, Mandarin; CLS German, French, Japanese; HatEval Spanish, are: random prompting [0.380, 0.761, 0.663, 0.526, 0.362, 0.682, 0.412, 0.609, 0.435]; semantic alignment [0.458, 0.783, 0.762, 0.608, 0.450, 0.677, 0.505, 0.691, 0.493]; task-based alignment [0.355, 0.888, 0.826, 0.727, 0.333, 0.620, 0.696, 0.752, 0.499]; automated aligner [0.531, 0.792, 0.699, 0.599, 0.350, 0.721, 0.430, 0.610, 0.438]. The generated aligner beats random prompting on several targets and is competitive with semantic alignment in some cases, but does not consistently match task-based alignment. The authors propose that the generated bridge lacks task-specific signals and may be suboptimal because mT5 and XGLM have different pretraining distributions.

  10. Knowl 10 — Static aligners and LLM reasoning limits constrain X-InSTA

    limitation

    X-InSTA relies on manually written, static task aligners, requiring task-specific human intervention and limiting the prompt’s ability to adapt to nuances in an individual example. The paper’s error analysis identifies misleading semantic matches when the same slur is used differently, missing cultural knowledge (including context about migration), long inputs that exceed the 1,024-token limit, and failures involving sarcasm despite an apparently good prompt. The authors attribute most observed errors to limitations of the LLM itself and note that X-InSTA does not resolve broader in-context-learning weaknesses such as limited commonsense reasoning or factual inconsistency. Resource constraints also prevented validation on larger or commercially available LLMs, so the reported benefits are not established for those models.

Coverage note — The full source-by-target score matrices and individual prompt examples are omitted; the main performance comparisons are summarized with source-averaged scores, and the recurring error patterns are included without reproducing each example.

References

  1. 1.Sweta Agrawal, Chunting Zhou, Mike Lewis, Luke Zettlemoyer, and Marjan Ghazvininejad. 2022. In-context examples selection for machine translation.
  2. 2.Valerio Basile, Cristina Bosco, Elisabetta Fersini, Debora Nozza, Viviana Patti, Francisco Manuel Rangel Pardo, Paolo Rosso, and Manuela Sanguinetti. 2019. SemEval-2019 task 5: Multilingual detection of hate speech against immigrants and women in Twitter. In Proceedings of the 13th International Workshop on Semantic Evaluation, pages 54–63, Minneapolis, Minnesota, USA. Association for Computational Linguistics.
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  4. 4.Tyler A Chang, Zhuowen Tu, and Benjamin K Bergen. 2022. The geometry of multilingual language model representations. arXiv preprint arXiv:2205.10964.
  5. 5.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  6. 6.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding.
  8. 8.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding.
  9. 9.Phillip Keung, Yichao Lu, György Szarvas, and Noah A. Smith. 2020. The multilingual Amazon reviews corpus. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4563–4568, Online. Association for Computational Linguistics.
  10. 10.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  11. 11.Yingcong Li, M Emrullah Ildiz, Dimitris Papailiopoulos, and Samet Oymak. 2023. Transformers as algorithms: Generalization and implicit model selection in in-context learning. arXiv preprint arXiv:2301.07067.
  12. 12.Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, et al. 2021. Few-shot learning with multilingual language models. arXiv preprint arXiv:2112.10668.
  13. 13.Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2022. What makes good in-context examples for GPT-3? In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pages 100–114, Dublin, Ireland and Online. Association for Computational Linguistics.
  14. 14.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual denoising pre-training for neural machine translation.
  15. 15.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach.
  16. 16.Farhad Nooralahzadeh, Giannis Bekoulis, Johannes Bjerva, and Isabelle Augenstein. 2020. Zero-shot cross-lingual transfer with meta learning. arXiv preprint arXiv:2003.02739.
  17. 17.Marinela Parovic, Goran Glavaš, Ivan Vulić, and Anna Korhonen. 2022. BAD-X: Bilingual adapters improve zero-shot cross-lingual transfer. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1791–1799, Seattle, United States. Association for Computational Linguistics.
  18. 18.Peter Prettenhofer and Benno Stein. 2010. Cross-language text classification using structural correspondence learning. In Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics, pages 1118–1127, Uppsala, Sweden. Association for Computational Linguistics.
  19. 19.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer.
  20. 20.Nils Reimers and Iryna Gurevych. 2020. Making monolingual sentence embeddings multilingual using knowledge distillation. arXiv preprint arXiv:2004.09813.
  21. 21.Peng Shi, Rui Zhang, He Bai, and Jimmy Lin. 2022. Xricl: Cross-lingual retrieval-augmented in-context learning for cross-lingual text-to-sql semantic parsing. arXiv preprint arXiv:2210.13693.
  22. 22.Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov. 2022. Transformers learn in-context by gradient descent. arXiv preprint arXiv:2212.07677.
  23. 23.Albert Webson and Ellie Pavlick. 2022. Do prompt-based models really understand the meaning of their prompts? In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2300–2344, Seattle, United States. Association for Computational Linguistics.
  24. 24.Genta Indra Winata, Andrea Madotto, Zhaojiang Lin, Rosanne Liu, Jason Yosinski, and Pascale Fung. 2021. Language models are few-shot multilingual learners. In Proceedings of the 1st Workshop on Multilingual Representation Learning, pages 1–15, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  25. 25.Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. 2022. An explanation of in-context learning as implicit bayesian inference. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  26. 26.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2020. mt5: A massively multilingual pre-trained text-to-text transformer.
  27. 27.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online. Association for Computational Linguistics.
  28. 28.Ningyu Zhang, Luoqiu Li, Xiang Chen, Shumin Deng, Zhen Bi, Chuanqi Tan, Fei Huang, and Huajun Chen. 2021. Differentiable prompt makes pre-trained language models better few-shot learners. arXiv preprint arXiv:2108.13161.
  29. 29.Tony Z. Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models.

Citation

MLA
Tanwar, E., et al. “Multilingual LLMs Are Better Cross-lingual In-context Learners with Alignment”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 6292–307, https://doi.org/10.18653/v1/2023.acl-long.346.
APA
Tanwar, E., Dutta, S., Borthakur, M., & Chakraborty, T. (2023). Multilingual LLMs are Better Cross-lingual In-context Learners with Alignment. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6292–6307. https://doi.org/10.18653/v1/2023.acl-long.346
Chicago
Tanwar, E., S. Dutta, M. Borthakur, and T. Chakraborty. 2023. “Multilingual LLMs Are Better Cross-lingual In-context Learners with Alignment”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6292–6307. https://doi.org/10.18653/v1/2023.acl-long.346.
Harvard
Tanwar, E. et al. (2023) “Multilingual LLMs are Better Cross-lingual In-context Learners with Alignment”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 6292–6307. Available at: https://doi.org/10.18653/v1/2023.acl-long.346.
Vancouver
1. Tanwar E, Dutta S, Borthakur M, Chakraborty T (2023) Multilingual LLMs are Better Cross-lingual In-context Learners with Alignment. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 6292–6307

BibTeX

@inproceedings{tanwar-etal-2023-multilingual,
    title = "Multilingual {LLM}s are Better Cross-lingual In-context Learners with Alignment",
    author = "Tanwar, Eshaan  and
      Dutta, Subhabrata  and
      Borthakur, Manish  and
      Chakraborty, Tanmoy",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.346/",
    doi = "10.18653/v1/2023.acl-long.346",
    pages = "6292--6307"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/