Memory-efficient NLLB-200: Language-specific Expert Pruning of a Massively Multilingual Machine Translation Model

Yeskendir KoishekenovAlexandre BerardVassilina Nikoulina

article2023ACL63 citations

Proposes an inference-time pruning method that removes up to 80% of experts from the 54.5-billion-parameter NLLB-200 translation model without fine-tuning, reducing hardware requirements from multiple devices down to a single 32GB GPU while maintaining translation quality.

Listen

State-of-the-art machine translation across hundreds of languages increasingly relies on massive artificial intelligence models. A leading example is NLLB-200, a translation model covering 202 languages whose largest variant contains 54.5 billion parameters structured as a Mixture of Experts—an architecture that divides model capacity into specialized sub-networks called experts and activates only a subset for each piece of text. While this design delivers industry-leading translation quality, running the full model requires specialized hardware with at least four 32-gigabyte graphics processing units (GPUs), creating high computational costs and severe deployment bottlenecks for production environments.

The article investigates whether multilingual models can be compressed at inference time without requiring computationally expensive retraining or fine-tuning. Specifically, the objective is to evaluate whether specific experts naturally specialize in individual languages, and to demonstrate a pruning strategy that eliminates redundant experts so the model can operate on a single 32-gigabyte GPU while maintaining translation quality.

The researchers designed importance metrics based on how frequently and confidently the model selects each expert when translating validation text. They tested various pruning algorithms and retention rates across 53 representative languages spanning high-, low-, and very low-resource categories, validating their final setup across all 202 languages and 40,602 translation directions using the standard FLORES-200 benchmark dataset.

The key findings reveal that up to 80% of all experts can be safely pruned without fine-tuning while incurring a negligible average translation quality drop of approximately 0.28 points out of 100 on standard translation metrics. The analysis demonstrates that an unbalanced pruning ratio keeping three times as many experts in the encoder (input processing) as in the decoder (output generation) delivers the best performance. Furthermore, the experiments prove that language-specific specialization genuinely emerges in the model: decoder experts show strong specialization for the target language (sharing 68% to 87% similarity for the same language) and naturally cluster along linguistic family lines, while encoder experts remain largely independent of the target language.

These findings have direct operational and economic implications. Organizations can reduce the hardware footprint required to run top-tier multilingual translation by 75%, cutting infrastructure demands from four GPUs to a single 32-gigabyte GPU. When four GPUs remain available, the pruned model doubles the translation throughput. This eliminates major cost barriers and enables efficient localized translation services across both major and very low-resource languages without degrading quality below dense baseline models.

For deployment, the article recommends adopting a language-specific pruning configuration using fixed allocations per layer with a 3:1 encoder-to-decoder expert ratio. Practitioners should select encoder experts based on the source language and decoder experts based on the target language, utilizing the publicly released expert statistics to avoid recomputing routing metrics. Further development should focus on addressing a minor post-pruning tendency toward sentence repetition and slight hallucination by identifying and preserving experts specialized in signaling the end of generated sequences.

Decision-makers should note certain limitations: the conclusions are drawn exclusively from the NLLB-200 architecture and its particular training objective, so pruning patterns may differ for other models. While the evaluation across all 202 languages provides high confidence in general quality preservation, teams deploying pruned models for sensitive or automated downstream tasks should monitor output length ratios and translation biases.

arXiv: 2212.09811naver/nllb-pruning
Cover for Memory-efficient NLLB-200: Language-specific Expert Pruning of a Massively Multilingual Machine Translation Model

Abstract

The recently released NLLB-200 is a set of multilingual Neural Machine Translation models that cover 202 languages. The largest model is based on a Mixture of Experts architecture and achieves SoTA results across many language pairs. It contains 54.5B parameters and requires at least four 32GB GPUs just for inference. In this work, we propose a pruning method that enables the removal of up to 80% of experts without further finetuning and with a negligible loss in translation quality, which makes it feasible to run the model on a single 32GB GPU. Further analysis suggests that our pruning metrics can identify language-specific experts.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Background
  • 3.1 Mixture-of-Experts models
  • 3.2 NLLB-200
  • 4 Our Approach
  • 4.1 Expert pruning metrics
  • 4.2 Expert statistics granularity
  • 4.3 Expert pruning algorithm
  • 5 Experiments
  • 5.1 Evaluation settings
  • 5.2 Results
  • 6 Discussion
  • 6.1 Inference speed and compute budget
  • 6.2 Similarity of selected experts
  • 6.3 Similarity of languages based on the importance metric
  • 6.4 Discrepancy between chrF++ and spBLEU scores
  • 7 Conclusion
  • 8 Risks and Limitations
  • Acknowledgement
  • References
  • A Discrepancy between chrF++ and spBLEU scores

Knowls

  1. Knowl 1 — Language-specific fixed-budget pruning enables single-GPU NLLB-200 inference

    model/method

    The 54.5B-parameter NLLB-200 Mixture-of-Experts translation model can be pruned at inference without fine-tuning by selecting experts from gate statistics specific to a source or target language. For each encoder expert layer, rank experts by importance for the source language and retain the top 36; for each decoder expert layer, rank experts by importance for the target language and retain the top 12. NLLB-200 has six expert-bearing layers on each side, so this keeps 216 encoder experts and 72 decoder experts (288 total, reported as 80% pruning). Remove the other experts and use the resulting model for inference. The per-language configuration allows the pruned model to run on one 32GB GPU.

  2. Knowl 2 — Gate-based importance metric for ranking experts

    equation

    For an expert ee, let top1⁡(e)\operatorname{top1}(e) be the fraction of input tokens for which the gating network ranks ee first. Let conf⁡(e)\operatorname{conf}(e) be the mean gate value assigned to ee over tokens for which it is ranked first. Both quantities are dimensionless and lie in [0,1][0,1]. The pruning importance score is

    imp⁡(e)=top1⁡(e)exp⁡(conf⁡(e)).\operatorname{imp}(e)=\operatorname{top1}(e)\exp(\operatorname{conf}(e)).

    Experts with higher scores are retained before those with lower scores. The paper also evaluates first-choice activity alone, activity over the top two choices, load balancing (first-choice activity multiplied by mean gate value), and the unexponentiated score top1⁡(e)conf⁡(e)\operatorname{top1}(e)\operatorname{conf}(e).

  3. Knowl 3 — Gate-statistic collection and evaluation protocol

    experimental setup

    The study uses FLORES-200, whose parallel translations cover 3,001 English sentences and 201 other languages. Gate statistics are first collected on FLORES-200 dev data while decoding with beam size 4; final translation scores are measured on FLORES-200 devtest. Intermediate comparisons use 30 languages, with 10 languages in each of the high-, low-, and very-low-resource categories. A 53-language subset is used for final subset evaluation, and language-specific pruning is also evaluated across all 202 languages. To obtain per-language statistics across all languages, the authors decode 25 randomly sampled line pairs per language direction using teacher forcing, which they report performs as well as beam-search decoding for this purpose. Translation quality is evaluated primarily with chrF++; spBLEU is also reported.

  4. Knowl 4 — Language-specific pruning preserves quality across all 202 languages

    empirical result

    On FLORES-200 devtest across all 40,602 translation directions, importance-based language-specific pruning retaining 216 encoder and 72 decoder experts achieved an average chrF++ of 35.46. The unpruned 54.5B model scored 35.74, while the 3.3B dense model scored 34.64. Thus, the pruned model was 0.28 chrF++ below the full model and 0.82 above the dense baseline, while using the paper's configuration for single-32GB-GPU decoding.

  5. Knowl 5 — Language-specific expert selection beats one global configuration

    empirical result

    On the 53-language FLORES-200 devtest subset, with importance-based pruning reported as 80%, the average chrF++ scores were 36.81 for the unpruned 54.5B model, 35.81 for the 3.3B dense model, 36.59 for language-pair-specific expert selection, 36.61 for language-specific selection, and 35.34 for a single globally selected expert set. Language-specific selection therefore matched language-pair selection closely while requiring fewer configurations, and both outperformed the global selection in this comparison.

  6. Knowl 6 — Importance ranking and global thresholding perform best in 75%-pruning validation comparisons

    empirical result

    In validation experiments on 30 languages, retaining 25% of the experts, the unpruned model averaged 35.07 chrF++ and the dense baseline 34.06. With a balanced fixed number of experts per layer, importance scoring reached 34.66 chrF++, compared with 34.64 for top-1 activity, 33.93 for top-2 activity, 34.01 for the load-balancing score, and 33.00 for the unexponentiated importance score. Global-threshold selection with importance reached 35.09 chrF++, the highest score in this comparison. Global thresholding allocates experts according to a shared metric threshold rather than a fixed per-layer count, keeping at least four experts per layer in the experiments. The authors nevertheless selected fixed per-layer pruning for subsequent experiments because global thresholding gives language-dependent memory requirements, requires rebuilding or reloading configurations across directions, and was more sensitive to over-generation.

  7. Knowl 7 — Global-threshold pruning allocates far more experts to the encoder

    empirical result

    With importance-based global-threshold pruning at 75% on the validation directions, the selected 384 experts averaged 335 in the encoder and 49 in the decoder. Across language-resource pair categories, encoder allocations ranged from 314 to 348 experts and decoder allocations from 36 to 70. This allocation pattern indicates that, under this pruning criterion and setting, the encoder generally needs a larger share of the retained expert capacity than the decoder.

  8. Knowl 8 — Decoder expert selections show target-language specialization

    empirical result

    The authors compared Jaccard overlap among the top 25% of importance-ranked experts selected separately for language pairs, using equal encoder and decoder expert counts. Decoder selections for language pairs sharing a target language had 68–87% Jaccard similarity; selections for different target languages had 13–39% similarity. Encoder selections were independent of target language, consistent with the model introducing the target-language code only on the decoder side, and encoder overlap between different source languages was still 30–50%. Clustering decoder importance profiles also grouped some linguistically related languages, including Yue Chinese with Korean and Japanese, Russian with Belarusian, and Portuguese with Asturian and French. These analyses support the authors’ interpretation that language-specific experts emerge in NLLB-200, particularly in the decoder.

  9. Knowl 9 — Pruned-model inference speed and hardware requirements

    empirical result

    In a FLORES validation-set speed benchmark translating 29 languages into English, the 80%-pruned model with 36 experts per encoder layer and 12 per decoder layer ran on one V100 GPU with batch size 4k at 172 words per second (WPS), taking 135 seconds. The 3.3B dense model, also on one GPU with batch size 4k, reached 246 WPS and took 86 seconds. The full 54.5B model on four GPUs with batch size 16k reached 156 WPS and took 131 seconds; the pruned model on four GPUs with batch size 16k reached 299 WPS and took 79 seconds. The different batch sizes mean the one-GPU comparison is not a like-for-like throughput comparison, but the results demonstrate that the pruned model fits on one GPU and can be substantially faster than the full model when run on four GPUs.

  10. Knowl 10 — Pruning can increase over-generation, and findings are model-specific

    limitation

    The authors observed occasional repetitions, paraphrases, and slight hallucinations after pruning, which they hypothesize contribute to a larger performance drop under spBLEU than under chrF++. In their validation analysis, 80% global-threshold pruning had a 1.75-point spBLEU drop relative to the full model, compared with a 0.53-point drop for fixed-per-layer pruning. The corresponding average output-length-ratio difference was 0.16 for global thresholding and 0.04 for fixed-per-layer pruning, with positive values indicating longer outputs. The study evaluates only NLLB-200, so its conclusions may depend on this model and its training; the authors also did not fine-tune pruned models, and note that pruning could amplify biases already present in the full model.

Coverage note — Detailed per-direction scores and decoding outputs, exhaustive language lists and similarity matrices, and the full GPU-hour accounting are omitted because they are supporting artifacts or exhaustive diagnostics rather than additional central findings.

References

  1. 1.Roee Aharoni, Melvin Johnson, and Orhan Firat. 2019. Massively multilingual neural machine translation. arXiv preprint arXiv:1903.00089.
  2. 2.Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, Bryan Catanzaro, et al. 2022. ediffi: Text-to-image diffusion models with an ensemble of expert denoisers. arXiv preprint arXiv:2211.01324.
  3. 3.Tianyu Chen, Shaohan Huang, Yuan Xie, Binxing Jiao, Daxin Jiang, Haoyi Zhou, Jianxin Li, and Furu Wei. 2022. Task-specific expert pruning for sparse mixture-of-experts. arXiv preprint arXiv:2206.00277.
  4. 4.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Unsupervised cross-lingual representation learning at scale. arXiv preprint arXiv:1911.02116.
  5. 5.Marta R Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al. 2022. No language left behind: Scaling human-centered machine translation. arXiv preprint arXiv:2207.04672.
  6. 6.Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. 2019. Transformer-xl: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860.
  7. 7.Nan Du, Yanping Huang, Andrew M Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, Barret Zoph, Liam Fedus, Maarten P Bosma, Zongwei Zhou, Tao Wang, Emma Wang, Kellie Webster, Marie Pellat, Kevin Robinson, Kathleen Meier-Hellstern, Toju Duke, Lucas Dixon, Kun Zhang, Quoc Le, Yonghui Wu, Zhifeng Chen, and Claire Cui. 2022. GLaM: Efficient scaling of language models with mixture-of-experts. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 5547–5569. PMLR.
  8. 8.Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, et al. 2021. Beyond english-centric multilingual machine translation. J. Mach. Learn. Res., 22(107):1–48.
  9. 9.William Fedus, Jeff Dean, and Barret Zoph. 2022. A review of sparse expert models in deep learning. arXiv preprint arXiv:2209.01667.
  10. 10.William Fedus, Barret Zoph, and Noam Shazeer. 2021. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.
  11. 11.Zhida Feng, Zhenyu Zhang, Xintong Yu, Yewei Fang, Lanxin Li, Xuyi Chen, Yuxiang Lu, Jiaxiang Liu, Weichong Yin, Shikun Feng, et al. 2022. Ernie-vilg 2.0: Improving text-to-image diffusion model with knowledge-enhanced mixture-of-denoising-experts. arXiv preprint arXiv:2210.15257.
  12. 12.William Held and Diyi Yang. 2022. Shapley head pruning: Identifying and removing interference in multilingual transformers. arXiv preprint arXiv:2210.05709.
  13. 13.Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton. 1991. Adaptive mixtures of local experts. Neural computation, 3(1):79–87.
  14. 14.Jae-young Jo and Sung-Hyon Myaeng. 2020. Roles and utilization of attention heads in transformer-based neural language models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3404–3417.
  15. 15.Michael I Jordan and Robert A Jacobs. 1994. Hierarchical mixtures of experts and the em algorithm. Neural computation, 6(2):181–214.
  16. 16.Young Jin Kim, Ammar Ahmad Awan, Alexandre Muzio, Andres Felipe Cruz Salinas, Liyang Lu, Amr Hendy, Samyam Rajbhandari, Yuxiong He, and Hany Hassan Awadalla. 2021a. Scalable and efficient moe training for multitask multilingual models. arXiv preprint arXiv:2109.10465.
  17. 17.Zae Myung Kim, Laurent Besacier, Vassilina Nikoulina, and Didier Schwab. 2021b. Do multilingual neural machine translation models contain language pair specific attention heads? arXiv preprint arXiv:2105.14940.
  18. 18.Sneha Kudugunta, Yanping Huang, Ankur Bapna, Maxim Krikun, Dmitry Lepikhin, Minh-Thang Luong, and Orhan Firat. 2021. Beyond distillation: Task-level mixture-of-experts for efficient inference. arXiv preprint arXiv:2110.03742.
  19. 19.Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020. Gshard: Scaling giant models with conditional computation and automatic sharding. arXiv preprint arXiv:2006.16668.
  20. 20.Paul Michel, Omer Levy, and Graham Neubig. 2019. Are sixteen heads really better than one? In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  21. 21.Ali Mohammadshahi, Vassilina Nikoulina, Alexandre Berard, Caroline De Brun, James Henderson, and Laurent Besacier. 2022a. Small-100: Introducing shallow multilingual machine translation model for low-resource languages. ArXiv, abs/2210.11621.
  22. 22.Ali Mohammadshahi, Vassilina Nikoulina, Alexandre Berard, Caroline De Brun, James Henderson, and Laurent Besacier. 2022b. What do compressed multilingual machine translation models forget? ArXiv, abs/2205.10828.
  23. 23.Basil Mustafa, Carlos Riquelme, Joan Puigcerver, Rodolphe Jenatton, and Neil Houlsby. 2022. Multimodal contrastive learning with limoe: the language-image mixture of experts. arXiv preprint arXiv:2206.02770.
  24. 24.David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean. 2021. Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350.
  25. 25.Maja Popovic. 2015. chrf: character n-gram f-score for automatic mt evaluation. In Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 392–395.
  26. 26.Matt Post. 2018. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191, Brussels, Belgium. Association for Computational Linguistics.
  27. 27.Joan Puigcerver, Carlos Riquelme, Basil Mustafa, Cedric Renggli, André Susano Pinto, Sylvain Gelly, Daniel Keysers, and Neil Houlsby. 2020. Scalable transfer learning with expert models. arXiv preprint arXiv:2009.13239.
  28. 28.Carlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann, Rodolphe Jenatton, André Susano Pinto, Daniel Keysers, and Neil Houlsby. 2021. Scaling vision with sparse mixture of experts. Advances in Neural Information Processing Systems, 34:8583–8595.
  29. 29.Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538.
  30. 30.Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019. Energy and policy considerations for deep learning in nlp. arXiv preprint arXiv:1906.02243.
  31. 31.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2818–2826.
  32. 32.Yuqing Tang, Chau Tran, Xian Li, Peng-Jen Chen, Naman Goyal, Vishrav Chaudhary, Jiatao Gu, and Angela Fan. 2020. Multilingual translation with extensible multilingual pretraining and finetuning. arXiv preprint arXiv:2008.00401.
  33. 33.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  34. 34.Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. arXiv preprint arXiv:1905.09418.
  35. 35.Hongyu Wang, Shuming Ma, Li Dong, Shaohan Huang, Dongdong Zhang, and Furu Wei. 2022. Deepnet: Scaling transformers to 1,000 layers. arXiv preprint arXiv:2203.00555.
  36. 36.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019. Xlnet: Generalized autoregressive pretraining for language understanding. Advances in neural information processing systems, 32.
  37. 37.Zhao You, Shulin Feng, Dan Su, and Dong Yu. 2021. Speechmoe: Scaling to large acoustic models with dynamic routing mixture of experts. arXiv preprint arXiv:2105.03036.
  38. 38.Seniha Esen Yuksel, Joseph N Wilson, and Paul D Gader. 2012. Twenty years of mixture of experts. IEEE transactions on neural networks and learning systems, 23(8):1177–1193.
  39. 39.Biao Zhang, Philip Williams, Ivan Titov, and Rico Sennrich. 2020. Improving massively multilingual neural machine translation and zero-shot translation. arXiv preprint arXiv:2004.11867.
  40. 40.Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. 2022. St-moe: Designing stable and transferable sparse expert models, 2022. URL https://arxiv.org/abs/2202.08906.

Citation

MLA
Koishekenov, Y., et al. “Memory-efficient NLLB-200: Language-specific Expert Pruning of a Massively Multilingual Machine Translation Model”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 3567–85, https://doi.org/10.18653/v1/2023.acl-long.198.
APA
Koishekenov, Y., Bérard, A., & Nikoulina, V. (2023). Memory-efficient NLLB-200: Language-specific Expert Pruning of a Massively Multilingual Machine Translation Model. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3567–3585. https://doi.org/10.18653/v1/2023.acl-long.198
Chicago
Koishekenov, Y., A. Bérard, and V. Nikoulina. 2023. “Memory-efficient NLLB-200: Language-specific Expert Pruning of a Massively Multilingual Machine Translation Model”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3567–85. https://doi.org/10.18653/v1/2023.acl-long.198.
Harvard
Koishekenov, Y., Bérard, A. and Nikoulina, V. (2023) “Memory-efficient NLLB-200: Language-specific Expert Pruning of a Massively Multilingual Machine Translation Model”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3567–3585. Available at: https://doi.org/10.18653/v1/2023.acl-long.198.
Vancouver
1. Koishekenov Y, Bérard A, Nikoulina V (2023) Memory-efficient NLLB-200: Language-specific Expert Pruning of a Massively Multilingual Machine Translation Model. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 3567–3585

BibTeX

@inproceedings{koishekenov-etal-2023-memory,
    title = "Memory-efficient {NLLB}-200: Language-specific Expert Pruning of a Massively Multilingual Machine Translation Model",
    author = "Koishekenov, Yeskendir  and
      Berard, Alexandre  and
      Nikoulina, Vassilina",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.198/",
    doi = "10.18653/v1/2023.acl-long.198",
    pages = "3567--3585"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/