Revisiting Machine Translation for Cross-lingual Classification

Mikel ArtetxeVedanuj GoswamiShruti BhosaleAngela FanLuke Zettlemoyer

article2023EMNLP61 citations

Demonstrates that modern machine translation combined with domain adaptation allows simple translate-test pipelines to outperform multilingual pretrained models on complex cross-lingual classification benchmarks like XNLI.

Listen

Organizations deploying artificial intelligence across global markets must accurately classify text across dozens of languages, yet training data is often available only in English. Prevailing industry practice relies on massively multilingual pretrained models that directly process non-English inputs (zero-shot transfer) or learn from machine-translated training data (translate-train). Meanwhile, the simpler pipeline of translating non-English text back into English at runtime to use a strong English-only classifier (translate-test) has been largely dismissed as inferior due to translation errors and poor baseline benchmarks.

The article systematically evaluates the role of modern machine translation systems in cross-lingual classification. Its primary objective is to demonstrate that translate-test can outperform state-of-the-art multilingual models when using high-capacity translation engines and mitigating data mismatch, while explaining why optimal deployment strategies vary substantially across different types of tasks.

To establish these findings, the authors conducted extensive experiments across six multilingual classification benchmarks covering diverse language families and task complexities, including natural language inference, paraphrase detection, sentiment analysis, and multiple-choice commonsense reasoning. They tested leading monolingual and multilingual language models alongside a high-capacity 3.3-billion-parameter translation engine. In addition to testing standard configurations, the authors developed techniques to align translation style across training and inference, including document-level translation fine-tuning and roundtrip back-translation data augmentation. They also introduced a diagnostic framework using English datasets to measure and isolate specific sources of performance degradation, such as information loss, translation style mismatch, representation misalignment, and language representation quality.

The investigation produced several key findings. First, translate-test performance has been severely underestimated; combining advanced translation with data adaptation techniques improved overall accuracy by an average of 2.0 to 2.5 percentage points across models. Second, an optimized translate-test pipeline utilizing strong English-only models outperformed top multilingual models on the majority of benchmarks, beating the previous benchmark standard on the XNLI inference task by 8.1 percentage points and achieving an average score of 76.2% with DeBERTa compared to 70.9% for the best multilingual translate-train setup. Third, translate-test demonstrated significant superiority on complex tasks requiring commonsense and world knowledge (e.g., outperforming multilingual translate-train by 13.7 points on causal reasoning), because such tasks benefit heavily from the richer representations available in English-only models. Conversely, multilingual translate-train performed best on linguistically simpler tasks like sentiment review classification, where translation errors and domain noise outweigh the benefits of advanced English models.

These results have direct operational and cost implications for engineering and product leadership. Relying exclusively on complex multilingual model architectures is not always the best technical strategy. Translating inputs to English at test time allows organizations to deploy best-in-class, specialized English language models that deliver superior accuracy on nuanced reasoning tasks. However, runtime machine translation introduces latency, higher computational costs per inference query, and translation error risks on highly specialized terminology. Deploying multilingual models remains the most efficient, cost-effective choice for surface-level classification tasks or applications with strict real-time latency budgets.

Decision-makers should avoid a one-size-fits-all approach and instead select cross-lingual architectures based on task complexity. Engineering teams should run initial pilot benchmarks using the article's diagnostic methodology on English-only data to determine whether target tasks are sensitive to representation quality or translation noise. For high-reasoning tasks adopting translate-test, teams should apply roundtrip translation to English training sets or fine-tune translation models at the document level to minimize distribution mismatch.

These findings should be interpreted with awareness of certain boundaries. The evaluation was limited to text classification and encoder-only architectures, meaning conclusions cannot be directly generalized to generative tasks or sequence labeling without further validation. In addition, machine translation quality remains lower for low-resource languages, which introduces performance uncertainty in those specific dialects.

arXiv: 2305.14240
Cover for Revisiting Machine Translation for Cross-lingual Classification

Abstract

Machine Translation (MT) has been widely used for cross-lingual classification, either by translating the test set into English and running inference with a monolingual model (translate-test), or translating the training set into the target languages and finetuning a multilingual model (translate-train). However, most research in the area focuses on the multilingual models rather than the MT component. We show that, by using a stronger MT system and mitigating the mismatch between training on original text and running inference on machine translated text, translate-test can do substantially better than previously assumed. The optimal approach, however, is highly task dependent, as we identify various sources of cross-lingual transfer gap that affect different tasks and approaches differently. Our work calls into question the dominance of multilingual models for cross-lingual classification, and prompts to pay more attention to MT-based baselines.

Table of Contents

  • 1 Introduction
  • 2 Experimental setup
  • 3 Main results
  • 3.1 Revisiting translate-test
  • 3.1.1 MTadaptation
  • 3.1.2 Training data adaptation
  • 3.2 Revisiting translate-train
  • 3.3 Reconsidering the state-of-the-art
  • 4 Analyzing the variance across tasks
  • 4.1 Sources of cross-lingual transfer gap
  • 4.2 Methodology
  • 4.3 Results and discussion
  • 5 Related work
  • 6 Conclusions
  • Limitations
  • References

Knowls

  1. Knowl 1 — Framework for Decomposing Sources of Cross-Lingual Transfer Gap

    model/method

    Let TsrcT_{\text{src}} denote downstream training data in a source language (e.g., English), Ttgt=MTsrc→tgt(Tsrc)T_{\text{tgt}} = \text{MT}_{\text{src} \to \text{tgt}}(T_{\text{src}}) its machine translation into a target language, and Tbt=MTtgt→src(Ttgt)T_{\text{bt}} = \text{MT}_{\text{tgt} \to \text{src}}(T_{\text{tgt}}) its back-translation into the source language. Let EsrcE_{\text{src}}, EtgtE_{\text{tgt}}, and EbtE_{\text{bt}} denote evaluation sets defined analogously. Let acc(T,E)\text{acc}(T, E) denote the test accuracy of a multilingual model trained on dataset TT and evaluated on dataset EE, and let accmono(T,E)\text{acc}_{\text{mono}}(T, E) denote the accuracy of an equivalent monolingual source-language model trained on TT and evaluated on EE.

    The cross-lingual performance transfer gap can be estimated and decomposed into five distinct underlying sources:

    1. MT Information Lost (ΔinfoMT\Delta^{\text{MT}}_{\text{info}}): The accuracy drop due to semantic information lost during machine translation, assuming symmetric error across forward and backward translation steps: ΔinfoMT=acc(Tsrc,Esrc)−acc(Tbt,Ebt)2\Delta^{\text{MT}}_{\text{info}} = \frac{\text{acc}(T_{\text{src}}, E_{\text{src}}) - \text{acc}(T_{\text{bt}}, E_{\text{bt}})}{2}

    2. MT Distribution Shift (ΔdistMT\Delta^{\text{MT}}_{\text{dist}}): The performance degradation caused by stylistic and distributional mismatch between original human text and machine-translated text: ΔdistMT=acc(Tbt,Ebt)−acc(Tsrc,Ebt)2\Delta^{\text{MT}}_{\text{dist}} = \frac{\text{acc}(T_{\text{bt}}, E_{\text{bt}}) - \text{acc}(T_{\text{src}}, E_{\text{bt}})}{2}

    3. Source Representation Quality (Δsrcrep\Delta^{\text{rep}}_{\text{src}}): The performance advantage of a monolingual source-language model over a multilingual model when both train and evaluate on original source-language data: Δsrcrep=accmono(Tsrc,Esrc)−acc(Tsrc,Esrc)\Delta^{\text{rep}}_{\text{src}} = \text{acc}_{\text{mono}}(T_{\text{src}}, E_{\text{src}}) - \text{acc}(T_{\text{src}}, E_{\text{src}})

    4. Target Representation Quality (Δtgtrep\Delta^{\text{rep}}_{\text{tgt}}): The representation deficit of the target language within the multilingual model relative to the source language, subtracting translation error: Δtgtrep=acc(Tsrc,Esrc)−acc(Ttgt,Etgt)−ΔinfoMT\Delta^{\text{rep}}_{\text{tgt}} = \text{acc}(T_{\text{src}}, E_{\text{src}}) - \text{acc}(T_{\text{tgt}}, E_{\text{tgt}}) - \Delta^{\text{MT}}_{\text{info}}

    5. Representation Misalignment (Δalignrep\Delta^{\text{rep}}_{\text{align}}): The generalization penalty due to imperfect alignment between source and target representations in the multilingual model, subtracting the distribution shift error: Δalignrep=acc(Ttgt,Etgt)−acc(Tsrc,Etgt)−ΔdistMT\Delta^{\text{rep}}_{\text{align}} = \text{acc}(T_{\text{tgt}}, E_{\text{tgt}}) - \text{acc}(T_{\text{src}}, E_{\text{tgt}}) - \Delta^{\text{MT}}_{\text{dist}}

  2. Knowl 2 — Performance Comparison of Cross-Lingual Transfer Paradigms Across Multiple Backbones

    data/table

    Cross-lingual classification accuracy across six benchmarks demonstrates that an improved translate-test strategy (combining document-level machine translation adaptation with roundtrip training data adaptation using NLLB 3.3B) enables monolingual English models to outperform multilingual models.

    Model Method XNLI PAWS-X MARC XCOPA XStoryCloze EXAMS Avg
    XLM-R zero-shot 80.1 87.1 60.6 69.1 84.6 36.0 69.6
    XLM-R translate-train 84.3 90.7 60.8 67.6 86.7 35.2 70.9
    XLM-R translate-test (vanilla) 79.3 86.9 58.0 68.6 84.8 34.9 68.8
    XLM-R translate-test (improved) 84.6 89.3 58.8 69.0 87.9 35.0 70.8
    RoBERTa translate-test (vanilla) 79.9 87.3 57.6 72.9 89.3 36.3 70.6
    RoBERTa translate-test (improved) 85.9 89.3 59.1 75.7 91.2 36.4 72.9
    DeBERTaV3 translate-test (vanilla) 81.0 87.1 58.2 77.7 92.1 46.1 73.7
    DeBERTaV3 translate-test (improved) 86.7 90.3 59.2 81.3 93.8 46.0 76.2

    All models utilize 304M backbone parameters. All systems use the NLLB 3.3B translation model. Improved translate-test outperforms vanilla translate-test across all pretrained backbones by an average of 2.0 points for XLM-R, 2.3 points for RoBERTa, and 2.5 points for DeBERTaV3. Using DeBERTaV3 in the improved translate-test configuration achieves the highest overall average score (76.2%), outperforming XLM-R translate-train (70.9%) by 5.3 percentage points.

  3. Knowl 3 — Machine Translation Adaptation for Translate-Test Cross-Lingual Classification

    model/method

    To mitigate the train/test mismatch in translate-test caused by domain differences between general-purpose machine translation (MT) and task data, the MT engine (NLLB 3.3B) can be fine-tuned directly on synthetic parallel downstream data:

    1. Sentence-Level Domain Adaptation (+dom adapt): Downstream English training data is segmented into sentences (using spaCy's xx_sent_ud_sm model) and back-translated into all target languages using beam search (k=4k=4). The MT model is fine-tuned to translate from the target languages into English on these sentence pairs using a learning rate of 5×10−55 \times 10^{-5}, batch size of 32k32\text{k} tokens, disabled dropout, and checkpoint selection at 25k25\text{k} steps.

    2. Document-Level Adaptation (+doc level): The back-translated target sentences for an instance are concatenated, and all input fields (e.g., premise and hypothesis in NLI, or questions and options in multiple choice) are separated by a special <sep> token. The MT model is fine-tuned to translate the full concatenated example jointly and performs document-level translation at inference time.

    On RoBERTa-large translate-test across six benchmarks, document-level fine-tuning yields the highest accuracy: 83.8%83.8\% on XNLI (vs. 79.9%79.9\% for vanilla NLLB and 76.8%76.8\% for official benchmark translations), 87.8%87.8\% on PAWS-X (vs. 87.3%87.3\%), 58.3%58.3\% on MARC (vs. 57.6%57.6\%), 76.3%76.3\% on XCOPA (vs. 72.9%72.9\%), and 91.3%91.3\% on XStoryCloze (vs. 89.3%89.3\%).

  4. Knowl 4 — Training Data Adaptation via Roundtrip Back-Translation for Translate-Test

    model/method

    To adapt downstream classifier fine-tuning to the distribution of machine-translated text encountered at test time, training examples are augmented via roundtrip translation:

    1. For a downstream dataset with kk English examples and nn target languages, each training instance is translated from English into each target language and then back-translated into English using the same MT model (NLLB 3.3B) used at test time.
    2. Forward translation (out of English) utilizes beam search (k=4k=4) for multiple-choice tasks (to avoid noisy short options) and nucleus sampling (p=0.8p=0.8) for classification tasks (to enhance diversity). Backward translation (into English) uses beam search (k=4k=4) across all tasks.
    3. The resulting k×nk \times n back-translated examples are concatenated with the original kk English examples to create an augmented dataset of k×(n+1)k \times (n+1) instances.
    4. The classifier is fine-tuned on this augmented dataset for a single epoch.

    Using RoBERTa-large with vanilla NLLB test-time translation, roundtrip data adaptation increases test accuracy over original English training data from 79.9%79.9\% to 85.2%85.2\% on XNLI, 87.3%87.3\% to 89.9%89.9\% on PAWS-X, 57.6%57.6\% to 58.8%58.8\% on MARC, 72.9%72.9\% to 74.3%74.3\% on XCOPA, and 89.3%89.3\% to 91.2%91.2\% on XStoryCloze.

  5. Knowl 5 — Task Complexity and Resource Level Effects on Cross-Lingual Transfer Gaps

    empirical result

    Decomposing the cross-lingual transfer gap across 15 target languages (5 high-resource related, 5 high-resource unrelated, and 5 low-resource unrelated) explains why different downstream tasks favor different transfer paradigms:

    1. Commonsense and Knowledge-Intensive Tasks: Tasks such as XStoryCloze and OpenBookQA/EXAMS exhibit high sensitivity to representation quality (Δsrcrep\Delta^{\text{rep}}_{\text{src}} and Δtgtrep\Delta^{\text{rep}}_{\text{tgt}}). Because English-only models (RoBERTa, DeBERTaV3) learn substantially higher-quality representations than multilingual models (XLM-R), translate-test achieves superior performance on these benchmarks (e.g., DeBERTaV3 translate-test exceeds XLM-R translate-train by 13.7 points on XCOPA).
    2. Shallow Semantic Tasks: Tasks such as sentiment analysis (MARC) are largely insensitive to source and target representation quality. For these tasks, the translation noise and information lost during MT (ΔinfoMT\Delta^{\text{MT}}_{\text{info}}) outweigh the benefits of a stronger monolingual backbone, making multilingual translate-train preferable.
    3. Representation Alignment: Representation misalignment (Δalignrep\Delta^{\text{rep}}_{\text{align}}) is generally minor compared to representation quality across languages, indicating that multilingual encoders align shared feature spaces effectively during fine-tuning. However, absolute target representation quality degrades drastically on low-resource and unrelated target languages.
    4. MT Degradation by Task Domain: MT information loss is high on tasks containing domain-specific jargon or isolated, context-free candidate options (e.g., OpenBookQA science questions), but low on coherent narrative text (e.g., XStoryCloze).
  6. Knowl 6 — Impact of Translation Quality and Decoding Strategy on Translate-Train

    empirical result

    Evaluating translate-train using XLM-R large across different MT configurations yields three key empirical findings:

    1. Nucleus Sampling vs. Beam Search: In tasks where translate-train outperforms zero-shot transfer, generating training translations via nucleus sampling (p=0.8p=0.8) outperforms beam search (k=4k=4). XLM-R translate-train with sampling achieves 84.3%84.3\% on XNLI (vs. 83.5%83.5\% for beam search), 90.7%90.7\% on PAWS-X (vs. 90.4%90.4\%), and 86.7%86.7\% on XStoryCloze (vs. 86.4%86.4\%).
    2. Task-Dependent Helpfulness: Translate-train improves accuracy over zero-shot transfer on XNLI (84.3%84.3\% vs. 80.1%80.1\%), PAWS-X (90.7%90.7\% vs. 87.1%87.1\%), and XStoryCloze (86.7%86.7\% vs. 84.6%84.6\%), but performs on par or worse than zero-shot on MARC (60.8%60.8\% vs. 60.6%60.6\%), XCOPA (67.6%67.6\% vs. 69.1%69.1\%), and EXAMS (35.2%35.2\% vs. 36.0%36.0\%).
    3. Differential MT Sensitivity: Translate-test is far more sensitive to MT engine quality than translate-train. On XNLI, moving from official translations to vanilla NLLB creates an accuracy difference of 3.13.1 points for translate-train (favouring official) versus a 3.13.1 point gap for translate-test in favor of NLLB (76.8%→79.9%76.8\% \to 79.9\%).
  7. Knowl 7 — Experimental Setup for Cross-Lingual Classification Benchmarking

    experimental setup

    The experimental evaluation across cross-lingual classification paradigms uses the following unified setup:

    • Pretrained Encoders: XLM-R large (multilingual), RoBERTa large (monolingual English), and DeBERTaV3 large (monolingual English). All models share a 304M backbone parameter count.
    • Machine Translation Engine: NLLB-200 3.3B parameter model.
    • Evaluation Benchmarks: Six datasets: XNLI (15 languages; trained on MultiNLI), PAWS-X (7 languages; trained on PAWS), MARC (6 languages; English training set), XCOPA (10 evaluated languages; Quechua and Haitian Creole excluded due to lack of NLLB or XLM-R support), XStoryCloze (11 languages), and EXAMS (16 languages). For multiple-choice datasets lacking native English training sets (XCOPA, XStoryCloze, EXAMS), models are trained on a combined English pool comprising Social IQa, SWAG, COPA, OpenBookQA, ARC, and PIQA.
    • Fine-Tuning Protocol: HuggingFace Transformers library; Adam optimizer; batch size 64; sequence truncation at 256 tokens; learning rate 6×10−66 \times 10^{-6} with linear decay and 50 warmup steps; final checkpoint used without task-specific model selection. Multiple-choice tasks evaluate input-candidate pairs independently via first-token projection and softmax scoring. Training duration is 2 epochs for original English datasets and 1 epoch for back-translated augmented datasets. Reported scores are accuracies averaged across all supported languages and 5 random seeds.
  8. Knowl 8 — Scope and Methodological Limitations of the Cross-Lingual MT Study

    limitation

    The empirical findings and transfer-gap analysis are subject to three specific scope limitations:

    1. Task Modality: The study is restricted to sentence classification and multiple-choice tasks. It does not examine structured prediction (e.g., sequence labeling, which necessitates cross-lingual label projection) or free-form text generation.
    2. Translation Artifacts in Existing Benchmarks: Many multilingual evaluation datasets were constructed via human translation from English. While prior work shows that translation artifacts favor translate-train and disadvantage translate-test, subtle artifacts (such as translating text that was already a translation) may influence results.
    3. Model Paradigm: Experiments are limited to encoder-only masked language models within the pretrain-finetune paradigm, leaving open the interaction of translate-test and translate-train with modern multilingual autoregressive decoder-only language models.

Coverage note — None was omitted; all key contributions, methodology definitions, empirical tables, and qualitative transfer gap analyses were included.

References

  1. 1.Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2020. Translation artifacts in cross-lingual transfer learning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7674–7684, Online. Association for Computational Linguistics.
  2. 2.Carmen Banea, Rada Mihalcea, Janyce Wiebe, and Samer Hassan. 2008. Multilingual subjectivity analysis using machine translation. In Proceedings of the 2008 Conference on Empirical Methods in Natural Language Processing, pages 127–135, Honolulu, Hawaii. Association for Computational Linguistics.
  3. 3.Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. 2020. PIQA: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, pages 7432–7439.
  4. 4.Zewen Chi, Shaohan Huang, Li Dong, Shuming Ma, Bo Zheng, Saksham Singhal, Payal Bajaj, Xia Song, Xian-Ling Mao, Heyan Huang, and Furu Wei. 2022. XLM-E: Cross-lingual language model pre-training via ELECTRA. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6170–6182, Dublin, Ireland. Association for Computational Linguistics.
  5. 5.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge.
  6. 6.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  7. 7.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2475–2485, Brussels, Belgium. Association for Computational Linguistics.
  8. 8.Sumanth Doddapaneni, Gowtham Ramesh, Mitesh M. Khapra, Anoop Kunchukuttan, and Pratyush Kumar. 2021. A primer on pretrained multilingual language models.
  9. 9.Kevin Duh, Akinori Fujino, and Masaaki Nagata. 2011. Is machine translation ripe for cross-lingual sentiment classification? In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 429–433, Portland, Oregon, USA. Association for Computational Linguistics.
  10. 10.Sergey Edunov, Myle Ott, Michael Auli, and David Grangier. 2018. Understanding back-translation at scale. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 489–500, Brussels, Belgium. Association for Computational Linguistics.
  11. 11.Yuwei Fang, Shuohang Wang, Zhe Gan, Siqi Sun, and Jingjing Liu. 2021. Filter: An enhanced fusion method for cross-lingual language understanding. Proceedings of the AAAI Conference on Artificial Intelligence, 35(14):12776–12784.
  12. 12.Hao Fei, Meishan Zhang, and Donghong Ji. 2020. Cross-lingual semantic role labeling with high-quality translated training corpus. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7014–7026, Online. Association for Computational Linguistics.
  13. 13.Blaz Fortuna and John Shawe-Taylor. 2005. The use of machine translation tools for cross-lingual text mining. In Proceedings of the ICML Workshop on Learning with Multiple Views.
  14. 14.Iker García-Ferrero, Rodrigo Agerri, and German Rigau. 2022a. Model and data transfer for cross-lingual sequence labelling in zero-resource settings. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 6403–6416, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  15. 15.Iker García-Ferrero, Rodrigo Agerri, and German Rigau. 2022b. T-projection: High quality annotation projection for sequence labeling tasks.
  16. 16.Momchil Hardalov, Todor Mihaylov, Dimitrina Zlatkova, Yoan Dinkov, Ivan Koychev, and Preslav Nakov. 2020. EXAMS: A multi-subject high school examinations dataset for cross-lingual and multilingual question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5427–5444, Online. Association for Computational Linguistics.
  17. 17.Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2021. Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing.
  18. 18.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. In International Conference on Learning Representations.
  19. 19.Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020. XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 4411–4421. PMLR.
  20. 20.Tim Isbister, Fredrik Carlsson, and Magnus Sahlgren. 2021. Should we stop training more monolingual models, and simply use machine translation instead? In Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa), pages 385–390, Reykjavik, Iceland (Online). Linköping University Electronic Press, Sweden.
  21. 21.Alankar Jain, Bhargavi Paranjape, and Zachary C. Lipton. 2019. Entity projection via machine translation for cross-lingual NER. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1083–1092, Hong Kong, China. Association for Computational Linguistics.
  22. 22.Phillip Keung, Yichao Lu, György Szarvas, and Noah A. Smith. 2020. The multilingual Amazon reviews corpus. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4563–4568, Online. Association for Computational Linguistics.
  23. 23.Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, and Xian Li. 2021. Few-shot learning with multilingual language models.
  24. 24.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach.
  25. 25.Zihan Liu, Genta Indra Winata, Andrea Madotto, and Pascale Fung. 2021. Preserving cross-linguality of pre-trained models via continual learning. In Proceedings of the 6th Workshop on Representation Learning for NLP (RepL4NLP-2021), pages 64–71, Online. Association for Computational Linguistics.
  26. 26.Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? a new dataset for open book question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2381–2391, Brussels, Belgium. Association for Computational Linguistics.
  27. 27.NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. 2022. No language left behind: Scaling human-centered machine translation.
  28. 28.Jaehoon Oh, Jongwoo Ko, and Se-Young Yun. 2022. Synergy with translation artifacts for training and inference in multilingual tasks. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 6747–6754, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  29. 29.Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulić, and Anna Korhonen. 2020. XCOPA: A multilingual dataset for causal commonsense reasoning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2362–2376, Online. Association for Computational Linguistics.
  30. 30.Edoardo Maria Ponti, Julia Kreutzer, Ivan Vulić, and Siva Reddy. 2021. Modelling latent translations for cross-lingual transfer.
  31. 31.Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon. 2011. Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In AAAI Spring Symposium: Logical Formalizations of Commonsense Reasoning, pages 90–95.
  32. 32.Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, and Melvin Johnson. 2021. XTREME-R: Towards more challenging and nuanced multilingual evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10215–10245, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  33. 33.Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019. Social IQa: Commonsense reasoning about social interactions. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4463–4473, Hong Kong, China. Association for Computational Linguistics.
  34. 34.Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. 2022. Language models are multilingual chain-of-thought reasoners.
  35. 35.Lei Shi, Rada Mihalcea, and Mingjun Tian. 2010. Cross language text classification by model translation and semi-supervised learning. In Proceedings of the 2010 Conference on Empirical Methods in Natural Language Processing, pages 1057–1067, Cambridge, MA. Association for Computational Linguistics.
  36. 36.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122, New Orleans, Louisiana. Association for Computational Linguistics.
  37. 37.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2019. Huggingface’s transformers: State-of-the-art natural language processing.
  38. 38.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online. Association for Computational Linguistics.
  39. 39.Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. 2019. PAWS-X: A cross-lingual adversarial dataset for paraphrase identification. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3687–3692, Hong Kong, China. Association for Computational Linguistics.
  40. 40.Tao Yu and Shafiq Joty. 2021. Effective fine-tuning methods for cross-lingual adaptation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8492–8501, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  41. 41.Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018. SWAG: A large-scale adversarial dataset for grounded commonsense inference. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 93–104, Brussels, Belgium. Association for Computational Linguistics.
  42. 42.Yuan Zhang, Jason Baldridge, and Luheng He. 2019. PAWS: Paraphrase adversaries from word scrambling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1298–1308, Minneapolis, Minnesota. Association for Computational Linguistics.
  43. 43.Bo Zheng, Li Dong, Shaohan Huang, Wenhui Wang, Zewen Chi, Saksham Singhal, Wanxiang Che, Ting Liu, Xia Song, and Furu Wei. 2021. Consistency regularization for cross-lingual fine-tuning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3403–3417, Online. Association for Computational Linguistics.

Citation

MLA
Artetxe, M., et al. “Revisiting Machine Translation for Cross-lingual Classification”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 6489–99, https://doi.org/10.18653/v1/2023.emnlp-main.399.
APA
Artetxe, M., Goswami, V., Bhosale, S., Fan, A., & Zettlemoyer, L. (2023). Revisiting Machine Translation for Cross-lingual Classification. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 6489–6499. https://doi.org/10.18653/v1/2023.emnlp-main.399
Chicago
Artetxe, M., V. Goswami, S. Bhosale, A. Fan, and L. Zettlemoyer. 2023. “Revisiting Machine Translation for Cross-lingual Classification”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 6489–99. https://doi.org/10.18653/v1/2023.emnlp-main.399.
Harvard
Artetxe, M. et al. (2023) “Revisiting Machine Translation for Cross-lingual Classification”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 6489–6499. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.399.
Vancouver
1. Artetxe M, Goswami V, Bhosale S, Fan A, Zettlemoyer L (2023) Revisiting Machine Translation for Cross-lingual Classification. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 6489–6499

BibTeX

@inproceedings{artetxe-etal-2023-revisiting,
    title = "Revisiting Machine Translation for Cross-lingual Classification",
    author = "Artetxe, Mikel  and
      Goswami, Vedanuj  and
      Bhosale, Shruti  and
      Fan, Angela  and
      Zettlemoyer, Luke",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.399/",
    doi = "10.18653/v1/2023.emnlp-main.399",
    pages = "6489--6499"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/