Revisiting Machine Translation for Cross-lingual Classification
Mikel ArtetxeVedanuj GoswamiShruti BhosaleAngela FanLuke Zettlemoyer
Demonstrates that modern machine translation combined with domain adaptation allows simple translate-test pipelines to outperform multilingual pretrained models on complex cross-lingual classification benchmarks like XNLI.
Organizations deploying artificial intelligence across global markets must accurately classify text across dozens of languages, yet training data is often available only in English. Prevailing industry practice relies on massively multilingual pretrained models that directly process non-English inputs (zero-shot transfer) or learn from machine-translated training data (translate-train). Meanwhile, the simpler pipeline of translating non-English text back into English at runtime to use a strong English-only classifier (translate-test) has been largely dismissed as inferior due to translation errors and poor baseline benchmarks.
The article systematically evaluates the role of modern machine translation systems in cross-lingual classification. Its primary objective is to demonstrate that translate-test can outperform state-of-the-art multilingual models when using high-capacity translation engines and mitigating data mismatch, while explaining why optimal deployment strategies vary substantially across different types of tasks.
To establish these findings, the authors conducted extensive experiments across six multilingual classification benchmarks covering diverse language families and task complexities, including natural language inference, paraphrase detection, sentiment analysis, and multiple-choice commonsense reasoning. They tested leading monolingual and multilingual language models alongside a high-capacity 3.3-billion-parameter translation engine. In addition to testing standard configurations, the authors developed techniques to align translation style across training and inference, including document-level translation fine-tuning and roundtrip back-translation data augmentation. They also introduced a diagnostic framework using English datasets to measure and isolate specific sources of performance degradation, such as information loss, translation style mismatch, representation misalignment, and language representation quality.
The investigation produced several key findings. First, translate-test performance has been severely underestimated; combining advanced translation with data adaptation techniques improved overall accuracy by an average of 2.0 to 2.5 percentage points across models. Second, an optimized translate-test pipeline utilizing strong English-only models outperformed top multilingual models on the majority of benchmarks, beating the previous benchmark standard on the XNLI inference task by 8.1 percentage points and achieving an average score of 76.2% with DeBERTa compared to 70.9% for the best multilingual translate-train setup. Third, translate-test demonstrated significant superiority on complex tasks requiring commonsense and world knowledge (e.g., outperforming multilingual translate-train by 13.7 points on causal reasoning), because such tasks benefit heavily from the richer representations available in English-only models. Conversely, multilingual translate-train performed best on linguistically simpler tasks like sentiment review classification, where translation errors and domain noise outweigh the benefits of advanced English models.
These results have direct operational and cost implications for engineering and product leadership. Relying exclusively on complex multilingual model architectures is not always the best technical strategy. Translating inputs to English at test time allows organizations to deploy best-in-class, specialized English language models that deliver superior accuracy on nuanced reasoning tasks. However, runtime machine translation introduces latency, higher computational costs per inference query, and translation error risks on highly specialized terminology. Deploying multilingual models remains the most efficient, cost-effective choice for surface-level classification tasks or applications with strict real-time latency budgets.
Decision-makers should avoid a one-size-fits-all approach and instead select cross-lingual architectures based on task complexity. Engineering teams should run initial pilot benchmarks using the article's diagnostic methodology on English-only data to determine whether target tasks are sensitive to representation quality or translation noise. For high-reasoning tasks adopting translate-test, teams should apply roundtrip translation to English training sets or fine-tune translation models at the document level to minimize distribution mismatch.
These findings should be interpreted with awareness of certain boundaries. The evaluation was limited to text classification and encoder-only architectures, meaning conclusions cannot be directly generalized to generative tasks or sequence labeling without further validation. In addition, machine translation quality remains lower for low-resource languages, which introduces performance uncertainty in those specific dialects.
- Paper: XNLI: Evaluating Cross-lingual Sentence Representations, Alexis Conneau et al. (2018). This benchmark establishes the standard translate-train and translate-test paradigms for cross-lingual classification that the source paper critically revisits and improves upon.
- Paper: Unsupervised Cross-lingual Representation Learning at Scale, Alexis Conneau et al. (2019). It introduces XLM-RoBERTa, establishing the dominant multilingual masked language model paradigm that the source contrasts with modern machine translation baselines.
- Paper: mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer, Linting Xue et al. (2020). It establishes massively multilingual text-to-text pre-training and standardizes zero-shot versus translate-train evaluation across varied language benchmarks.
- Paper: Cross-lingual Language Model Pretraining, Guillaume Lample et al. (2019). It provides foundational principles for cross-lingual language model pre-training and early comparative evaluation between multilingual representations and translation-based transfer.
- Paper: How Multilingual is Multilingual BERT?, Telmo Pires et al. (2019). It provides key empirical analysis of how multilingual encoders generalize zero-shot across languages, illuminating the baseline behavior investigated by the source.
- Paper: Multilingual Denoising Pre-training for Neural Machine Translation, Yinhan Liu et al. (2020). It demonstrates sequence-to-sequence multilingual pre-training that directly powers high-quality neural machine translation systems used in modern cross-lingual pipelines.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). It extends the evaluation of translate-test workflows to modern generative large language models across 70 typologically diverse languages.
- Paper: MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks, Sanchit Ahuja et al. (2024). It broadens the benchmark scope by assessing state-of-the-art multilingual LLMs across diverse tasks and modalities, building on insights into cross-lingual transfer gaps.
- Paper: MMTEB: Massive Multilingual Text Embedding Benchmark, Kenneth C. Enevoldsen et al. (2025). It significantly scales up cross-lingual representation evaluation across hundreds of languages and tasks, offering a comprehensive arena to test multilingual embedding transfer versus translation.
- Paper: Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages, Zihao Li et al. (2025). It develops representation-level metrics to quantify cross-lingual alignment gaps in language models, directly continuing the analysis of language-transfer bottlenecks.
- Paper: Omnilingual MT: Machine Translation for 1,600 Languages, Omnilingual MT Team et al. (2026). It scales machine translation coverage to more than 1,600 languages, expanding the viability of MT-based cross-lingual transfer to extreme low-resource regimes.
