XNLI: Evaluating Cross-lingual Sentence Representations
Alexis ConneauGuillaume LampleRuty RinottAdina WilliamsSamuel R. BowmanHolger SchwenkVeselin Stoyanov
Introduces the Cross-lingual Natural Language Inference (XNLI) benchmark across 15 diverse languages, establishing a standard evaluation suite and competitive baselines for cross-lingual sentence representations and zero-shot transfer.
Modern natural language processing systems rely heavily on large volumes of annotated data to learn complex tasks. However, these datasets are typically available only in English, leaving international applications constrained because manually labeling data across dozens of languages is prohibitively expensive. To address this challenge, researchers have focused on cross-lingual language understanding, where a model trains primarily on English data and evaluates on other target languages. Despite growing interest, progress has been constrained by the absence of a standardized, large-scale sentence-level evaluation benchmark.
The article introduces and evaluates the Cross-lingual Natural Language Inference (XNLI) corpus, a standardized evaluation benchmark designed to measure cross-lingual sentence understanding across 15 diverse languages. The article demonstrates how well different cross-lingual approaches—ranging from machine translation pipelines to shared multilingual sentence encoders—can transfer knowledge from English to non-English tasks without requiring target-language training data.
To build this benchmark, the authors collected 7,500 new English sentence pairs across ten genres using established crowdsourcing procedures, then hired professional translators to translate them into 14 additional languages, including lower-resource languages such as Swahili and Urdu. This process yielded a fully aligned evaluation suite of 112,500 annotated pairs. The authors then evaluated baseline architectures: translation-based methods that translate data either during training or at test time, and multilingual sentence encoders that align target-language sentence representations directly to English using parallel corpora and a specialized alignment loss function.
The evaluation revealed several key findings. First, the highest overall performance across all languages was achieved by translating incoming foreign test sentences directly into English at inference time, reaching up to 70.7% accuracy in Spanish compared to the English baseline of 73.7%. Second, machine translation quality heavily dictated task accuracy; languages with high translation quality regularly exceeded 70% accuracy, whereas low-resource languages like Swahili and Urdu scored lower, between 58% and 62%. Third, multilingual sentence encoders trained directly with parallel text yielded competitive results (68.9% in Greek and 68.7% in Spanish) without requiring a runtime translation system, though they trailed test-time translation by up to 6 percentage points in lower-resource settings. Finally, bidirectional neural sentence encoders that pooled representations across all hidden states consistently outperformed pretrained word-averaging techniques across every tested language.
These results present clear operational trade-offs for organizations deploying multilingual systems. While test-time machine translation yields the highest accuracy, it is computationally expensive and introduces runtime latency. In contrast, multilingual sentence encoders significantly reduce computational overhead at inference time because they map foreign text directly into a shared space without intermediate translation. Organizations can therefore choose the appropriate approach based on their balance between compute costs, latency requirements, and accuracy targets.
Teams developing multilingual systems should use the XNLI corpus to evaluate and benchmark multilingual representations. For near-term deployments where accuracy is critical and infrastructure permits, test-time machine translation remains the recommended option. For cost-sensitive, low-latency environments, teams should invest in multilingual sentence encoders. Future development should focus on joint encoder training, shared parameters, and gathering more parallel data for lower-resource languages to close the performance gap with translation pipelines.
A primary limitation of the study is that XNLI was created by translating English sentences, meaning it does not fully capture natural cultural nuances, idioms, or stylistic variations found in native text. Additionally, performance in lower-resource languages remains constrained by the limited availability of parallel training data. Readers can have high confidence in the relative comparisons between model architectures, but should exercise caution when assuming these models will automatically generalize to culturally divergent, colloquial native text.
- Paper: A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference, Adina Williams et al. (2016). This work introduces the Multi-Genre Natural Language Inference (MultiNLI) corpus, whose development and test sets serve as the direct foundation extended into 15 languages to construct XNLI.
- Paper: A large annotated corpus for learning natural language inference, Samuel R. Bowman et al. (2015). This foundational paper introduces the SNLI corpus and establishes the modern natural language inference task framing that MultiNLI and subsequently XNLI rely on.
- Paper: Supervised Learning of Universal Sentence Representations from Natural Language Inference Data, Alexis Conneau et al. (2017). This study demonstrates that supervised NLI training yields powerful universal sentence representations, providing the conceptual motivation for using NLI as the benchmark task for evaluating cross-lingual sentence encoders.
- Paper: Word Translation Without Parallel Data, Alexis Conneau et al. (2017). This paper establishes unsupervised cross-lingual alignment techniques for embedding spaces that directly inform cross-lingual sentence representation baselines.
- Paper: Google’s Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation, Melvin Johnson et al. (2016). This paper presents Google's multilingual neural machine translation system, demonstrating zero-shot cross-lingual transfer principles that underlie multilingual representation learning.
- Paper: Improving Neural Machine Translation Models with Monolingual Data, Rico Sennrich et al. (2016). This work introduces back-translation for machine translation, an essential technique leveraged in translation-based multilingual baselines and cross-lingual data pipelines.
- Paper: Cross-lingual Language Model Pretraining, Guillaume Lample et al. (2019). This paper introduces cross-lingual language model pretraining (XLM) and utilizes XNLI as the primary benchmark to prove the efficacy of Translation Language Modeling.
- Paper: Unsupervised Cross-lingual Representation Learning at Scale, Alexis Conneau et al. (2019). This work scales unsupervised multilingual pretraining to 100 languages with XLM-R and evaluates cross-lingual transfer performance directly on the XNLI benchmark.
- Paper: How Multilingual is Multilingual BERT?, Telmo Pires et al. (2019). This paper systematically probes the zero-shot cross-lingual mechanisms of Multilingual BERT using cross-lingual transfer frameworks popularized by benchmarks like XNLI.
- Paper: mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer, Linting Xue et al. (2020). This paper introduces mT5 and evaluates its massively multilingual text-to-text representations across cross-lingual understanding tasks such as XNLI within XTREME.
- Paper: Multilingual Denoising Pre-training for Neural Machine Translation, Yinhan Liu et al. (2020). This work develops mBART for multilingual sequence-to-sequence pre-training, extending multilingual evaluation methods onto downstream cross-lingual transfer tasks.
