XNLI: Evaluating Cross-lingual Sentence Representations

Alexis ConneauGuillaume LampleRuty RinottAdina WilliamsSamuel R. BowmanHolger SchwenkVeselin Stoyanov

article2018EMNLP1,671 citations

Introduces the Cross-lingual Natural Language Inference (XNLI) benchmark across 15 diverse languages, establishing a standard evaluation suite and competitive baselines for cross-lingual sentence representations and zero-shot transfer.

Listen

Modern natural language processing systems rely heavily on large volumes of annotated data to learn complex tasks. However, these datasets are typically available only in English, leaving international applications constrained because manually labeling data across dozens of languages is prohibitively expensive. To address this challenge, researchers have focused on cross-lingual language understanding, where a model trains primarily on English data and evaluates on other target languages. Despite growing interest, progress has been constrained by the absence of a standardized, large-scale sentence-level evaluation benchmark.

The article introduces and evaluates the Cross-lingual Natural Language Inference (XNLI) corpus, a standardized evaluation benchmark designed to measure cross-lingual sentence understanding across 15 diverse languages. The article demonstrates how well different cross-lingual approaches—ranging from machine translation pipelines to shared multilingual sentence encoders—can transfer knowledge from English to non-English tasks without requiring target-language training data.

To build this benchmark, the authors collected 7,500 new English sentence pairs across ten genres using established crowdsourcing procedures, then hired professional translators to translate them into 14 additional languages, including lower-resource languages such as Swahili and Urdu. This process yielded a fully aligned evaluation suite of 112,500 annotated pairs. The authors then evaluated baseline architectures: translation-based methods that translate data either during training or at test time, and multilingual sentence encoders that align target-language sentence representations directly to English using parallel corpora and a specialized alignment loss function.

The evaluation revealed several key findings. First, the highest overall performance across all languages was achieved by translating incoming foreign test sentences directly into English at inference time, reaching up to 70.7% accuracy in Spanish compared to the English baseline of 73.7%. Second, machine translation quality heavily dictated task accuracy; languages with high translation quality regularly exceeded 70% accuracy, whereas low-resource languages like Swahili and Urdu scored lower, between 58% and 62%. Third, multilingual sentence encoders trained directly with parallel text yielded competitive results (68.9% in Greek and 68.7% in Spanish) without requiring a runtime translation system, though they trailed test-time translation by up to 6 percentage points in lower-resource settings. Finally, bidirectional neural sentence encoders that pooled representations across all hidden states consistently outperformed pretrained word-averaging techniques across every tested language.

These results present clear operational trade-offs for organizations deploying multilingual systems. While test-time machine translation yields the highest accuracy, it is computationally expensive and introduces runtime latency. In contrast, multilingual sentence encoders significantly reduce computational overhead at inference time because they map foreign text directly into a shared space without intermediate translation. Organizations can therefore choose the appropriate approach based on their balance between compute costs, latency requirements, and accuracy targets.

Teams developing multilingual systems should use the XNLI corpus to evaluate and benchmark multilingual representations. For near-term deployments where accuracy is critical and infrastructure permits, test-time machine translation remains the recommended option. For cost-sensitive, low-latency environments, teams should invest in multilingual sentence encoders. Future development should focus on joint encoder training, shared parameters, and gathering more parallel data for lower-resource languages to close the performance gap with translation pipelines.

A primary limitation of the study is that XNLI was created by translating English sentences, meaning it does not fully capture natural cultural nuances, idioms, or stylistic variations found in native text. Additionally, performance in lower-resource languages remains constrained by the limited availability of parallel training data. Readers can have high confidence in the relative comparisons between model architectures, but should exercise caution when assuming these models will automatically generalize to culturally divergent, colloquial native text.

Cover for XNLI: Evaluating Cross-lingual Sentence Representations

Abstract

State-of-the-art natural language processing systems rely on supervision in the form of annotated data to learn competent models. These models are generally trained on data in a single language (usually English), and cannot be directly used beyond that language. Since collecting data in every language is not realistic, there has been a growing interest in cross-lingual language understanding (XLU) and low-resource cross-language transfer. In this work, we construct an evaluation set for XLU by extending the development and test sets of the Multi-Genre Natural Language Inference Corpus (MultiNLI) to 15 languages, including low-resource languages such as Swahili and Urdu. We hope that our dataset, dubbed XNLI, will catalyze research in cross-lingual sentence understanding by providing an informative standard evaluation task. In addition, we provide several baselines for multilingual sentence understanding, including two based on machine translation systems, and two that use parallel data to train aligned multilingual bag-of-words and LSTM encoders. We find that XNLI represents a practical and challenging evaluation suite, and that directly translating the test data yields the best performance among available baselines.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 The XNLI Corpus
  • 3.1 Data Collection
  • 3.2 The Resulting Corpus
  • 4 Cross-Lingual NLI
  • 4.1 Translation-Based Approaches
  • 4.2 Multilingual Sentence Encoders
  • 4.2.1 Aligning Word Embeddings
  • 4.2.2 Universal Multilingual Sentence Embeddings
  • 4.2.3 Aligning Sentence Embeddings
  • 5 Experiments and Results
  • 5.1 Training details
  • 5.2 Parallel Datasets
  • 5.3 Analysis
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Cross-Lingual Natural Language Inference (XNLI) Dataset

    definition

    The Cross-lingual Natural Language Inference (XNLI) dataset is a multilingual benchmark for evaluating cross-lingual sentence understanding across 15 languages: English (en), French (fr), Spanish (es), German (de), Greek (el), Bulgarian (bg), Russian (ru), Turkish (tr), Arabic (ar), Vietnamese (vi), Thai (th), Chinese (zh), Hindi (hi), Swahili (sw), and Urdu (ur).

    The dataset extends the Multi-Genre Natural Language Inference (MultiNLI) corpus. It contains 7,500 human-annotated English sentence pairs—sampled equally (750 pairs each) from 10 genres: Face-To-Face, Telephone, Government, 9/11, Letters, Oxford University Press (OUP), Slate, Verbatim, Government, and Fiction (Captain Blood). Each premise is paired with three crowdsourced hypotheses corresponding to the three NLI labels: entailment, neutral, and contradiction.

    To construct the cross-lingual corpus, professional translators translated premises and hypotheses independently into the 14 target languages, yielding 2,500 development pairs and 5,000 test pairs per language (112,500 annotated pairs in total). Gold labels from the English source are directly mapped to the translations. In quality validation, 93% of sentence pairs achieved a three-vote consensus among five annotators (the remaining 7% lack consensus and are marked with label -). Bilingual re-annotation of 100 development examples achieved 85% label agreement with English consensus on English data and 83% agreement on translated French data.

  2. Knowl 2 — Sentence Embedding Alignment Loss for Cross-Lingual Transfer

    equation

    To align sentence encoders across different languages into a shared embedding space without generating translations, the target language sentence encoder is trained to map target sentences to the fixed representation of their English translations using the alignment loss:

    Lalign(x,y)=∥x−y∥2−λ(∥xc−y∥2+∥x−yc∥2)\mathcal{L}_{\text{align}}(x, y) = \|x - y\|_2 - \lambda \left(\|x_c - y\|_2 + \|x - y_c\|_2\right)

    where x∈Rdx \in \mathbb{R}^d is the source (English) sentence embedding, y∈Rdy \in \mathbb{R}^d is the target language sentence embedding for a parallel sentence pair, xc∈Rdx_c \in \mathbb{R}^d and yc∈Rdy_c \in \mathbb{R}^d are contrastive (negative) sentence embeddings sampled randomly from other sentences in the batch, ∥⋅∥2\|\cdot\|_2 denotes the Euclidean distance, and λ≥0\lambda \ge 0 is a scalar hyperparameter weighting the negative sampling penalty (optimal performance is obtained at λ=0.25\lambda = 0.25).

    During training, gradients are backpropagated exclusively through the target language encoder while keeping the pretrained source language (English) encoder fixed, ensuring all target encoders project into the source coordinate space.

  3. Knowl 3 — Failure of Margin-Based Ranking Loss for Shared-Classifier Cross-Lingual Transfer

    theoretical result

    In cross-lingual transfer where a frozen classifier trained on source language (English) embeddings is applied directly to aligned target language sentence embeddings, standard margin-based ranking losses fail to produce competitive classifiers. The margin ranking loss:

    Lrank(x,y)=max⁡(0,α−∥x−yc∥2+∥x−y∥2)+max⁡(0,α−∥xc−y∥2+∥x−y∥2)\mathcal{L}_{\text{rank}}(x, y) = \max(0, \alpha - \|x - y_c\|_2 + \|x - y\|_2) + \max(0, \alpha - \|x_c - y\|_2 + \|x - y\|_2)

    where α>0\alpha > 0 is a margin parameter, only enforces relative distances—ensuring true translation pairs (x,y)(x, y) are closer to each other than contrastive negative pairs (xc,y)(x_c, y) and (x,yc)(x, y_c).

    Because Lrank\mathcal{L}_{\text{rank}} does not constrain the absolute Euclidean distance ∥x−y∥2\|x - y\|_2 between translations, target embeddings can shift arbitrarily in metric space relative to the fixed decision boundaries of the English classifier. Conversely, the metric alignment loss Lalign(x,y)=∥x−y∥2−λ(∥xc−y∥2+∥x−yc∥2)\mathcal{L}_{\text{align}}(x, y) = \|x - y\|_2 - \lambda(\|x_c - y\|_2 + \|x - y_c\|_2) explicitly minimizes absolute Euclidean distance, forcing target sentence vectors into the exact geometric region where the English classifier makes accurate predictions.

  4. Knowl 4 — Cross-Lingual BiLSTM Sentence Encoders (X-BiLSTM)

    model/method

    The X-BiLSTM architecture aligns bidirectional LSTM encoders across multiple languages to perform zero-shot cross-lingual natural language inference:

    1. Encoder Architecture: A BiLSTM with 512 hidden units per direction, initialized with 300-dimensional FastText word embeddings aligned cross-lingually via the MUSE framework. Two feature extraction pooling methods are defined: BiLSTM-last (concatenation of initial and final hidden states) and BiLSTM-max (element-wise max-pooling over all hidden states).

    2. Monolingual English Training: An English BiLSTM encoder and an NLI classifier are trained on MultiNLI training data. Given premise vector u∈Rdu \in \mathbb{R}^d and hypothesis vector v∈Rdv \in \mathbb{R}^d, the classifier receives the feature vector [u,v,∣u−v∣,u∗v]∈R4d[u, v, |u - v|, u * v] \in \mathbb{R}^{4d}, where ∗* denotes element-wise multiplication. The classifier is a feed-forward network with one hidden layer of 128 hidden units, regularized with dropout (rate 0.1) and trained using Adam.

    3. Cross-Lingual Alignment: The English BiLSTM and classifier are frozen. For each target language, a separate BiLSTM with the identical architecture is trained on parallel sentence data using Lalign\mathcal{L}_{\text{align}} (with λ=0.25\lambda = 0.25). The target language word lookup table is fine-tuned while backpropagating only through the target encoder.

  5. Knowl 5 — Cross-Lingual Continuous Bag-of-Words Encoder (X-CBOW)

    model/method

    The X-CBOW method provides a lightweight baseline for universal cross-lingual sentence representations by aligning continuous bag-of-words (CBOW) embeddings across languages:

    1. Sentence embeddings are computed by averaging 300-dimensional FastText word vectors over all tokens in the sentence.
    2. The English word embedding space is held fixed.
    3. For each target language, word vector lookup tables are fine-tuned by minimizing the sentence alignment loss Lalign\mathcal{L}_{\text{align}} over parallel sentence pairs, ensuring the average word vector of a target sentence matches the average word vector of its English translation.
    4. An NLI classification network (one hidden layer with 128 units operating on [u,v,∣u−v∣,u∗v][u, v, |u-v|, u*v]) is trained on frozen English CBOW representations and evaluated directly on target language CBOW sentence embeddings without machine translation at inference time.
  6. Knowl 6 — Machine Translation Baselines for Cross-Lingual NLI

    model/method

    Cross-lingual natural language inference is benchmarked against two machine-translation-based paradigms:

    1. TRANSLATE TRAIN: The English MultiNLI training corpus is translated offline into each target language using neural machine translation systems. A separate monolingual BiLSTM encoder and classifier are trained independently from scratch on the translated training data for each target language.

    2. TRANSLATE TEST: A single monolingual BiLSTM encoder and classifier are trained exclusively on the original English MultiNLI training set. At evaluation time, foreign-language premise and hypothesis pairs from the XNLI test set are translated into English using neural machine translation systems and classified by the English model.

  7. Knowl 7 — Cross-Lingual NLI Benchmark Evaluation Results

    data/table

    Test accuracy (%) on the 15 languages of the XNLI benchmark across machine translation baselines and cross-lingual sentence embedding models:

    Method en fr es de el bg ru tr ar vi th zh hi sw ur
    Machine translation baselines (TRANSLATE TRAIN)
    BiLSTM-last 71.0 66.7 67.0 65.7 65.3 65.6 65.1 61.9 63.9 63.1 61.3 65.7 61.3 55.2 55.2
    BiLSTM-max 73.7 68.3 68.8 66.5 66.4 67.4 66.5 64.5 65.8 66.0 62.8 67.0 62.1 58.2 56.6
    Machine translation baselines (TRANSLATE TEST)
    BiLSTM-last 71.0 68.3 68.7 66.9 67.3 68.1 66.2 64.9 65.8 64.3 63.2 66.5 61.8 60.1 58.1
    BiLSTM-max 73.7 70.4 70.7 68.7 69.1 70.4 67.8 66.3 66.8 66.5 64.4 68.3 64.2 61.8 59.3
    Multilingual sentence encoders (in-domain alignment)
    X-BiLSTM-last 71.0 65.2 67.8 66.6 66.3 65.7 63.7 64.2 62.7 65.6 62.7 63.7 62.8 54.1 56.4
    X-BiLSTM-max 73.7 67.7 68.7 67.7 68.9 67.9 65.4 64.2 64.8 66.4 64.1 65.8 64.1 55.7 58.4
    Pretrained multilingual sentence encoders (transfer learning)
    X-CBOW 64.5 60.3 60.7 61.0 60.5 60.4 57.8 58.7 57.5 58.8 56.9 58.8 56.3 50.4 52.2

    Key takeaways:

    1. TRANSLATE TEST with BiLSTM-max achieves the highest accuracy across every non-English language (ranging from 59.3% in Urdu to 70.7% in Spanish).
    2. TRANSLATE TEST consistently outperforms TRANSLATE TRAIN by 1.3% to 3.6% accuracy across all languages.
    3. Max-pooling (BiLSTM-max) outperforms last-state pooling (BiLSTM-last) by ~2.7% on English, and this performance advantage transfers consistently across all 14 non-English languages.
    4. Aligned sentence encoders (X-BiLSTM-max) achieve competitive performance with TRANSLATE TRAIN (e.g., 68.7% vs 68.8% in Spanish; 68.9% vs 66.4% in Greek) without requiring machine translation at inference time, but lag behind on low-resource languages (Swahili, Urdu) where parallel alignment data is scarce.
  8. Knowl 8 — Correlation Between Machine Translation Quality and Cross-Lingual NLI Accuracy

    empirical result

    Cross-lingual NLI classification performance correlates directly with translation quality as quantified by BLEU scores:

    • Languages with the highest XX-to-English translation BLEU scores—Spanish (45.8 BLEU), Greek (42.1 BLEU), French (41.2 BLEU), and Bulgarian (38.7 BLEU)—achieve the highest TRANSLATE TEST NLI accuracies (70.7%, 69.1%, 70.4%, and 70.4% respectively using BiLSTM-max).
    • Lower-resource languages with lower BLEU scores—Swahili (21.3 BLEU) and Urdu (24.4 BLEU)—yield lower NLI accuracy (61.8% and 59.3% respectively under TRANSLATE TEST BiLSTM-max).
    • Monolingual English NLI performance (73.7%) exceeds the best cross-lingual translation transfer performance (70.7% on Spanish) by 3.0 percentage points, reflecting translation errors, stylistic shifts, and machine translation artifacts.
  9. Knowl 9 — Hyperparameter Sensitivity in Cross-Lingual BiLSTM Sentence Alignment

    empirical result

    Validation accuracy (%) for X-BiLSTM-max under variations in negative sampling weight λ\lambda and target embedding fine-tuning ft∈{0,1}ft \in \{0, 1\}:

    Configuration French (fr) Russian (ru) Chinese (zh)
    ft=1,λ=0.25ft = 1, \lambda = 0.25 (default) 68.9 66.4 67.9
    ft=1,λ=0.0ft = 1, \lambda = 0.0 (no negatives) 67.8 66.2 66.3
    ft=1,λ=0.5ft = 1, \lambda = 0.5 64.5 61.3 63.7
    ft=0,λ=0.25ft = 0, \lambda = 0.25 68.5 66.3 67.7

    Key observations:

    1. Negative Term Weight λ\lambda: Setting λ=0.25\lambda = 0.25 improves validation accuracy over omitting contrastive negative terms entirely (λ=0.0\lambda = 0.0), with gains up to +1.6% in Chinese (67.9% vs 66.3%). Setting λ=0.5\lambda = 0.5 over-penalizes negatives and causes severe performance degradation (-3.5% to -5.1%).
    2. Embedding Fine-Tuning ftft: Freezing target language word embeddings (ft=0ft = 0) results in minimal performance loss (0.1% to 0.4%), demonstrating that the BiLSTM recurrent parameters alone are sufficient to align cross-lingual sentence representations into the English embedding space.
    3. Convergence Dynamics: Tracking Lalign\mathcal{L}_{\text{align}} on parallel development sets across training epochs exhibits a strong inverse correlation with XNLI development accuracy: lower alignment loss directly predicts higher cross-lingual classification accuracy.
  10. Knowl 10 — Limitations of Translation-Derived Multilingual Benchmarks

    limitation

    The XNLI corpus was created by translating English sentences into target languages rather than generating native premises and hypotheses within each linguistic and cultural context. While professional translation preserves logical semantic relations (entailment, neutrality, contradiction) across languages with minimal label corruption, it fails to reflect cultural differences, style shifts, and language-specific domain adaptation challenges inherent to native multilingual text.

Coverage note — None. All major contributed dataset specifications, alignment methods, model architectures, baseline translation strategies, empirical benchmarks, ablation results, and stated limitations are covered.

References

  1. 1.Željko Agić and Natalie Schluter. 2018. Baselines and test data for cross-lingual inference. LREC.
  2. 2.Waleed Ammar, George Mulcaire, Yulia Tsvetkov, Guillaume Lample, Chris Dyer, and Noah A Smith. 2016. Massively multilingual word embeddings. arXiv preprint arXiv:1602.01925.
  3. 3.Kunchukuttan Anoop, Mehta Pratik, and Bhattacharyya Pushpak. 2018. The iit bombay english-hindi parallel corpus. In LREC.
  4. 4.Sanjeev Arora, Yingyu Liang, and Tengyu Ma. 2017. A simple but tough-to-beat baseline for sentence embeddings. ICLR.
  5. 5.Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2017. Learning bilingual word embeddings with (almost) no bilingual data. In ACL, pages 451–462.
  6. 6.Mikel Artetxe, Gorka Labaka, Eneko Agirre, and Kyunghyun Cho. 2018. Unsupervised neural machine translation. In ICLR.
  7. 7.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In EMNLP.
  8. 8.Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia. 2017. Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation. In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), pages 1–14.
  9. 9.Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, et al. 2018. Universal sentence encoder. arXiv preprint arXiv:1803.11175.
  10. 10.Sarath Chandar, Mitesh M. Khapra, Balaraman Ravindran, Vikas Raykar, and Amrita Saha. 2013. Multilingual deep learning. In NIPS, Workshop Track.
  11. 11.Pi-Chuan Chang, Michel Galley, and Christopher D Manning. 2008. Optimizing chinese word segmentation for machine translation performance. In Proceedings of the third workshop on statistical machine translation, pages 224–232.
  12. 12.Ronan Collobert and Jason Weston. 2008. A unified architecture for natural language processing: Deep neural networks with multitask learning. In ICML, pages 160–167. ACM.
  13. 13.Alexis Conneau and Douwe Kiela. 2018. Senteval: An evaluation toolkit for universal sentence representations. LREC.
  14. 14.Alexis Conneau, Douwe Kiela, Holger Schwenk, Loïc Barrault, and Antoine Bordes. 2017. Supervised learning of universal sentence representations from natural language inference data. In EMNLP, pages 670–680.
  15. 15.Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018a. What you can cram into a single vector: Probing sentence embeddings for linguistic properties. In ACL.
  16. 16.Alexis Conneau, Guillaume Lample, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jegou. 2018b. Word translation without parallel data. In ICLR.
  17. 17.Cristina España-Bonet, Ádám Csaba Varga, Alberto Barrón-Cedeño, and Josef van Genabith. 2017. An empirical analysis of nmt-derived interlingual embeddings and their use in parallel sentence identification. IEEE Journal of Selected Topics in Signal Processing, pages 1340–1348.
  18. 18.Manaal Faruqui and Chris Dyer. 2014. Improving vector space word representations using monolingual correlation. In EACL.
  19. 19.Yichen Gong, Heng Luo, and Jian Zhang. 2018. Natural language inference over interaction space. ICLR.
  20. 20.S. Gouews, Y. Bengio, and G. Corrado. 2014. Bilbowa: Fast bilingual distributed representations without word alignments. In ICML.
  21. 21.Edouard Grave, Piotr Bojanowski, Prakhar Gupta, Armand Joulin, and Tomas Mikolov. 2018. Learning word vectors for 157 languages. LREC.
  22. 22.Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel R Bowman, and Noah A Smith. 2018. Annotation artifacts in natural language inference data. NAACL.
  23. 23.Karl Moritz Hermann and Phil Blunsom. 2014. Multilingual models for compositional distributed semantics. In ACL, pages 58–68.
  24. 24.Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short-term memory. Neural computation, 9(8):1735–1780.
  25. 25.Jeremy Howard and Sebastian Ruder. 2018. Fine-tuned language models for text classification. In ACL.
  26. 26.Melvin Johnson, Mike Schuster, Quoc V Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda Viégas, Martin Wattenberg, Greg Corrado, et al. 2016. Google’s multilingual neural machine translation system: enabling zero-shot translation. arXiv preprint arXiv:1611.04558.
  27. 27.Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. In ICLR.
  28. 28.Ryan Kiros, Yukun Zhu, Ruslan R Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Skip-thought vectors. In NIPS, pages 3294–3302.
  29. 29.A. Klementiev, I. Titov, and B. Bhattarai. 2012. Inducing crosslingual distributed representations of words. In COLING.
  30. 30.T. Kociský, K.M. Hermann, and P. Blunsom. 2014. Learning bilingual word representations by marginalizing alignments. In ACL, pages 224–229.
  31. 31.Philipp Koehn. 2005. Europarl: A parallel corpus for statistical machine translation. In MT summit, volume 5, pages 79–86.
  32. 32.Guillaume Lample, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2018a. Unsupervised machine translation using monolingual corpora only. In ICLR.
  33. 33.Guillaume Lample, Myle Ott, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2018b. Phrase-based & neural unsupervised machine translation. In EMNLP.
  34. 34.Quoc Le and Tomas Mikolov. 2014. Distributed representations of sentences and documents. In ICML, pages 1188–1196.
  35. 35.Yashar Mehdad, Matteo Negri, and Marcello Federico. 2011. Using bilingual parallel corpora for crosslingual textual entailment. In ACL, pages 1336–1345.
  36. 36.Tomas Mikolov, Quoc V Le, and Ilya Sutskever. 2013a. Exploiting similarities among languages for machine translation. arXiv preprint arXiv:1309.4168.
  37. 37.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013b. Distributed representations of words and phrases and their compositionality. In NIPS, pages 3111–3119.
  38. 38.Saif M. Mohammad, Mohammad Salameh, and Svetlana Kiritchenko. 2016. How translation alters sentiment. J. Artif. Int. Res., 55(1):95–130.
  39. 39.Matteo Negri, Luisa Bentivogli, Yashar Mehdad, Danilo Giampiccolo, and Alessandro Marchetti. 2011. Divide and conquer: Crowdsourcing the creation of cross-lingual textual entailment corpora. In EMNLP, pages 670–679.
  40. 40.Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. In NAACL, volume 1, pages 2227–2237.
  41. 41.Hieu Pham, Minh-Thang Luong, and Christopher D. Manning. 2015. Learning distributed representations for multilingual text sequences. In Workshop on Vector Space Modeling for NLP.
  42. 42.Lison Pierre and Tiedemann Jörg. 2016. Pierre lison and jörg tiedemann, 2016, opensubtitles2016: Extracting large parallel corpora from movie and tv subtitles. In LREC.
  43. 43.Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018. Hypothesis only baselines in natural language inference. In NAACL.
  44. 44.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding by generative pre-training.
  45. 45.Tim Rocktäschel, Edward Grefenstette, Karl Moritz Hermann, Tomáš Kočiský, and Phil Blunsom. 2016. Reasoning about entailment with neural attention. ICLR.
  46. 46.Rafael Sabatini. 1922. Captain Blood. Houghton Mifflin Company.
  47. 47.Holger Schwenk and Xian Li. 2018. A corpus for multilingual document classification in eight languages. In LREC, pages 3548–3551.
  48. 48.Holger Schwenk, Ke Tran, Orhan Firat, and Matthijs Douze. 2017. Learning joint multilingual sentence representations with neural machine translation. In ACL workshop, Repl4NLP.
  49. 49.Laura Smith, Salvatore Giorgi, Rishi Solanki, Johannes C. Eichstaedt, H. Andrew Schwartz, Muhammad Abdul-Mageed, Anneke Buffone, and Lyle H. Ungar. 2016. Does ’well-being’ translate on twitter? In EMNLP, pages 2042–2047.
  50. 50.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014. Dropout: A simple way to prevent neural networks from overfitting. The Journal of Machine Learning Research, 15(1):1929–1958.
  51. 51.Sandeep Subramanian, Adam Trischler, Yoshua Bengio, and Christopher J Pal. 2018. Learning general purpose distributed sentence representations via large scale multi-task learning. In ICLR.
  52. 52.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014. Sequence to sequence learning with neural networks. In NIPS, pages 3104–3112.
  53. 53.Jörg Tiedemann. 2012. Parallel data, tools and interfaces in opus. In LREC, Istanbul, Turkey. European Language Resources Association (ELRA).
  54. 54.Masatoshi Tsuchiya. 2018. Performance impact caused by hidden bias of training data for recognizing textual entailment. In LREC.
  55. 55.Alex Wang, Amapreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461.
  56. 56.Jason Weston, Samy Bengio, and Nicolas Usunier. 2011. Wsabie: Scaling up to large vocabulary image annotation. In IJCAI.
  57. 57.John Wieting, Mohit Bansal, Kevin Gimpel, and Karen Livescu. 2016. Towards universal paraphrastic sentence embeddings. ICLR.
  58. 58.Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2017. A broad-coverage challenge corpus for sentence understanding through inference. In NAACL.
  59. 59.Min Xiao and Yuhong Guo. 2014. Distributed word representation learning for cross-lingual dependency parsing. In CoNLL, pages 119–129.
  60. 60.Chao Xing, Dong Wang, Chao Liu, and Yiye Lin. 2015. Normalized word embedding and orthogonal transform for bilingual word translation. NAACL.
  61. 61.Yuan Zhang, David Gaddy, Regina Barzilay, and Tommi Jaakkola. 2016. Ten pairs to tag–multilingual pos tagging via coarse mapping between embeddings. In NAACL, pages 1307–1317.
  62. 62.Xinjie Zhou, Xiaojun Wan, and Jianguo Xiao. 2016. Cross-lingual sentiment classification with bilingual document representation learning. In ACL.
  63. 63.Michal Ziemski, Marcin Junczys-Dowmunt, and Bruno Pouliquen. 2016. The united nations parallel corpus v1. 0. In LREC.

Citation

MLA
Conneau, A., et al. “XNLI: Evaluating Cross-lingual Sentence Representations”. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 2018, pp. 2475–85, https://doi.org/10.18653/v1/D18-1269.
APA
Conneau, A., Rinott, R., Lample, G., Williams, A., Bowman, S., Schwenk, H., & Stoyanov, V. (2018). XNLI: Evaluating Cross-lingual Sentence Representations. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 2475–2485. https://doi.org/10.18653/v1/D18-1269
Chicago
Conneau, A., R. Rinott, G. Lample, et al. 2018. “XNLI: Evaluating Cross-lingual Sentence Representations”. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 2475–85. https://doi.org/10.18653/v1/D18-1269.
Harvard
Conneau, A. et al. (2018) “XNLI: Evaluating Cross-lingual Sentence Representations”, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 2475–2485. Available at: https://doi.org/10.18653/v1/D18-1269.
Vancouver
1. Conneau A, Rinott R, Lample G, Williams A, Bowman S, Schwenk H, Stoyanov V (2018) XNLI: Evaluating Cross-lingual Sentence Representations. In: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 2475–2485

BibTeX

@inproceedings{Conneau_2018, title={XNLI: Evaluating Cross-lingual Sentence Representations}, url={http://dx.doi.org/10.18653/v1/D18-1269}, DOI={10.18653/v1/d18-1269}, booktitle={Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing}, publisher={Association for Computational Linguistics}, author={Conneau, Alexis and Rinott, Ruty and Lample, Guillaume and Williams, Adina and Bowman, Samuel and Schwenk, Holger and Stoyanov, Veselin}, year={2018}, pages={2475–2485} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/