Annotation Artifacts in Natural Language Inference Data

Suchin GururanganSwabha SwayamdiptaOmer LevyRoy SchwartzSamuel R. BowmanNoah A. Smith

article2018NAACL1,339 citations

Demonstrates that widely used natural language inference benchmarks contain crowdsourced artifacts that allow models to predict labels from the hypothesis alone without observing the premise, proving that machine reasoning performance has been substantially overestimated.

Listen

Natural language inference is a core benchmark for artificial intelligence, testing whether a model can determine if a given premise sentence logically entails, contradicts, or is neutral toward a hypothesis sentence. However, the standard crowdsourcing process used to generate these datasets introduces unintended patterns and shortcuts, known as annotation artifacts. These artifacts allow machine learning models to appear capable of complex reasoning while merely exploiting surface-level statistical cues.

The main objective of the article is to demonstrate the widespread existence of these annotation artifacts in two prominent natural language inference datasets, identify the human heuristics that cause them, and evaluate how heavily state-of-the-art models rely on these shortcuts rather than genuine logical reasoning.

To evaluate this, the authors applied a simple text classifier that predicted labels using only the generated hypothesis sentences, without ever seeing the premise sentences. They analyzed lexical associations and sentence lengths across the datasets to diagnose human writing habits. Finally, they split test datasets into "Easy" examples (which the hypothesis-only model solved) and "Hard" examples (which it failed), re-evaluating three leading natural language inference systems across these segments.

The investigation produced four central findings. First, a simple classifier ignoring the premise entirely achieved 67.0% accuracy on the Stanford Natural Language Inference dataset and 52.3% to 53.9% on the Multi-Genre Natural Language Inference dataset, far exceeding the roughly 35% baseline of random guessing. Second, crowd workers consistently applied specific heuristics: negation words (such as "no," "never," and "nobody") strongly correlated with contradictions, purpose clauses and modifiers correlated with neutral statements, and general hypernyms (such as "animal" for "dog") correlated with entailments. Third, sentence length served as a distinct cue; neutral hypotheses were systematically longer, while 60% of entailments were seven words or fewer. Fourth, when top-performing models were evaluated on the "Hard" test subset, their performance dropped sharply—falling by 12 to 17 percentage points compared to their standard benchmark scores.

These findings indicate that the perceived reasoning capabilities of current AI models have been substantially overstated. Rather than performing true textual reasoning, these systems are largely exploiting subtle artifacts present in crowdsourced benchmarks. Relying on such benchmark scores presents significant risks for downstream applications—such as automated question answering and summarization—where systems may fail unexpectedly on real-world inputs that lack these specific statistical shortcuts.

To address this challenge, the article recommends that AI practitioners evaluate existing and future models on the newly partitioned "Hard" datasets alongside standard benchmarks. Simply filtering out easy examples from existing data is insufficient, as it introduces new statistical distortions. Instead, dataset creators should improve crowdsourcing protocols by diversifying prompt instructions and utilizing automated or adversarial systems during data collection to balance heuristic patterns across all classes.

The findings are supported with high statistical confidence across multiple standard benchmarks and model architectures. However, the authors caution that identifying and balancing every heuristic during data collection remains difficult. Stakeholders should recognize that addressing data artifacts is an ongoing challenge, and current natural language inference remains an open, unsolved problem.

arXiv: 1803.02324
Cover for Annotation Artifacts in Natural Language Inference Data

Abstract

Large-scale datasets for natural language inference are created by presenting crowd workers with a sentence (premise), and asking them to generate three new sentences (hypotheses) that it entails, contradicts, or is logically neutral with respect to. We show that, in a significant portion of such data, this protocol leaves clues that make it possible to identify the label by looking only at the hypothesis, without observing the premise. Specifically, we show that a simple text categorization model can correctly classify the hypothesis alone in about 67% of SNLI (Bowman et. al, 2015) and 53% of MultiNLI (Williams et. al, 2017). Our analysis reveals that specific linguistic phenomena such as negation and vagueness are highly correlated with certain inference classes. Our findings suggest that the success of natural language inference models to date has been overestimated, and that the task remains a hard open problem.

Table of Contents

  • 1 Introduction
  • 2 Annotation Artifacts are Common
  • 3 Characteristics of Annotation Artifacts
  • 3.1 Lexical Choice
  • 3.2 Sentence Length
  • 4 Re-evaluating NLI Models
  • 5 Discussion
  • References

Knowls

  1. Knowl 1 — Performance Degradation of NLI Models on Hypothesis-Hard vs. Easy Test Splits

    data/table

    When state-of-the-art Natural Language Inference (NLI) models are evaluated on test instances where a premise-oblivious hypothesis-only classifier fails (the "Hard" test set), their classification accuracy drops dramatically compared to their accuracy on the full test sets or on instances where the hypothesis-only classifier succeeds (the "Easy" test set). This demonstrates that a large portion of model accuracy on standard NLI benchmarks comes from exploiting hypothesis artifacts rather than joint semantic reasoning over premise-hypothesis pairs.

    Model SNLI MultiNLI Matched MultiNLI Mismatched
    Full Hard Easy Full Hard Easy Full Hard Easy
    DAM 84.7 69.4 92.4 72.0 55.8 85.3 72.1 56.2 85.7
    ESIM 85.8 71.3 92.6 74.1 59.3 86.2 73.1 58.9 85.2
    DIIN 86.5 72.7 93.4 77.0 64.1 87.6 76.5 64.4 86.8

    The evaluated models are:

    • DAM: Decomposable Attention Model
    • ESIM: Enhanced Sequential Inference Model
    • DIIN: Densely Interactive Inference Network

    All models were trained on the original full training sets of SNLI or MultiNLI alone and evaluated across the Full, Hard, and Easy partitions.

  2. Knowl 2 — Premise-Oblivious Hypothesis Classification Accuracy on SNLI and MultiNLI

    data/table

    A fastText bag-of-words and bigrams classifier trained exclusively on hypothesis sentences—without observing the premise sentences—achieves classification accuracy well above the majority class baseline on both SNLI and MultiNLI (matched and mismatched evaluation sets).

    Model SNLI MultiNLI
    Matched Mismatched
    majority class 34.3 35.4 35.2
    fastText 67.0 53.9 52.3

    For MultiNLI, the fastText model additionally utilizes character 4-grams and filters out words appearing fewer than 10 times in the training data. The ability to correctly classify over 67%67\% of SNLI and over 52%52\% of MultiNLI hypotheses without seeing the premise establishes that the crowd-authored hypotheses contain strong statistical artifacts correlating directly with the inference labels.

  3. Knowl 3 — Partitioning NLI Benchmarks into Easy and Hard Evaluation Subsets

    model/method

    To measure the extent to which standard NLI models rely on hypothesis-only artifacts, test sets are partitioned into two disjoint subsets based on the performance of a premise-oblivious fastText hypothesis classifier:

    1. Easy Test Set: The subset of test examples whose inference label is correctly predicted by the hypothesis-only classifier.
    2. Hard Test Set: The subset of test examples that the hypothesis-only classifier fails to predict correctly.

    Evaluating full premise-hypothesis NLI models on the Hard subset measures model performance on instances where shallow lexical and structural artifacts in the hypothesis are insufficient to determine the entailment relationship.

  4. Knowl 4 — Smoothed Pointwise Mutual Information for Quantifying Lexical Class Artifacts

    data/table

    To identify specific words that disproportionately correlate with each NLI class, the pointwise mutual information (PMI) between word tokens and class labels c∈{Entailment,Neutral,Contradiction}c \in \{\text{Entailment}, \text{Neutral}, \text{Contradiction}\} is computed over the training corpus:

    PMI(word,class)=log⁡p(word,class)p(word,⋅)p(⋅,class)\text{PMI}(\text{word}, \text{class}) = \log \frac{p(\text{word}, \text{class})}{p(\text{word}, \cdot) p(\cdot, \text{class})}

    Add-100 smoothing is applied to the raw counts to highlight highly discriminative word-class associations. The top words ranked by smoothed PMI, along with the percentage of training sentences in each class containing that word, are as follows:

    Dataset Entailment Neutral Contradiction
    SNLI outdoors 2.8% tall 0.7% nobody 0.1%
    least 0.2% first 0.6% sleeping 3.2%
    instrument 0.5% competition 0.7% no 1.2%
    outside 8.0% sad 0.5% tv 0.4%
    animal 0.7% favorite 0.4% cat 1.3%
    MultiNLI some 1.6% also 1.4% never 5.0%
    yes 0.1% because 4.1% no 7.6%
    something 0.9% popular 0.7% nothing 1.4%
    sometimes 0.2% many 2.2% any 4.1%
    various 0.1% most 1.8% none 0.1%
  5. Knowl 5 — Crowd-Worker Heuristic Strategies Producing NLI Annotation Artifacts

    empirical result

    Statistical analysis reveals consistent generation heuristics adopted by crowd workers when constructing hypotheses for each NLI relation:

    • Entailment: Annotators frequently generalize specific premise entities into broad hypernyms (e.g., replacing dog with animal, guitar with instrument, or scenery mentions with outdoors), replace exact counts with approximate quantifiers (some, at least, various), and remove explicit gender references (person, human).
    • Neutral: Annotators commonly invent plausible but unconfirmed details by adding descriptive modifiers (tall, sad, popular), superlatives (first, favorite, most), or causal/purpose clauses introduced by discourse markers like because or infinitive clauses (e.g., to help provide for her family).
    • Contradiction: Annotators predominantly rely on overt negation words (no, never, nothing, nobody), state universal activity negations (such as describing a subject as sleeping to contradict any active premise), or invoke opposite states (such as naked to contradict clothing descriptions, or cat when the premise involves dogs).

    Many of these heuristics align directly with the specific illustrative examples provided in the crowd-worker annotation guidelines, indicating that guideline prompts primed annotator generation behavior.

  6. Knowl 6 — Hypothesis Length and Token Subsumption Biases in NLI Data

    empirical result

    Hypothesis sentence length and token containment vary systematically across NLI classes:

    1. Sentence Length Skew: Neutral hypotheses are systematically longer (median token length of 9 in SNLI; half of all hypotheses containing 12 or more tokens belong to the neutral class). Entailed hypotheses are substantially shorter (60%60\% of entailments contain 7 tokens or fewer, and hypotheses of length 5 and under are predominantly entailments).
    2. Premise Subsumption: When represented as bags of words, 8.8%8.8\% of entailed hypotheses in SNLI are completely contained within their corresponding premise sentence, compared to only 0.2%0.2\% of neutral hypotheses and 0.2%0.2\% of contradiction hypotheses. MultiNLI exhibits matching patterns.
  7. Knowl 7 — Pitfalls of Post-Hoc Dataset Filtering to Eliminate Annotation Artifacts

    limitation

    Filtering out "Easy" (artifact-exploitable) examples from NLI datasets does not yield an effective, artifact-free training benchmark due to two structural issues:

    1. Creation of Inverted Artifacts: Removing examples where a strong lexical cue indicates its primary class (e.g., removing all contradiction examples containing no) leaves remaining instances of no predominantly in the neutral and entailment classes, creating a reverse artifact that models can exploit.
    2. Loss of Valid Semantic Phenomena: Artifacts reflect a skewed sample distribution rather than incorrect annotations. "Easy" examples often contain legitimate semantic relationships (such as hypernymy between dog and animal). Removing them deprives models of data needed to learn genuine inference rules.

Coverage note — None was omitted; all key empirical findings, artifact analyses, partitioning methodologies, and theoretical limitations contributed by the paper are covered.

References

  1. 1.Aishwarya Agrawal, Dhruv Batra, and Devi Parikh. 2016. Analyzing the behavior of visual question answering models. In Proc. of EMNLP. https://aclweb.org/anthology/D16-1203.
  2. 2.Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. VQA: Visual question answering. In Proc. of ICCV. https://arxiv.org/abs/1506.00278.
  3. 3.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proc. of EMNLP. https://doi.org/10.18653/v1/D15-1075.
  4. 4.Zheng Cai, Lifu Tu, and Kevin Gimpel. 2017. Pay attention to the ending:strong neural baselines for the ROC Story Cloze task. In Proc. of ACL. https://doi.org/10.18653/v1/P17-2097.
  5. 5.Danqi Chen, Jason Bolton, and Christopher D. Manning. 2016. A thorough examination of the CNN/Daily Mail reading comprehension task. In Proc. of ACL. https://doi.org/10.18653/v1/P16-1223.
  6. 6.Qian Chen, Xiaodan Zhu, Zhen-Hua Ling, and Diana Inkpen. 2017. Natural language inference with external knowledge. arXiv:1711.04289. https://arxiv.org/abs/1711.04289.
  7. 7.Volkan Cirik, Louis-Philippe Morency, and Taylor Berg-Kirkpatrick. 2018. Visual referring expression recognition: What do our systems actually learn? In Proc. of NAACL.
  8. 8.Alexis Conneau, Douwe Kiela, Holger Schwenk, Loic Barrault, and Antoine Bordes. 2017. Supervised learning of universal sentence representations from natural language inference data. In Proc. of EMNLP. https://doi.org/10.18653/v1/D17-1070.
  9. 9.Ido Dagan, Oren Glickman, and Bernardo Magnini. 2006. The PASCAL recognising textual entailment challenge. Machine Learning Challenges pages 177–190. https://doi.org/10.1007/11736790_9.
  10. 10.Ishita Dasgupta, Demi Guo, Andreas Stuhlmüller, Samuel J Gershman, and Noah D Goodman. 2018. Evaluating compositionality in sentence embeddings. https://arxiv.org/abs/1802.04302.
  11. 11.Yichen Gong, Heng Luo, and Jian Zhang. 2018. Natural language inference over interaction space. In Proc. of ICLR. https://arxiv.org/abs/1709.04348.
  12. 12.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In Proc. of CVPR. https://arxiv.org/abs/1612.00837.
  13. 13.Karl Moritz Hermann, Tomáš Koćiský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. In Proc. of NIPS. http://dl.acm.org/citation.cfm?id=2969239.2969428.
  14. 14.Allan Jabri, Armand Joulin, and Laurens van der Maaten. 2016. Revisiting visual question answering baselines. In Proc. of ECCV. https://doi.org/10.1007/978-3-319-46484-8_44.
  15. 15.Robin Jia and Percy Liang. 2017. Adversarial examples for evaluating reading comprehension systems. In Proc. of EMNLP. https://www.aclweb.org/anthology/D17-1215.
  16. 16.Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2017. Bag of tricks for efficient text classification. In Proc. of EACL. https://doi.org/10.18653/v1/E17-2068.
  17. 17.Alice Lai and Julia Hockenmaier. 2014. Illinois-LH: A denotational and distributional approach to semantics. In Proc. of SemEval. https://doi.org/10.3115/v1/S14-2055.
  18. 18.Omer Levy and Ido Dagan. 2016. Annotating relation inference in context via question answering. In Proc. of ACL. https://doi.org/10.18653/v1/P16-2041.
  19. 19.Omer Levy, Steffen Remus, Chris Biemann, and Ido Dagan. 2015. Do supervised distributional methods really learn lexical inference relations? In Proc. of NAACL. https://doi.org/10.3115/v1/N15-1098.
  20. 20.Dekang Lin and Patrick Pantel. 2001. Discovery of inference rules for question-answering. Natural Language Engineering 7(4):343–360. https://doi.org/10.1017/S1351324901002765.
  21. 21.Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli. 2014. A SICK cure for the evaluation of compositional distributional semantic models. In Proc. of LREC. pages 216–223. https://doi.org/10.3115/v1/S14-2001.
  22. 22.Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. 2016. A corpus and Cloze evaluation for deeper understanding of commonsense stories. In Proc. of NAACL. https://doi.org/10.18653/v1/N16-1098.
  23. 23.Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit. 2016. A decomposable attention model for natural language inference. In Proc. of EMNLP. https://doi.org/10.18653/v1/D16-1244.
  24. 24.Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018. Hypothesis only baselines for natural language inference. In *Proc of SEM.
  25. 25.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proc. of EMNLP. https://doi.org/10.18653/v1/D16-1264.
  26. 26.Rachel Rudinger, Chandler May, and Benjamin Van Durme. 2017. Social bias in elicited natural language inferences. In Proc. of EthNLP. http://www.aclweb.org/anthology/W17-1609.
  27. 27.Roy Schwartz, Maarten Sap, Ioannis Konstas, Li Zilles, Yejin Choi, and Noah A. Smith. 2017. The effect of different writing tasks on linguistic style: A case study of the ROC Story Cloze task. In Proc. of CoNLL. https://doi.org/10.18653/v1/K17-1004.
  28. 28.Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proc. of NAACL.
  29. 29.Naomi Zeichner, Jonathan Berant, and Ido Dagan. 2012. Crowdsourcing inference-rule evaluation. In Proc. of ACL. http://www.aclweb.org/anthology/P12-2031.

Citation

MLA
Gururangan, S., et al. “Annotation Artifacts in Natural Language Inference Data”. arXiv, 2018, http://arxiv.org/abs/1803.02324v2.
APA
Gururangan, S., Swayamdipta, S., Levy, O., Schwartz, R., Bowman, S. R., & Smith, N. A. (2018). Annotation Artifacts in Natural Language Inference Data. arXiv. http://arxiv.org/abs/1803.02324v2
Chicago
Gururangan, S., S. Swayamdipta, O. Levy, R. Schwartz, S. R. Bowman, and N. A. Smith. 2018. “Annotation Artifacts in Natural Language Inference Data”. arXiv. http://arxiv.org/abs/1803.02324v2.
Harvard
Gururangan, S. et al. (2018) “Annotation Artifacts in Natural Language Inference Data”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1803.02324v2.
Vancouver
1. Gururangan S, Swayamdipta S, Levy O, Schwartz R, Bowman SR, Smith NA (2018) Annotation Artifacts in Natural Language Inference Data. arXiv

BibTeX

@article{gururangan2018annotation,
  title = {Annotation Artifacts in Natural Language Inference Data},
  author = {Gururangan, Suchin and Swayamdipta, Swabha and Levy, Omer and Schwartz, Roy and Bowman, Samuel R. and Smith, Noah A.},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1803.02324v2},
  eprint = {1803.02324}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/