Annotation Artifacts in Natural Language Inference Data
Suchin GururanganSwabha SwayamdiptaOmer LevyRoy SchwartzSamuel R. BowmanNoah A. Smith
Demonstrates that widely used natural language inference benchmarks contain crowdsourced artifacts that allow models to predict labels from the hypothesis alone without observing the premise, proving that machine reasoning performance has been substantially overestimated.
Natural language inference is a core benchmark for artificial intelligence, testing whether a model can determine if a given premise sentence logically entails, contradicts, or is neutral toward a hypothesis sentence. However, the standard crowdsourcing process used to generate these datasets introduces unintended patterns and shortcuts, known as annotation artifacts. These artifacts allow machine learning models to appear capable of complex reasoning while merely exploiting surface-level statistical cues.
The main objective of the article is to demonstrate the widespread existence of these annotation artifacts in two prominent natural language inference datasets, identify the human heuristics that cause them, and evaluate how heavily state-of-the-art models rely on these shortcuts rather than genuine logical reasoning.
To evaluate this, the authors applied a simple text classifier that predicted labels using only the generated hypothesis sentences, without ever seeing the premise sentences. They analyzed lexical associations and sentence lengths across the datasets to diagnose human writing habits. Finally, they split test datasets into "Easy" examples (which the hypothesis-only model solved) and "Hard" examples (which it failed), re-evaluating three leading natural language inference systems across these segments.
The investigation produced four central findings. First, a simple classifier ignoring the premise entirely achieved 67.0% accuracy on the Stanford Natural Language Inference dataset and 52.3% to 53.9% on the Multi-Genre Natural Language Inference dataset, far exceeding the roughly 35% baseline of random guessing. Second, crowd workers consistently applied specific heuristics: negation words (such as "no," "never," and "nobody") strongly correlated with contradictions, purpose clauses and modifiers correlated with neutral statements, and general hypernyms (such as "animal" for "dog") correlated with entailments. Third, sentence length served as a distinct cue; neutral hypotheses were systematically longer, while 60% of entailments were seven words or fewer. Fourth, when top-performing models were evaluated on the "Hard" test subset, their performance dropped sharply—falling by 12 to 17 percentage points compared to their standard benchmark scores.
These findings indicate that the perceived reasoning capabilities of current AI models have been substantially overstated. Rather than performing true textual reasoning, these systems are largely exploiting subtle artifacts present in crowdsourced benchmarks. Relying on such benchmark scores presents significant risks for downstream applications—such as automated question answering and summarization—where systems may fail unexpectedly on real-world inputs that lack these specific statistical shortcuts.
To address this challenge, the article recommends that AI practitioners evaluate existing and future models on the newly partitioned "Hard" datasets alongside standard benchmarks. Simply filtering out easy examples from existing data is insufficient, as it introduces new statistical distortions. Instead, dataset creators should improve crowdsourcing protocols by diversifying prompt instructions and utilizing automated or adversarial systems during data collection to balance heuristic patterns across all classes.
The findings are supported with high statistical confidence across multiple standard benchmarks and model architectures. However, the authors caution that identifying and balancing every heuristic during data collection remains difficult. Stakeholders should recognize that addressing data artifacts is an ongoing challenge, and current natural language inference remains an open, unsolved problem.
- Paper: A large annotated corpus for learning natural language inference, Samuel R. Bowman et al. (2015). It introduces the Stanford Natural Language Inference (SNLI) corpus, which serves as the primary dataset whose crowdsourced hypothesis generation protocol is directly analyzed for annotation artifacts.
- Paper: A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference, Adina Williams et al. (2016). It establishes the Multi-Genre Natural Language Inference (MultiNLI) corpus, the other major benchmark evaluated by the source for hypothesis-only statistical shortcuts.
- Paper: A Decomposable Attention Model for Natural Language Inference, Ankur P. Parikh et al. (2016). It develops a canonical decomposable attention architecture for natural language inference that provides essential context for how standard premise-hypothesis matching models operate on NLI benchmarks.
- Paper: Supervised Learning of Universal Sentence Representations from Natural Language Inference Data, Alexis Conneau et al. (2017). It demonstrates the use of supervised NLI data to train sentence encoders, establishing the high benchmark scores that the source reveals to be partly driven by dataset artifacts.
- Paper: Adversarial Examples for Evaluating Reading Comprehension Systems, Robin Jia et al. (2017). It introduces adversarial evaluation methodologies to expose how NLP models exploit superficial heuristics rather than true language comprehension.
- Paper: Baselines and Bigrams: Simple, Good Sentiment and Topic Classification, Sida I. Wang et al. (2012). It provides foundational baselines for simple n-gram text classification, motivating the source's use of simple text categorization models to detect hypothesis-only cues.
- Paper: Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference, R. Thomas McCoy et al. (2019). It builds directly on the findings of dataset artifacts by introducing the HANS diagnostic dataset to test whether NLI models rely on specific structural and lexical heuristics.
- Paper: Beyond Accuracy: Behavioral Testing of NLP Models with CheckList, Marco Túlio Ribeiro et al. (2020). It generalizes behavioral and capability-driven diagnostic testing across NLP models to identify linguistic failure modes like negation handling that bypass benchmark accuracy.
- Paper: HellaSwag: Can a Machine Really Finish Your Sentence?, Rowan Zellers et al. (2019). It directly implements adversarial filtering to counteract the crowdsourcing annotation artifacts highlighted by the source when constructing robust NLI benchmarks.
- Paper: XNLI: Evaluating Cross-lingual Sentence Representations, Alexis Conneau et al. (2018). It extends the MultiNLI dataset across 15 languages, establishing a cross-lingual benchmark where evaluation must be aware of underlying English collection artifacts.
- Paper: GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding, Alex Wang et al. (2018). It integrates NLI datasets into a broader diagnostic platform and multi-task benchmark to systematically assess general natural language understanding beyond isolated task artifacts.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). It investigates how superficial heuristics and subtle biases can distort modern reasoning models and their chain-of-thought explanations.
