Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets
Yuxiang WuMatt GardnerPontus StenetorpPradeep Dasigi
Proposes a data-centric debiasing framework that pairs unlikelihood-trained text generators with statistical feature filtering to create synthetic natural language inference datasets that improve out-of-distribution and adversarial generalization without modifying model architectures.
Machine learning models for natural language processing frequently rely on spurious correlations—accidental statistical shortcuts between non-essential features and target labels—rather than developing genuine language comprehension. These shortcuts are typically introduced during the data annotation and collection process. While models trained on such datasets perform well in controlled testing, they fail when deployed across new distributions or exposed to adversarial inputs. This creates operational risks, unpredictable failures, and unreliability in production applications.
The article demonstrates a dataset-centric framework to resolve this issue by automatically generating and filtering synthetic training data. Instead of modifying model architectures or training objectives, the authors evaluate whether replacing standard training corpora with newly generated, debiased datasets can enhance generalisation and robustness across diverse, out-of-distribution evaluation suites.
The authors implemented a generative pipeline using pretrained language models to produce synthetic text samples for natural language inference benchmarks. To ensure data quality, the pipeline applied unlikelihood loss objectives to enforce logical consistency between text and labels, followed by confidence filtering to remove ungrammatical samples. The generated samples were subsequently filtered using standardized statistical metrics (z-statistics) to identify and eliminate instances exhibiting high correlations with superficial shortcuts, such as word overlap, sentence length, and predictive single-word cues. The resulting debiased datasets were evaluated on standard benchmark splits, challenge evaluation sets, syntactic diagnostic suites, and an adversarial attack benchmark across standard baseline classifiers and larger modern architectures.
Models trained on the debiased datasets consistently outperformed baseline models trained on original datasets across all challenging evaluations. Syntactic robustness improved by up to 13.3 percentage points on diagnostic tests, while accuracy on adversarial challenge sets increased by up to 2.8 percentage points. In adversarial attack benchmarks covering multiple linguistic stress tests, debiased datasets provided an average improvement of up to 4.2 percentage points, exceeding manual augmentation heuristics. Furthermore, the approach proved complementary to existing model-centric debiasing algorithms: combining the debiased data with ensembling methods set new performance records on challenge benchmarks. The benefits also scaled successfully to larger language models, yielding average gains of 1.1 to 2.3 percentage points.
These findings indicate that addressing data-level artifacts is an effective, non-invasive alternative to complex architectural changes. Organizations can improve system robustness without redesigning downstream production pipelines, altering existing training objectives, or maintaining specialized model components. This separation of data curation from model deployment reduces operational complexity and engineering overhead.
Decision-makers should consider adopting data-level debiasing pipelines during training data preparation, particularly for safety-critical and customer-facing language processing tasks. When maximum robustness is required, technical teams can combine debiased corpora with existing algorithmic debiasing methods. Future development should focus on testing this generative debiasing strategy beyond sentence inference tasks, expanding into broader enterprise text classification, information extraction, and decision-support systems.
The findings are supported by consistent results across multiple model sizes and diverse benchmark suites. However, potential limitations remain regarding the computational overhead required to generate and filter synthetic samples at scale. Additionally, the approach relies on identifying known categories of bias to track and filter statistical shortcuts; unmapped or entirely novel spurious features may still persist in the final datasets.
- Paper: Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference, R. Thomas McCoy et al. (2019). This diagnostic study established how NLI models exploit lexical-overlap and related syntactic heuristics, motivating the source’s targeted identification and filtering of shortcut correlations.
- Paper: Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization, Shiori Sagawa et al. (2019). Its analysis of spurious correlations and distribution shifts in NLI provides foundational context for why the source seeks data-level remedies to improve out-of-distribution performance.
- Paper: Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human Interventions, John Joon Young Chung et al. (2023). It extends synthetic-data debiasing into broader text classification, showing how human checks can balance generated-data diversity against label correctness.
- Paper: Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations, Zhuoyan Li et al. (2023). It broadens the synthetic-training-data question beyond NLI, testing when LLM-generated examples can support text classification across tasks with differing subjectivity.
- Paper: Exploiting Asymmetry for Synthetic Training Data Generation: SynthIE and the Case of Information Extraction, Martin Josifoski et al. (2023). It carries synthetic data generation into information extraction, using structured facts to create training text for models trained on generated examples.
