Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets

Yuxiang WuMatt GardnerPontus StenetorpPradeep Dasigi

article2022ACL77 citations

Proposes a data-centric debiasing framework that pairs unlikelihood-trained text generators with statistical feature filtering to create synthetic natural language inference datasets that improve out-of-distribution and adversarial generalization without modifying model architectures.

Listen

Machine learning models for natural language processing frequently rely on spurious correlations—accidental statistical shortcuts between non-essential features and target labels—rather than developing genuine language comprehension. These shortcuts are typically introduced during the data annotation and collection process. While models trained on such datasets perform well in controlled testing, they fail when deployed across new distributions or exposed to adversarial inputs. This creates operational risks, unpredictable failures, and unreliability in production applications.

The article demonstrates a dataset-centric framework to resolve this issue by automatically generating and filtering synthetic training data. Instead of modifying model architectures or training objectives, the authors evaluate whether replacing standard training corpora with newly generated, debiased datasets can enhance generalisation and robustness across diverse, out-of-distribution evaluation suites.

The authors implemented a generative pipeline using pretrained language models to produce synthetic text samples for natural language inference benchmarks. To ensure data quality, the pipeline applied unlikelihood loss objectives to enforce logical consistency between text and labels, followed by confidence filtering to remove ungrammatical samples. The generated samples were subsequently filtered using standardized statistical metrics (z-statistics) to identify and eliminate instances exhibiting high correlations with superficial shortcuts, such as word overlap, sentence length, and predictive single-word cues. The resulting debiased datasets were evaluated on standard benchmark splits, challenge evaluation sets, syntactic diagnostic suites, and an adversarial attack benchmark across standard baseline classifiers and larger modern architectures.

Models trained on the debiased datasets consistently outperformed baseline models trained on original datasets across all challenging evaluations. Syntactic robustness improved by up to 13.3 percentage points on diagnostic tests, while accuracy on adversarial challenge sets increased by up to 2.8 percentage points. In adversarial attack benchmarks covering multiple linguistic stress tests, debiased datasets provided an average improvement of up to 4.2 percentage points, exceeding manual augmentation heuristics. Furthermore, the approach proved complementary to existing model-centric debiasing algorithms: combining the debiased data with ensembling methods set new performance records on challenge benchmarks. The benefits also scaled successfully to larger language models, yielding average gains of 1.1 to 2.3 percentage points.

These findings indicate that addressing data-level artifacts is an effective, non-invasive alternative to complex architectural changes. Organizations can improve system robustness without redesigning downstream production pipelines, altering existing training objectives, or maintaining specialized model components. This separation of data curation from model deployment reduces operational complexity and engineering overhead.

Decision-makers should consider adopting data-level debiasing pipelines during training data preparation, particularly for safety-critical and customer-facing language processing tasks. When maximum robustness is required, technical teams can combine debiased corpora with existing algorithmic debiasing methods. Future development should focus on testing this generative debiasing strategy beyond sentence inference tasks, expanding into broader enterprise text classification, information extraction, and decision-support systems.

The findings are supported by consistent results across multiple model sizes and diverse benchmark suites. However, potential limitations remain regarding the computational overhead required to generate and filter synthetic samples at scale. Additionally, the approach relies on identifying known categories of bias to track and filter statistical shortcuts; unmapped or entirely novel spurious features may still persist in the final datasets.

Wu et al (2022).pdf
Cover for Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets

Abstract

Natural language processing models often exploit spurious correlations between task-independent features and labels in datasets to perform well only within the distributions they are trained on, while not generalising to different task distributions. We propose to tackle this problem by generating a debiased version of a dataset, which can then be used to train a debiased, off-the-shelf model, by simply replacing its training data. Our approach consists of 1) a method for training data generators to generate high-quality, label-consistent data samples; and 2) a filtering mechanism for removing data points that contribute to spurious correlations, measured in terms of z-statistics. We generate debiased versions of the SNLI and MNLI datasets, and we evaluate on a large suite of debiased, out-of-distribution, and adversarial test sets. Results show that models trained on our debiased datasets generalise better than those trained on the original datasets in all settings. On the majority of the datasets, our method outperforms or performs comparably to previous state-of-the-art debiasing strategies, and when combined with an orthogonal technique, product-of-experts, it improves further and outperforms previous best results of SNLI-hard and MNLI-hard.

Table of Contents

  • 1 Introduction
  • 2 Generating High-Quality Data Samples
  • 2.1 Finetuning Pretrained Language Model to Generate NLI Samples
  • 2.2 Improving Data Generation Quality
  • 2.2.1 Unlikelihood Training to Improve Label Consistency
  • 2.2.2 Filtering Based on Model Confidence
  • 3 Mitigating Spurious Correlations using z-filtering
  • 3.1 Identifying and Measuring Spurious Correlations
  • 3.2 z-filtering
  • 4 Constructing Debiased NLI Datasets via Data Generation
  • 4.1 Learning to Generate Unbiased Samples
  • 4.2 Combining with z-filtering to Construct the Debiased NLI Datasets
  • 5 Experiments
  • 5.1 Experimental Setup
  • 5.2 Hypothesis-only Bias in NLI
  • 5.3 Syntactic Bias in NLI
  • 5.4 Adversarial Tests for Combating Distinct Biases in NLI
  • 5.5 Generalisation to Larger Pretrained Language Models
  • 6 Related Work
  • 7 Conclusions
  • Acknowledgments
  • References
  • A Hyperparameters
  • A.1 Hyperparameters of Our Proposed Method
  • A.2 Hyperparameter Tuning of PoE
  • B Task-independent Features
  • C Description of Adversarial Test (Liu et al., 2020b) Subcategories
  • D Visualisation of z-statistics
  • E Ablation Study
  • F Generated Samples of Debiased Dataset

Knowls

  1. Knowl 1 — Autoregressive generation of premise–label–hypothesis examples

    model/method

    The data generator is a GPT-2 large language model fine-tuned on an NLI dataset to generate each example in premise, label, hypothesis order. For an example with premise PP, label ll, and hypothesis HH, the training objective is the negative log-likelihood

    LMLE=−∑i=1∣D0∣log⁡pG(P(i))pG(l(i)∣P(i))pG(H(i)∣l(i),P(i)),\mathcal{L}_{\mathrm{MLE}}=-\sum_{i=1}^{|D_0|}\log p_G(P^{(i)})p_G(l^{(i)}\mid P^{(i)})p_G(H^{(i)}\mid l^{(i)},P^{(i)}),

    where D0D_0 is the source NLI dataset, ii indexes its examples, and pGp_G is the generator's probability distribution. This factorization trains the model to produce a premise first, then its label, and finally a hypothesis conditioned on both. The paper reports using this sequence order because it performed better in a preliminary comparison with alternative orders. Generator training used learning rate 10−510^{-5}, batch size 24, five epochs, Adam, and maximum sequence length 128.

  2. Knowl 2 — Label-consistency training and confidence filtering improve generated examples

    model/method

    To discourage a generator from producing hypotheses that conflict with their assigned NLI labels, the method creates negative training examples by replacing each example's true label l(i)l^{(i)} with an incorrect label l′(i)l'^{(i)}. It applies token-level unlikelihood training to the original hypothesis tokens Ht(i)H_t^{(i)}:

    Lconsistency=−∑i∑t=1∣H(i)∣log⁡(1−pG(Ht(i)∣l′(i),P(i),H<t(i))),\mathcal{L}_{\mathrm{consistency}}=-\sum_i\sum_{t=1}^{|H^{(i)}|}\log\left(1-p_G\left(H_t^{(i)}\mid l'^{(i)},P^{(i)},H_{<t}^{(i)}\right)\right),

    where H<t(i)H_{<t}^{(i)} denotes the hypothesis tokens preceding token tt. The generator is trained with LG=LMLE+λLconsistency\mathcal{L}_G=\mathcal{L}_{\mathrm{MLE}}+\lambda\mathcal{L}_{\mathrm{consistency}}; the experiments use λ=0.5\lambda=0.5. Generated examples are then confidence-filtered: retain (P,H,l)(P,H,l) only when an NLI classifier MM, trained on the source dataset, assigns the label probability pM(l∣P,H)>τp_M(l\mid P,H)>\tau. The experiments use RoBERTa-large as MM and τ=0.95\tau=0.95. The paper reports that rejected examples often have ungrammatical text or incorrect labels.

  3. Knowl 3 — Task-independent NLI features are measured by deviations from uniform label rates

    equation

    The debiasing method assumes that each selected task-independent feature should have the same probability under every NLI label. The feature set comprises premise and hypothesis unigrams and bigrams treated separately; hypothesis length; the hypothesis-to-premise length ratio; lexical overlap, measured as the proportion of hypothesis tokens occurring in the premise; predictions from a BERT-base hypothesis-only classifier trained on the source dataset; and a null feature. For a feature xx occurring in nn examples and a label ll, let p^(l∣x)\hat p(l\mid x) be the fraction of those examples with label ll. Its standardized deviation from the uniform label rate is

    z∗(x,l)=p^(l∣x)−p0p0(1−p0)/n,z^*(x,l)=\frac{\hat p(l\mid x)-p_0}{\sqrt{p_0(1-p_0)/n}},

    where p0=1/3p_0=1/3 for the three NLI labels entailment, neutral, and contradiction. For each label, the method designates the kk features with the highest z-statistics as biased; the experiments use k=20k=20. The selected feature set is not intrinsic to the method: other task-independent features can be substituted or added.

  4. Knowl 4 — Iterative z-filtering constructs a subset that avoids label-associated features

    algorithm

    z-filtering builds an accepted dataset ZZ and a rejected set Z−Z^{-} from an input dataset D0D_0. For each label, it identifies the kk features with the largest z-statistics in the currently accepted data. An instance is accepted only if none of its features belongs to the biased-feature set for its label; statistics and biased-feature sets are updated as the method processes successive batches. Processing stops after all input examples have been considered. An optional seed dataset initializes ZZ; with a seed, input examples are tested against the seed's biased features, giving conditional z-filtering. The output is the accepted set ZZ and rejected set Z−Z^{-}.

    Input: dataset D0D_0, feature set XX, number of biased features kk, optional seed dataset DseedD_{seed}
    Output: accepted dataset ZZ, rejected dataset Z−Z^{-}
    Initialize ZZ to DseedD_{seed} if provided, otherwise to the empty set
    Initialize Z−Z^{-} to the empty set
    For each successive batch DtD_t from D0D_0:
        Compute or update z∗(x,l∣Z)z^*(x,l\mid Z) for features xx in XX and each NLI label ll
        For each label ll, let BZ(l)B_Z(l) be its kk highest-scoring features
        For each instance I=(P,H,l)I=(P,H,l) in DtD_t:
            Extract the feature set f(I)f(I)
            If f(I)f(I) has no feature in BZ(l)B_Z(l):
                Add II to ZZ
            Otherwise:
                Add II to Z−Z^{-}
    Return ZZ and Z−Z^{-}
  5. Knowl 5 — Biased-feature unlikelihood trains the generator to produce filter-acceptable data

    model/method

    Directly z-filtering generated data can discard most of it; for SNLI, the paper reports that filtering removed about 85% of samples from a generator fine-tuned on SNLI. To make acceptable examples more likely before post-hoc filtering, the method further trains the generator on z-filtered source data. It maximizes likelihood on accepted examples Z(D0)Z(D_0) and applies unlikelihood training to rejected examples Z−(D0)Z^{-}(D_0):

    Ldebias=LMLE(Z(D0))+αLUL(Z−(D0)).\mathcal{L}_{\mathrm{debias}}=\mathcal{L}_{\mathrm{MLE}}(Z(D_0))+\alpha\mathcal{L}_{\mathrm{UL}}(Z^{-}(D_0)).

    For each token It−I_t^{-} in a rejected example I−I^{-}, a mask mtm_t is zero if that token contributes to a biased feature for the example's label and one otherwise. The token loss is

    LUL(I−)=−∑tlog⁡(mtpG(It−∣I<t−)+(1−mt)(1−pG(It−∣I<t−))).\mathcal{L}_{\mathrm{UL}}(I^{-})=-\sum_t\log\left(m_t p_G(I_t^{-}\mid I_{<t}^{-})+(1-m_t)(1-p_G(I_t^{-}\mid I_{<t}^{-}))\right).

    Thus, tokens contributing to a biased feature receive an unlikelihood penalty, while the other tokens retain a likelihood term to reduce degenerate outputs. For unigram and bigram biases, only the corresponding tokens are marked as contributing; for other feature types, such as hypothesis length, all hypothesis tokens are marked. The experiments use α=1.0\alpha=1.0. The resulting generator is sampled and its output is confidence-filtered as described in the quality-control method.

  6. Knowl 6 — Three dataset-construction strategies combine filtered synthetic and source examples

    model/method

    Let D0D_0 be the original dataset, D^G∗\hat D_{G^*} the generated data after confidence filtering, and Z(A∣B)Z(A\mid B) the result of conditional z-filtering dataset AA using seed dataset BB. The method constructs three alternatives: Z-Aug keeps all of D0D_0 and adds Z(D^G∗∣D0)Z(\hat D_{G^*}\mid D_0); Par-Z independently z-filters D0D_0 and D^G∗\hat D_{G^*} and takes their union; Seq-Z first forms Z(D0)Z(D_0) and then adds Z(D^G∗∣Z(D0))Z(\hat D_{G^*}\mid Z(D_0)). The page 1 overview graphic depicts the generator-to-filter pipeline and represents labels by point shape and task-independent features by color.

    The experiments sampled 5,000,000 examples from the SNLI generator and 4,000,000 from the MNLI generator, and used 20 biased features per label. The resulting training-set sizes were:

    Dataset construction SNLI examples MNLI examples
    Original 549,367 382,702
    Z-Aug 1,142,475 744,326
    Par-Z 933,085 740,811
    Seq-Z 927,906 744,200

    These alternatives trade off retaining the original examples against removing source-data correlations: Z-Aug retains the original dataset, whereas Par-Z and Seq-Z filter it.

  7. Knowl 7 — Debiased training data improves performance on hypothesis-only hard sets

    empirical result

    BERT-base models trained with ordinary cross-entropy on the proposed datasets were evaluated for accuracy on SNLI and its hypothesis-only hard set, and on the MNLI matched and mismatched test sets and their corresponding hard sets. The hard sets are designed to challenge systems that rely on hypothesis-only predictions. Most reported non-PoE results are averages over five runs; MNLI test results use the checkpoint selected on the hard-set development data. The table reports test accuracy in percent.

    SNLI training data SNLI SNLI-hard
    Original 90.45 80.34 ±\pm 0.46
    Z-Aug 90.67 81.78 ±\pm 0.53
    Par-Z 88.11 82.81 ±\pm 0.37
    Seq-Z 88.08 82.82 ±\pm 0.15
    Original + PoE 90.25 82.92
    Seq-Z + PoE 87.65 84.48
    MNLI training data MNLI-m MNLI-mm
    Original 84.11 83.51
    Z-Aug 85.12 84.09
    Par-Z 83.27 82.95
    Seq-Z 83.41 83.17
    Original + PoE 84.69 83.75
    Z-Aug + PoE 85.38 84.53
    MNLI training data MNLI-m hard MNLI-mm hard
    Original 75.88 75.75
    Z-Aug 78.60 78.51
    Par-Z 79.19 78.49
    Seq-Z 79.19 78.44
    Original + PoE 77.54 78.33
    Z-Aug + PoE 80.03 79.28

    Compared with the original-data baselines, each debiased-data variant scores higher on the corresponding hard sets. Z-Aug also raises both MNLI test scores, while combining PoE with Seq-Z for SNLI or Z-Aug for MNLI further improves hard-set accuracy.

  8. Knowl 8 — Debiased MNLI data improves accuracy on the HANS syntactic challenge set

    empirical result

    HANS tests NLI models on examples where common syntactic heuristics, including lexical-overlap and subsequence heuristics, can lead to incorrect predictions. BERT-base trained on original MNLI reaches 54.36 ±\pm 2.56 accuracy on HANS; the debiased-data variants reach 62.57 ±\pm 5.91 with Z-Aug, 65.11 ±\pm 5.62 with Par-Z, and 67.69 ±\pm 3.53 with Seq-Z. These results are means and standard deviations over five runs. Seq-Z therefore gains 13.33 percentage points over the original-data BERT-base baseline. Adding PoE raises HANS accuracy from 63.40 with original MNLI to 68.75 with Z-Aug MNLI. RoBERTa-large also improves, from 75.74 ±\pm 2.82 on original MNLI to 78.65 ±\pm 2.26 on Z-Aug MNLI, showing that the HANS benefit is not limited to BERT-base.

  9. Knowl 9 — Debiased MNLI data raises scores across a diverse adversarial NLI benchmark

    empirical result

    The method was evaluated on seven categories of an adversarial NLI benchmark: PI-CD (classifier-detected partial-input artifacts), PI-SP (surface-pattern heuristics), IS-SD (syntactic diagnostics), IS-CS (lexically misleading cases), LI-LI (lexical inference), LI-TS (premise–hypothesis swap), and ST (stress tests of word overlap, negation, length mismatch, and spelling). The table gives accuracy in percent for BERT-base trained on original MNLI or each debiased MNLI variant; values are means and standard deviations over five runs.

    Training data PI-CD PI-SP IS-SD IS-CS LI-LI LI-TS ST Average
    Original MNLI 70.3 ±\pm 0.5 73.7 ±\pm 1.4 53.5 ±\pm 2.3 64.8 ±\pm 1.4 85.5 ±\pm 0.9 81.6 ±\pm 1.4 69.2 ±\pm 0.8 71.2 ±\pm 0.8
    Z-Aug 73.1 ±\pm 0.9 76.1 ±\pm 1.2 61.8 ±\pm 6.1 69.1 ±\pm 1.3 86.9 ±\pm 0.6 83.1 ±\pm 0.9 70.1 ±\pm 0.5 74.3 ±\pm 1.3
    Par-Z 72.0 ±\pm 0.9 78.7 ±\pm 1.2 64.5 ±\pm 5.8 70.7 ±\pm 1.7 88.5 ±\pm 0.7 82.6 ±\pm 0.3 69.6 ±\pm 1.0 75.2 ±\pm 1.4
    Seq-Z 71.7 ±\pm 0.9 77.8 ±\pm 1.2 66.9 ±\pm 3.7 71.1 ±\pm 0.7 89.1 ±\pm 1.0 82.3 ±\pm 0.9 69.3 ±\pm 0.8 75.4 ±\pm 0.8

    All three debiased-data variants have a higher mean score than the original-MNLI baseline in every listed category and on the benchmark average. Their average scores exceed the baseline by 3.1 to 4.2 percentage points.

  10. Knowl 10 — Debiased training data also benefits larger pretrained language models

    empirical result

    The authors tested whether data-level debiasing transfers beyond BERT-base by training RoBERTa-base, RoBERTa-large, and ALBERT-xxlarge on original or debiased data. For SNLI-hard, the debiased training set is Seq-Z SNLI; for the other evaluation sets it is Z-Aug MNLI. Accuracy is reported in percent. Results for RoBERTa are means and standard deviations over five runs; ALBERT-xxlarge was run once.

    Model Evaluation set Original data Debiased data
    RoBERTa-base SNLI-hard 82.02 ±\pm 0.24 83.71 ±\pm 0.31
    MNLI-m hard 81.74 ±\pm 0.44 83.14 ±\pm 0.25
    MNLI-mm hard 81.93 ±\pm 0.30 83.12 ±\pm 0.24
    HANS 71.17 ±\pm 2.95 76.15 ±\pm 1.52
    Adversarial benchmark average 77.63 ±\pm 0.49 79.89 ±\pm 0.38
    RoBERTa-large SNLI-hard 83.61 ±\pm 0.31 85.09 ±\pm 0.32
    MNLI-m hard 85.44 ±\pm 0.62 85.69 ±\pm 0.24
    MNLI-mm hard 85.37 ±\pm 0.63 85.94 ±\pm 0.21
    HANS 75.74 ±\pm 2.82 78.65 ±\pm 2.26
    Adversarial benchmark average 80.92 ±\pm 0.46 81.86 ±\pm 0.31
    ALBERT-xxlarge SNLI-hard 83.59 84.82
    MNLI-m hard 86.42 86.40
    MNLI-mm hard 86.38 86.82
    HANS 76.32 79.05
    Adversarial benchmark average 81.91 83.18

    Debiased data improves the adversarial-benchmark average for all three model families. The reported average gains are 2.30 points for RoBERTa-base, 1.23 for RoBERTa-large, and 1.13 for ALBERT-xxlarge; the ALBERT-xxlarge MNLI-m hard score is a small exception, decreasing from 86.42 to 86.40.

Coverage note — The supplementary ablation comparisons and generated-example displays are omitted: they provide supporting component analyses and illustrations, while the knowls retain the central method, construction choices, and main evaluation results.

References

  1. 1.Max Bartolo, Tristan Thrush, Robin Jia, Sebastian Riedel, Pontus Stenetorp, and Douwe Kiela. 2021. Improving question answering model robustness with synthetic adversarial data generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8830–8848, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  2. 2.Yonatan Belinkov, Adam Poliak, Stuart Shieber, Benjamin Van Durme, and Alexander Rush. 2019a. Don’t take the premise for granted: Mitigating artifacts in natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 877–891, Florence, Italy. Association for Computational Linguistics.
  3. 3.Yonatan Belinkov, Adam Poliak, Stuart Shieber, Benjamin Van Durme, and Alexander Rush. 2019b. On adversarial removal of hypothesis-only bias in natural language inference. In *Proceedings of the Eighth Joint Conference on Lexical and Computational Semantics (SEM 2019), pages 256–262, Minneapolis, Minnesota. Association for Computational Linguistics.
  4. 4.Prajjwal Bhargava, Aleksandr Drozd, and Anna Rogers. 2021. Generalization in NLI: Ways (not) to go beyond simple heuristics. In Proceedings of the Second Workshop on Insights from Negative Results in NLP, pages 125–135, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  5. 5.Samuel R. Bowman. 2021. When combating hype, proceed with caution. CoRR, abs/2110.08300.
  6. 6.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal. Association for Computational Linguistics.
  7. 7.Ronan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers, Matthew E. Peters, Ashish Sabharwal, and Yejin Choi. 2020. Adversarial filters of dataset biases. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research, pages 1078–1088. PMLR.
  8. 8.Danqi Chen, Jason Bolton, and Christopher D. Manning. 2016. A thorough examination of the CNN/Daily Mail reading comprehension task. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2358–2367, Berlin, Germany. Association for Computational Linguistics.
  9. 9.Christopher Clark, Mark Yatskar, and Luke Zettlemoyer. 2019. Don’t take the easy way out: Ensemble based methods for avoiding known dataset biases. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4069–4082, Hong Kong, China. Association for Computational Linguistics.
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  11. 11.Matt Gardner, William Merrill, Jesse Dodge, Matthew Peters, Alexis Ross, Sameer Singh, and Noah A. Smith. 2021. Competency problems: On finding and removing artifacts in language data. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1801–1813, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  12. 12.Mor Geva, Yoav Goldberg, and Jonathan Berant. 2019. Are we modeling the task or the annotator? an investigation of annotator bias in natural language understanding datasets. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1161–1166, Hong Kong, China. Association for Computational Linguistics.
  13. 13.Abbas Ghaddar, Phillippe Langlais, Mehdi Rezagholizadeh, and Ahmad Rashid. 2021. End-to-end self-debiasing framework for robust NLU training. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 1923–1929, Online. Association for Computational Linguistics.
  14. 14.Max Glockner, Vered Shwartz, and Yoav Goldberg. 2018. Breaking NLI systems with sentences that require simple lexical inferences. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 650–655, Melbourne, Australia. Association for Computational Linguistics.
  15. 15.Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018. Annotation artifacts in natural language inference data. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 107–112, New Orleans, Louisiana. Association for Computational Linguistics.
  16. 16.He He, Sheng Zha, and Haohan Wang. 2019. Unlearn dataset bias in natural language inference by fitting the residual. In Proceedings of the 2nd Workshop on Deep Learning Approaches for Low-Resource NLP (DeepLo 2019), pages 132–142, Hong Kong, China. Association for Computational Linguistics.
  17. 17.Geoffrey E Hinton. 2002. Training products of experts by minimizing contrastive divergence. Neural computation, 14(8):1771–1800.
  18. 18.Rabeeh Karimi Mahabadi, Yonatan Belinkov, and James Henderson. 2020. End-to-end bias mitigation by modelling biases in corpora. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8706–8716, Online. Association for Computational Linguistics.
  19. 19.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020. ALBERT: A lite BERT for self-supervised learning of language representations. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  20. 20.Minwoo Lee, Seungpil Won, Juae Kim, Hwanhee Lee, Cheoneum Park, and Kyomin Jung. 2021. Crossaug: A contrastive data augmentation method for debiasing fact verification models. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management, pages 3181–3185.
  21. 21.Patrick Lewis, Yuxiang Wu, Linqing Liu, Pasquale Minervini, Heinrich Kuttler, Aleksandra Piktus, Pontus Stenetorp, and Sebastian Riedel. 2021. PAQ: 65 million probably-asked questions and what you can do with them. Transactions of the Association for Computational Linguistics, 9:1098–1115.
  22. 22.Haochen Liu, Joseph Thekinen, Sinem Mollaoglu, Da Tang, Ji Yang, Youlong Cheng, Hui Liu, and Jiliang Tang. 2021. Toward annotator group bias in crowdsourcing. ArXiv preprint, abs/2110.08038.
  23. 23.Tianyu Liu, Zheng Xin, Baobao Chang, and Zhifang Sui. 2020a. HypoNLI: Exploring the artificial patterns of hypothesis-only bias in natural language inference. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 6852–6860, Marseille, France. European Language Resources Association.
  24. 24.Tianyu Liu, Zheng Xin, Xiaoan Ding, Baobao Chang, and Zhifang Sui. 2020b. An empirical study on model-agnostic debiasing strategies for robust natural language inference. In Proceedings of the 24th Conference on Computational Natural Language Learning, pages 596–608, Online. Association for Computational Linguistics.
  25. 25.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  26. 26.Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019. Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3428–3448, Florence, Italy. Association for Computational Linguistics.
  27. 27.Pasquale Minervini and Sebastian Riedel. 2018. Adversarially regularising neural NLI models to integrate logical background knowledge. In Proceedings of the 22nd Conference on Computational Natural Language Learning, pages 65–74, Brussels, Belgium. Association for Computational Linguistics.
  28. 28.Aakanksha Naik, Abhilasha Ravichander, Norman Sadeh, Carolyn Rose, and Graham Neubig. 2018. Stress test evaluation for natural language inference. In Proceedings of the 27th International Conference on Computational Linguistics, pages 2340–2353, Santa Fe, New Mexico, USA. Association for Computational Linguistics.
  29. 29.Yixin Nie, Yicheng Wang, and Mohit Bansal. 2019. Analyzing compositionality-sensitivity of NLI models. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, Honolulu, Hawaii, USA, pages 6867–6874. AAAI Press.
  30. 30.Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018. Hypothesis only baselines in natural language inference. In Proceedings of the Seventh Joint Conference on Lexical and Computational Semantics, pages 180–191, New Orleans, Louisiana. Association for Computational Linguistics.
  31. 31.Alexis Ross, Tongshuang Wu, Hao Peng, Matthew E Peters, and Matt Gardner. 2021. Tailor: Generating and perturbing text with semantic controls. ArXiv preprint, abs/2107.07150.
  32. 32.Victor Sanh, Thomas Wolf, Yonatan Belinkov, and Alexander M. Rush. 2021. Learning from others’ mistakes: Avoiding dataset biases without modeling them. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  33. 33.Timo Schick and Hinrich Schütze. 2021. Generating datasets with pretrained language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6943–6951, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  34. 34.Tal Schuster, Darsh Shah, Yun Jie Serene Yeo, Daniel Roberto Filizzola Ortiz, Enrico Santus, and Regina Barzilay. 2019. Towards debiasing fact verification models. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3419–3425, Hong Kong, China. Association for Computational Linguistics.
  35. 35.Roy Schwartz, Maarten Sap, Ioannis Konstas, Leila Zilles, Yejin Choi, and Noah A. Smith. 2017. The effect of different writing tasks on linguistic style: A case study of the ROC story cloze task. In Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017), pages 15–25, Vancouver, Canada. Association for Computational Linguistics.
  36. 36.Joe Stacey, Yonatan Belinkov, and Marek Rei. 2021. Supervising model attention with human explanations for robust natural language inference. ArXiv preprint, abs/2104.08142.
  37. 37.Joe Stacey, Pasquale Minervini, Haim Dubossarsky, Sebastian Riedel, and Tim Rocktäschel. 2020. Avoiding the Hypothesis-Only Bias in Natural Language Inference via Ensemble Adversarial Training. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8281–8291, Online. Association for Computational Linguistics.
  38. 38.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER: a large-scale dataset for fact extraction and VERification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 809–819, New Orleans, Louisiana. Association for Computational Linguistics.
  39. 39.Prasetya Ajie Utama, Nafise Sadat Moosavi, and Iryna Gurevych. 2020. Mind the trade-off: Debiasing NLU models without degrading the in-distribution performance. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8717–8729, Online. Association for Computational Linguistics.
  40. 40.Haohan Wang, Da Sun, and Eric P. Xing. 2019. What if we simply swap the two text fragments? A straightforward yet effective way to test the robustness of methods to confounding signals in nature language inference tasks. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, Honolulu, Hawaii, USA, pages 7136–7143. AAAI Press.
  41. 41.Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston. 2020. Neural text generation with unlikelihood training. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  42. 42.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122, New Orleans, Louisiana. Association for Computational Linguistics.
  43. 43.Yiben Yang, Chaitanya Malaviya, Jared Fernandez, Swabha Swayamdipta, Ronan Le Bras, Ji-Ping Wang, Chandra Bhagavatula, Yejin Choi, and Doug Downey. 2020. G-daug: Generative data augmentation for commonsense reasoning. In EMNLP (Findings), volume EMNLP 2020 of Findings of ACL, pages 1008–1025. Association for Computational Linguistics.
  44. 44.Xiang Zhou and Mohit Bansal. 2020. Towards robustifying NLI models against lexical dataset biases. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8759–8771, Online. Association for Computational Linguistics.

Citation

MLA
Wu, Y., et al. “Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 2660–76, https://doi.org/10.18653/v1/2022.acl-long.190.
APA
Wu, Y., Gardner, M., Stenetorp, P., & Dasigi, P. (2022). Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2660–2676. https://doi.org/10.18653/v1/2022.acl-long.190
Chicago
Wu, Y., M. Gardner, P. Stenetorp, and P. Dasigi. 2022. “Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2660–76. https://doi.org/10.18653/v1/2022.acl-long.190.
Harvard
Wu, Y. et al. (2022) “Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 2660–2676. Available at: https://doi.org/10.18653/v1/2022.acl-long.190.
Vancouver
1. Wu Y, Gardner M, Stenetorp P, Dasigi P (2022) Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 2660–2676

BibTeX

@inproceedings{wu-etal-2022-generating,
    title = "Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets",
    author = "Wu, Yuxiang  and
      Gardner, Matt  and
      Stenetorp, Pontus  and
      Dasigi, Pradeep",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.190/",
    doi = "10.18653/v1/2022.acl-long.190",
    pages = "2660--2676"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/