Supervising Model Attention with Human Explanations for Robust Natural Language Inference

Joe StaceyYonatan BelinkovMarek Rei

article2022AAAI54 citations

Proposes supervising transformer attention heads with human-provided explanations to simultaneously improve both in-distribution accuracy and out-of-distribution generalization in natural language inference models.

Listen

Natural language inference systems determine logical relationships between pairs of sentences, such as whether one statement supports, contradicts, or remains neutral toward another. While standard models perform well on familiar training data, they frequently rely on superficial statistical shortcuts and dataset biases rather than genuine linguistic understanding. This reliance causes them to fail when deployed on new, out-of-distribution text. Traditional debiasing techniques often penalize shortcuts directly, but this regularly degrades baseline accuracy on standard tasks.

The article demonstrates that directly teaching models to follow human reasoning patterns resolves this dilemma. Rather than focusing on suppressing specific biases, the authors supervise the internal attention mechanisms of transformer models using human-written explanations and word highlights from a large benchmark dataset. This guides the system to allocate focus to the specific words human annotators consider essential when establishing semantic relationships.

The authors implemented this approach on standard language architectures, including BERT and DeBERTa, across over 550,000 training examples. The supervision introduces an auxiliary loss that aligns the model's internal attention distribution with human explanations without requiring extra parameters or computational overhead during testing. The authors tested this method across several out-of-distribution challenge sets designed to expose superficial heuristics and hypothesis-only shortcuts.

The key findings show consistent performance and robustness gains across all benchmarks. Supervising the top three attention heads of an existing self-attention layer increased standard test accuracy on the primary benchmark by 0.40% and improved performance on its hardest subset by 0.79%. When applied to the state-of-the-art DeBERTa architecture, the method achieved a new benchmark record of 92.69% accuracy. On out-of-distribution evaluation sets, the approach boosted accuracy by up to 1.59% on heuristic challenge sets and nearly 1% on diverse multi-genre data, avoiding the typical trade-off between standard accuracy and generalizability. Attention analysis revealed that supervised models shifted focus away from punctuation and generic stop-words toward meaningful nouns, verbs, and premise context, more than doubling premise attention in the final layer.

These results demonstrate that incorporating human rationale data makes language models more reliable and interpretable without inflating deployment size or runtime inference costs. By balancing attention across both input sentences, models mitigate hypothesis-only blind spots and maintain high fidelity when encountering unfamiliar linguistic structures in production environments.

Engineering teams deploying language inference models should adopt selective attention supervision using human explanations rather than complex multi-model pipelines. Teams should specifically target a subset of attention heads rather than supervising all heads uniformly to preserve functional diversity across the network. Further research is recommended to expand explanation-guided attention methods to complex reasoning datasets with multi-sentence premises, where performance gains remain limited.

  • Paper: Annotation Artifacts in Natural Language Inference Data, Suchin Gururangan et al. (2018). This paper establishes the widespread presence of annotation artifacts and superficial shortcuts in standard NLI datasets, providing the core motivation for debiasing NLI models.
  • Paper: Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference, R. Thomas McCoy et al. (2019). This foundational work demonstrates that NLI models exploit shallow syntactic heuristics instead of genuine reasoning, motivating human explanation supervision to improve out-of-distribution robustness.
  • Paper: Attention is not Explanation, Sarthak Jain et al. (2019). This study analyzes how raw attention weights often fail to provide faithful explanations or reflect true feature importance, contextualizing the need to explicitly supervise attention distributions.
  • Paper: What Does BERT Look at? An Analysis of BERT’s Attention, Kevin Clark et al. (2019). This work reveals that unsupervised Transformer attention heads naturally attend heavily to uninformative tokens like separators, directly informing why supervising attention away from stop words and punctuation is necessary.
  • Paper: A Decomposable Attention Model for Natural Language Inference, Ankur P. Parikh et al. (2016). This early work introduces attention-based decomposition mechanisms for Natural Language Inference, framing the architectural premise of alignment in text inference.
Cover for Supervising Model Attention with Human Explanations for Robust Natural Language Inference

Abstract

Natural Language Inference (NLI) models are known to learn from biases and artefacts within their training data, impacting how well they generalise to other unseen datasets. Existing de-biasing approaches focus on preventing the models from learning these biases, which can result in restrictive models and lower performance. We instead investigate teaching the model how a human would approach the NLI task, in order to learn features that will generalise better to previously unseen examples. Using natural language explanations, we supervise the model's attention weights to encourage more attention to be paid to the words present in the explanations, significantly improving model performance. Our experiments show that the in-distribution improvements of this method are also accompanied by out-of-distribution improvements, with the supervised models learning from features that generalise better to other NLI datasets. Analysis of the model indicates that human explanations encourage increased attention on the important words, with more attention paid to words in the premise and less attention paid to punctuation and stop-words.

Table of Contents

  • Introduction
  • Related Work
  • Training NLI Models with Explanations
  • Training with Explanations Beyond NLI
  • Creating More Robust NLI Models
  • Attention Supervision Method
  • Supervising Self-Attention Layers
  • Selecting Attention Heads for Supervision
  • Supervising an Additional Attention Layer
  • Experimental Setup and Evaluation
  • Experiments: Performance in and out of Distribution
  • Experiments with DeBERTa
  • Comparing Results with Prior Work
  • Choosing Which Explanations to Use and Which Heads to Supervise
  • Performance when different heads are supervised
  • Analysis
  • Token Level Classification
  • Understanding the Changes in Attention
  • Words Receiving Most Attention
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Explanation-derived token supervision labels

    model/method

    The method converts e-SNLI human explanations into binary token-level supervision for an NLI premise–hypothesis pair. For each of the nn tokens in the concatenated pair, a label ei∈{0,1}e_i\in\{0,1\} indicates whether token ii is relevant to a human explanation. With free-text explanations, ei=1e_i=1 when the token corresponds to a word appearing in the explanation; stopwords are excluded from this matching. With annotator-provided highlighted words, ei=1e_i=1 when the corresponding premise or hypothesis word was highlighted, including highlighted stopwords. The explanation labels are used only during training: at test time, the model predicts the NLI class from the premise and hypothesis without an explanation.

  2. Knowl 2 — Supervising selected final-layer self-attention heads

    equation

    For an NLI token pair with explanation labels ei∈{0,1}e_i\in\{0,1\}, the desired attention distribution is obtained by normalizing the labels:

    di=ei∑k=1nek,i=1,…,n.d_i=\frac{e_i}{\sum_{k=1}^{n}e_k},\qquad i=1,\ldots,n.

    Here, did_i is the target attention assigned to token ii, and nn is the number of tokens in the premise–hypothesis input. For each of HH selected attention heads in the transformer’s final self-attention layer, the model’s attention from the [CLS] query to token ii is aiha_i^h. The training objective is the NLI cross-entropy loss plus a mean-squared attention loss:

    Ltotal=LNLI+λH∑h=1H∑i=1n(aih−di)2,\mathcal{L}_{\mathrm{total}}=\mathcal{L}_{\mathrm{NLI}}+\frac{\lambda}{H}\sum_{h=1}^{H}\sum_{i=1}^{n}(a_i^h-d_i)^2,

    where LNLI\mathcal{L}_{\mathrm{NLI}} is the classification cross-entropy, λ≥0\lambda\geq 0 weights explanation supervision, and HH is the number of supervised heads. For a head hh, the attention value is

    aih=exp⁡((qCLSh)Tkih/dk)∑j=1nexp⁡((qCLSh)Tkjh/dk),a_i^h=\frac{\exp\left((q_{\mathrm{CLS}}^h)^{\mathsf T}k_i^h/\sqrt{d_k}\right)}{\sum_{j=1}^{n}\exp\left((q_{\mathrm{CLS}}^h)^{\mathsf T}k_j^h/\sqrt{d_k}\right)},

    where qCLShq_{\mathrm{CLS}}^h is the query vector for [CLS], kihk_i^h is the key vector for token ii, and dkd_k is the key-vector dimensionality. The explanation loss encourages the selected heads to place attention on words identified by human annotators.

  3. Knowl 3 — Greedy selection of attention heads for supervision

    algorithm

    Because different multi-head attention heads can serve different functions, the method does not necessarily supervise every head. Each head is first supervised individually and evaluated on the NLI development set. The heads with the strongest individual improvements are ranked, and the top KK are supervised jointly. The candidate values are K∈{1,3,6,9,12}K\in\{1,3,6,9,12\}, with five random seeds used for each condition.

    Input: Transformer attention heads, training examples with explanation labels, candidate values K in {1, 3, 6, 9, 12}
    Output: A selected set of attention heads
    For each attention head h:
        Train the model while supervising only head h
        Evaluate development-set NLI accuracy
    Rank heads by their individual development-set accuracy
    For each candidate K:
        Select the K highest-ranked heads
        Train and evaluate the model while supervising those K heads jointly
    Return the K and corresponding head set with the highest development-set accuracy

    This greedy procedure is more efficient than evaluating every subset of heads, but it does not guarantee the globally optimal subset. In the experiments, supervising three heads was ultimately best; supervising all heads improved over no supervision but was worse than supervising the selected three.

  4. Knowl 4 — Supervised additional attention-layer architecture

    model/method

    An alternative method adds a trainable attention layer above the transformer sequence representations hih_i, where hih_i is the representation of token ii and i=1,…,ni=1,\ldots,n. The unnormalized scalar attention score is

    a~i=σ(Wh2tanh⁡(Wh1hi+bh1)+bh2),\widetilde a_i=\sigma\left(W_{h2}\tanh\left(W_{h1}h_i+b_{h1}\right)+b_{h2}\right),

    where Wh1W_{h1} and Wh2W_{h2} are trainable weight matrices, bh1b_{h1} and bh2b_{h2} are trainable bias vectors, and σ\sigma is the sigmoid function. The scores are normalized into attention weights,

    ai=exp⁡(a~i)∑k=1nexp⁡(a~k), a_i=\frac{\exp(\widetilde a_i)}{\sum_{k=1}^{n}\exp(\widetilde a_k)},

    and used to form a pooled representation

    c=∑i=1naihi. c=\sum_{i=1}^{n}a_i h_i.

    A linear classifier followed by softmax predicts the NLI class from cc. The single added attention head is trained with the same NLI loss plus mean-squared error against the explanation-derived target distribution. This approach improves over the baseline, but supervising selected heads in an existing transformer self-attention layer performs better.

  5. Knowl 5 — Training and robustness evaluation protocol

    experimental setup

    The attention-supervision methods were evaluated by fine-tuning BERT and DeBERTa on the e-SNLI NLI corpus, which contains 550,152 training observations in the reported large-scale setting. DeBERTa uses disentangled content and position representations when computing attention. The attention-loss weight λ\lambda was selected on development data from {0.2,0.4,…,1.8}\{0.2,0.4,\ldots,1.8\}; the selected values were λ=1.0\lambda=1.0 for BERT and λ=0.8\lambda=0.8 for DeBERTa. Main BERT results were averaged over 25 random seeds.

    SNLI development and test sets were treated as in-distribution evaluations. Robustness was assessed on SNLI-hard, MNLI matched, MNLI mismatched, ANLI, and HANS, which were treated as out-of-distribution or challenge evaluations. SNLI-hard contains SNLI examples misclassified by a hypothesis-only model, HANS targets failures of common syntactic heuristics, and ANLI contains intentionally difficult human-in-the-loop examples. Improvements over the baseline were assessed with two-tailed tt-tests.

  6. Knowl 6 — Attention supervision improves BERT in- and out-of-distribution accuracy

    data/table

    The main BERT experiment compares ordinary fine-tuning with two explanation-supervised variants: an additional supervised attention layer and supervision of three selected heads in an existing final self-attention layer. Values are mean accuracy over 25 random seeds. A dagger indicates statistical significance at p<0.05p<0.05; a double-dagger indicates significance after Bonferroni correction by a factor of 7.

    Could not parse LaTeX table

    Supervising existing self-attention improves SNLI-test accuracy by 0.400.40 percentage points and SNLI-hard accuracy by 0.790.79 points, while also improving MNLI mismatched, MNLI matched, and HANS by 0.840.84, 0.910.91, and 1.591.59 points, respectively. It does not improve ANLI. The additional-layer method also improves most datasets, but is consistently weaker than supervising existing attention. A control that randomly shuffled the desired attention distribution separately within the premise and hypothesis performed worse than the unsupervised baseline: it reached 89.50%89.50\% on SNLI-test, 78.84%78.84\% on SNLI-hard, 71.5%71.5\% on MNLI mismatched, and 71.23%71.23\% on MNLI matched. This control indicates that the gains are not explained by attention supervision acting only as generic regularization.

  7. Knowl 7 — Comparison with alternative explanation methods and DeBERTa

    empirical result

    The selected-head method improves a strong BERT baseline without adding inference-time parameters. On SNLI, BERT with existing-attention supervision reaches 90.17%90.17\%, compared with 89.77%89.77\% for the BERT baseline. An adaptation of LIREx reaches 90.79%90.79\%, but uses a four-model pipeline with 453 million parameters and loses 0.850.85 points on MNLI; the BERT baseline has 109 million parameters. An adaptation of the Pruthi et al. attention-supervision method improves SNLI by 0.220.22 points, MNLI by 0.870.87 points, and SNLI-hard by 0.540.54 points, whereas the proposed existing-attention method improves them by 0.400.40, 0.880.88, and 0.790.79 points.

    Could not parse LaTeX table

    DeBERTa without explanation supervision achieves 92.59%92.59\% on SNLI. Adding human-explanation attention supervision raises this to 92.69%92.69\%, a statistically significant 0.100.10-point gain and the paper’s reported new state-of-the-art SNLI result. The absolute gain is smaller than BERT’s 0.400.40 points because DeBERTa leaves less performance headroom.

  8. Knowl 8 — Combining explanation sources and selecting three heads gives the best development accuracy

    data/table

    The free-text and highlighted-word annotations provide complementary supervision. Development accuracy was averaged over five random seeds for BERT with existing-attention supervision.

    Could not parse LaTeX table

    When explanations contain only highlighted hypothesis words, the free-text-derived distribution is used as well so that the model receives supervision over both the premise and hypothesis. Combining the two explanation sources performs best. Supervising every attention head improves accuracy over no supervision, but supervising only the three individually strongest heads performs better than either supervising all heads or supervising other individual heads. This supports preserving diversity among heads whose functions do not benefit from being directed toward the same explanation distribution.

  9. Knowl 9 — Supervised attention predicts explanation-highlighted tokens

    empirical result

    The model’s normalized attention weights can be converted into token-level explanation predictions by applying a threshold to the average attention from the three supervised heads. The prediction task is to identify which premise and hypothesis tokens are highlighted in e-SNLI. Precision (PP), recall (RR), and F1 are reported separately for premise and hypothesis tokens.

    Could not parse LaTeX table

    The supervised attention model obtains hypothesis F1 69.1369.13, which is 3.1 points higher than the strongest listed prior hypothesis result, while also substantially improving over the unsupervised baseline. Its attention weights therefore serve not only as a training signal but also as effective token-level predictors of the words humans selected as explanatory.

  10. Knowl 10 — Explanation supervision shifts attention toward informative words and the premise

    data/table

    Analysis of the final [CLS] attention shows that explanation supervision changes what the NLI model attends to. In the baseline, the premise and first [SEP] token together receive only 22.86%22.86\% of attention; when all 12 heads are supervised, they receive 50.89%50.89\%. In the preceding, directly unsupervised layer, premise attention rises from 31.1%31.1\% in the baseline to 54.2%54.2\% with 12 supervised heads, and reaches 54.8%54.8\% when only the selected three top-layer heads are supervised. Thus, supervision propagates toward greater use of the premise, which is especially relevant to SNLI-hard examples designed to expose hypothesis-only behavior.

    The distribution of [CLS] attention across part-of-speech categories, averaged over five seeds, is:

    Could not parse LaTeX table

    The most-attended token in a sentence pair is most often punctuation or a stopword for the baseline, whereas it is usually a content noun for the supervised model:

    Could not parse LaTeX table

    Overall, supervision increases attention to nouns, verbs, and adjectives while reducing attention to punctuation, determiners, adpositions, stopwords, and special tokens. The resulting attention more often identifies important words in both the premise and hypothesis, making the model’s sentence interaction behavior more interpretable.

Coverage note — The paper’s two illustrative attention heatmaps and example-specific qualitative predictions were omitted because they repeat the aggregate attention and token-level findings captured above; no other substantial contributed material was omitted.

References

  1. 1.Andreas, J.; Klein, D.; and Levine, S. 2018. Learning with Latent Language. In NAACL.
  2. 2.Belinkov, Y.; Poliak, A.; Shieber, S.; Van Durme, B.; and Rush, A. 2019a. Don’t Take the Premise for Granted: Mitigating Artifacts in Natural Language Inference. In ACL.
  3. 3.Belinkov, Y.; Poliak, A.; Shieber, S.; Van Durme, B.; and Rush, A. 2019b. On Adversarial Removal of Hypothesis-only Bias in Natural Language Inference. In ACL.
  4. 4.Bowman, S. R.; Angeli, G.; Potts, C.; and Manning, C. D. 2015. A large annotated corpus for learning natural language inference. In EMNLP.
  5. 5.Bujel, K.; Yannakoudakis, H.; and Rei, M. 2021. Zero-shot Sequence Labeling for Transformer-based Sentence Classifiers. In RepL4NLP.
  6. 6.Camburu, O.-M.; Rocktaschel, T.; Lukasiewicz, T.; and Blunsom, P. 2018. e-SNLI: Natural Language Inference with Natural Language Explanations. In NeurIPS.
  7. 7.Clark, C.; Yatskar, M.; and Zettlemoyer, L. 2019. Don’t Take the Easy Way Out: Ensemble Based Methods for Avoiding Known Dataset Biases. In EMNLP-IJCNLP.
  8. 8.Clark, C.; Yatskar, M.; and Zettlemoyer, L. 2020. Learning to Model and Ignore Dataset Bias with Mixed Capacity Ensembles. In EMNLP Findings.
  9. 9.Clark, K.; Khandelwal, U.; Levy, O.; and Manning, C. D. 2019. What Does BERT Look at? An Analysis of BERT’s Attention. In BlackboxNLP@ACL.
  10. 10.Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL.
  11. 11.Gururangan, S.; Swayamdipta, S.; Levy, O.; Schwartz, R.; Bowman, S.; and Smith, N. A. 2018. Annotation Artifacts in Natural Language Inference Data. In NAACL.
  12. 12.Hase, P.; and Bansal, M. 2021. When Can Models Learn From Explanations? A Formal Framework for Understanding the Roles of Explanation Data. arXiv:2102.02201.
  13. 13.He, H.; Zha, S.; and Wang, H. 2019. Unlearn Dataset Bias in Natural Language Inference by Fitting the Residual. In DeepLo@EMNLP.
  14. 14.He, P.; Liu, X.; Gao, J.; and Chen, W. 2021. DeBERTa: Decoding-enhanced BERT with Disentangled Attention. In ICLR.
  15. 15.Kim, Y.; Jang, M.; and Allan, J. 2020. Explaining Text Matching on Neural Natural Language Inference. ACM Trans. Inf. Syst., 38(4).
  16. 16.Kumar, S.; and Talukdar, P. 2020. NILE : Natural Language Inference with Faithful Natural Language Explanations. In ACL. Online.
  17. 17.Lample, G.; Ballesteros, M.; Subramanian, S.; Kawakami, K.; and Dyer, C. 2016. Neural Architectures for Named Entity Recognition. In NAACL.
  18. 18.Liang, W.; Zou, J.; and Yu, Z. 2020. ALICE: Active Learning with Contrastive Natural Language Explanations. In EMNLP.
  19. 19.Liu, T.; Xin, Z.; Ding, X.; Chang, B.; and Sui, Z. 2020. An Empirical Study on Model-agnostic Debiasing Strategies for Robust Natural Language Inference. In CoNLL.
  20. 20.Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  21. 21.Mahabadi, R. K.; Belinkov, Y.; and Henderson, J. 2020. End-to-End Bias Mitigation by Modelling Biases in Corpora. In ACL.
  22. 22.Mahabadi, R. K.; Belinkov, Y.; and Henderson, J. 2021. Variational Information Bottleneck for Effective Low-Resource Fine-Tuning. In ICLR.
  23. 23.Mathew, B.; Saha, P.; Yimam, S. M.; Biemann, C.; Goyal, P.; and Mukherjee, A. 2021. HateXplain: A Benchmark Dataset for Explainable Hate Speech Detection. In AAAI.
  24. 24.McCoy, T.; Pavlick, E.; and Linzen, T. 2019. Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference. In ACL.
  25. 25.Min, J.; McCoy, R. T.; Das, D.; Pitler, E.; and Linzen, T. 2020. Syntactic Data Augmentation Increases Robustness to Inference Heuristics. In ACL.
  26. 26.Minervini, P.; and Riedel, S. 2018. Adversarially Regularising Neural NLI Models to Integrate Logical Background Knowledge. In CoNLL.
  27. 27.Mu, J.; Liang, P.; and Goodman, N. 2020. Shaping Visual Representations with Language for Few-Shot Classification. In ACL.
  28. 28.Murty, S.; Koh, P. W.; and Liang, P. 2020. ExpBERT: Representation Engineering with Natural Language Explanations. In ACL.
  29. 29.Nie, Y.; Williams, A.; Dinan, E.; Bansal, M.; Weston, J.; and Kiela, D. 2020. Adversarial NLI: A New Benchmark for Natural Language Understanding. In ACL.
  30. 30.Pilault, J.; Elhattami, A.; and Pal, C. 2021. Conditionally Adaptive Multi-Task Learning: Improving Transfer Learning in NLP Using Fewer Parameters & Less Data. In ICLR.
  31. 31.Poliak, A.; Naradowsky, J.; Haldar, A.; Rudinger, R.; and Van Durme, B. 2018. Hypothesis Only Baselines in Natural Language Inference. In SEM@NAACL.
  32. 32.Pruthi, D.; Dhingra, B.; Soares, L. B.; Collins, M.; Lipton, Z. C.; Neubig, G.; and Cohen, W. W. 2020. Evaluating Explanations: How much do explanations from the teacher aid students? arXiv:2012.00893.
  33. 33.Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8): 9.
  34. 34.Rajani, N. F.; McCann, B.; Xiong, C.; and Socher, R. 2019. Explain Yourself! Leveraging Language Models for Commonsense Reasoning. In ACL.
  35. 35.Rei, M.; and Søgaard, A. 2018. Zero-Shot Sequence Labeling: Transferring Knowledge from Sentences to Tokens. In Walker, M. A.; Ji, H.; and Stent, A., eds., NAACL.
  36. 36.Rei, M.; and Søgaard, A. 2019. Jointly Learning to Label Sentences and Tokens. In AAAI.
  37. 37.Ribeiro, M.; Singh, S.; and Guestrin, C. 2016. “Why Should I Trust You?”: Explaining the Predictions of Any Classifier. In NAACL.
  38. 38.Sanh, V.; Wolf, T.; Belinkov, Y.; and Rush, A. M. 2020. Learning from others’ mistakes: Avoiding dataset biases without modeling them. In ICLR.
  39. 39.Stacey, J.; Minervini, P.; Dubossarsky, H.; Riedel, S.; and Rocktaschel, T. 2020. Avoiding the Hypothesis-Only Bias in Natural Language Inference via Ensemble Adversarial Training. In EMNLP.
  40. 40.Sun, Z.; Fan, C.; Han, Q.; Sun, X.; Meng, Y.; Wu, F.; and Li, J. 2020. Self-Explaining Structures Improve NLP Models. arXiv:2012.01786.
  41. 41.Teney, D.; Abbasnedjad, E.; and van den Hengel, A. 2020. Learning What Makes a Difference from Counterfactual Examples and Gradient Supervision. arXiv:2004.09034.
  42. 42.Thorne, J.; Vlachos, A.; Christodoulopoulos, C.; and Mittal, A. 2019. Generating Token-Level Explanations for Natural Language Inference. In NAACL.
  43. 43.Tsuchiya, M. 2018. Performance Impact Caused by Hidden Bias of Training Data for Recognizing Textual Entailment. In LREC.
  44. 44.Tu, L.; Lalwani, G.; Gella, S.; and He, H. 2020. An Empirical Study on Robustness to Spurious Correlations using Pre-trained Language Models. TACL, 8: 621–633.
  45. 45.Utama, P. A.; Moosavi, N. S.; and Gurevych, I. 2020a. Mind the Trade-off: Debiasing NLU Models without Degrading the In-distribution Performance. In ACL.
  46. 46.Utama, P. A.; Moosavi, N. S.; and Gurevych, I. 2020b. Towards Debiasing NLU Models from Unknown Biases. In EMNLP. Online.
  47. 47.Vig, J.; and Belinkov, Y. 2019. Analyzing the Structure of Attention in a Transformer Language Model. In BlackboxNLP@ACL.
  48. 48.Williams, A.; Nangia, N.; and Bowman, S. 2018. A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. In NAACL.
  49. 49.Yaghoobzadeh, Y.; Mehri, S.; Tachet des Combes, R.; Hazen, T. J.; and Sordoni, A. 2021. Increasing Robustness to Spurious Correlations using Forgettable Examples. In EACL.
  50. 50.Zhang, Z.; Wu, Y.; Zhao, H.; Li, Z.; Zhang, S.; Zhou, X.; and Zhou, X. 2020. Semantics-Aware BERT for Language Understanding. In AAAI.
  51. 51.Zhao, X.; and Vydiswaran, V. G. V. 2021. LIREx: Augmenting Language Inference with Relevant Explanation. In AAAI.

Citation

MLA
Stacey, J., et al. “Supervising Model Attention with Human Explanations for Robust Natural Language Inference”. arXiv, 2021, http://arxiv.org/abs/2104.08142v3.
APA
Stacey, J., Belinkov, Y., & Rei, M. (2021). Supervising Model Attention with Human Explanations for Robust Natural Language Inference. arXiv. http://arxiv.org/abs/2104.08142v3
Chicago
Stacey, J., Y. Belinkov, and M. Rei. 2021. “Supervising Model Attention with Human Explanations for Robust Natural Language Inference”. arXiv. http://arxiv.org/abs/2104.08142v3.
Harvard
Stacey, J., Belinkov, Y. and Rei, M. (2021) “Supervising Model Attention with Human Explanations for Robust Natural Language Inference”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2104.08142v3.
Vancouver
1. Stacey J, Belinkov Y, Rei M (2021) Supervising Model Attention with Human Explanations for Robust Natural Language Inference. arXiv

BibTeX

@article{stacey2021supervising,
  title = {Supervising Model Attention with Human Explanations for Robust Natural Language Inference},
  author = {Stacey, Joe and Belinkov, Yonatan and Rei, Marek},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2104.08142v3},
  eprint = {2104.08142}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF