Supervising Model Attention with Human Explanations for Robust Natural Language Inference
Joe StaceyYonatan BelinkovMarek Rei
Proposes supervising transformer attention heads with human-provided explanations to simultaneously improve both in-distribution accuracy and out-of-distribution generalization in natural language inference models.
Natural language inference systems determine logical relationships between pairs of sentences, such as whether one statement supports, contradicts, or remains neutral toward another. While standard models perform well on familiar training data, they frequently rely on superficial statistical shortcuts and dataset biases rather than genuine linguistic understanding. This reliance causes them to fail when deployed on new, out-of-distribution text. Traditional debiasing techniques often penalize shortcuts directly, but this regularly degrades baseline accuracy on standard tasks.
The article demonstrates that directly teaching models to follow human reasoning patterns resolves this dilemma. Rather than focusing on suppressing specific biases, the authors supervise the internal attention mechanisms of transformer models using human-written explanations and word highlights from a large benchmark dataset. This guides the system to allocate focus to the specific words human annotators consider essential when establishing semantic relationships.
The authors implemented this approach on standard language architectures, including BERT and DeBERTa, across over 550,000 training examples. The supervision introduces an auxiliary loss that aligns the model's internal attention distribution with human explanations without requiring extra parameters or computational overhead during testing. The authors tested this method across several out-of-distribution challenge sets designed to expose superficial heuristics and hypothesis-only shortcuts.
The key findings show consistent performance and robustness gains across all benchmarks. Supervising the top three attention heads of an existing self-attention layer increased standard test accuracy on the primary benchmark by 0.40% and improved performance on its hardest subset by 0.79%. When applied to the state-of-the-art DeBERTa architecture, the method achieved a new benchmark record of 92.69% accuracy. On out-of-distribution evaluation sets, the approach boosted accuracy by up to 1.59% on heuristic challenge sets and nearly 1% on diverse multi-genre data, avoiding the typical trade-off between standard accuracy and generalizability. Attention analysis revealed that supervised models shifted focus away from punctuation and generic stop-words toward meaningful nouns, verbs, and premise context, more than doubling premise attention in the final layer.
These results demonstrate that incorporating human rationale data makes language models more reliable and interpretable without inflating deployment size or runtime inference costs. By balancing attention across both input sentences, models mitigate hypothesis-only blind spots and maintain high fidelity when encountering unfamiliar linguistic structures in production environments.
Engineering teams deploying language inference models should adopt selective attention supervision using human explanations rather than complex multi-model pipelines. Teams should specifically target a subset of attention heads rather than supervising all heads uniformly to preserve functional diversity across the network. Further research is recommended to expand explanation-guided attention methods to complex reasoning datasets with multi-sentence premises, where performance gains remain limited.
- Paper: Annotation Artifacts in Natural Language Inference Data, Suchin Gururangan et al. (2018). This paper establishes the widespread presence of annotation artifacts and superficial shortcuts in standard NLI datasets, providing the core motivation for debiasing NLI models.
- Paper: Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference, R. Thomas McCoy et al. (2019). This foundational work demonstrates that NLI models exploit shallow syntactic heuristics instead of genuine reasoning, motivating human explanation supervision to improve out-of-distribution robustness.
- Paper: Attention is not Explanation, Sarthak Jain et al. (2019). This study analyzes how raw attention weights often fail to provide faithful explanations or reflect true feature importance, contextualizing the need to explicitly supervise attention distributions.
- Paper: What Does BERT Look at? An Analysis of BERT’s Attention, Kevin Clark et al. (2019). This work reveals that unsupervised Transformer attention heads naturally attend heavily to uninformative tokens like separators, directly informing why supervising attention away from stop words and punctuation is necessary.
- Paper: A Decomposable Attention Model for Natural Language Inference, Ankur P. Parikh et al. (2016). This early work introduces attention-based decomposition mechanisms for Natural Language Inference, framing the architectural premise of alignment in text inference.
- Paper: Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes, Cheng-Yu Hsieh et al. (2023). This work extends the concept of using natural language reasoning steps as direct supervision signals by distilling rationales into compact downstream models.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). This paper investigates unfaithful chain-of-thought explanations swayed by subtle biases, offering a critical downstream perspective on relying on natural language rationales.
- Paper: GlobEnc: Quantifying Global Token Attribution by Incorporating the Whole Encoder Layer in Transformers, Ali Modarressi et al. (2022). This study advances beyond raw attention supervision by quantifying global token attribution across full Transformer encoder layers in NLI models.
- Paper: Discovering and Mitigating Visual Biases Through Keyword Explanation, Younghyun Kim et al. (2024). This research applies the idea of leveraging language explanations and keyword supervision to diagnose and mitigate model biases in computer vision.
