"Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification

Jasmijn BastingsSebastian EbertPolina ZablotskaiaAnders SandholmKatja Filippova

article2022EMNLP88 citations

Establishes a rigorous evaluation protocol using synthetic shortcut injection to benchmark the faithfulness of input salience methods, revealing that popular explanation techniques often fail to detect even simple lexical patterns used by text classifiers.

Listen

Modern natural language processing models often achieve high test accuracy by learning superficial shortcuts or spurious correlations rather than true linguistic patterns, leading to severe failures when deployed in real-world scenarios. While input salience methods—techniques that highlight the most important words influencing a model's prediction—are widely used for model debugging, different methods frequently yield contradictory explanations for the exact same input. The article establishes an objective, ground-truth evaluation protocol to determine how faithfully common salience methods identify known lexical shortcuts across various text classification architectures and datasets.

To establish an unambiguous ground truth, the authors augmented real text classification datasets (SST-2, IMDB, and Wikipedia Toxicity) with synthetic lexical shortcuts. These shortcuts spanned three complexity levels: single predictive tokens, contextual token combinations, and ordered token pairs. The evaluation benchmarked four primary classes of explainability techniques—raw Gradient, Gradient times Input, Integrated Gradients, and LIME—across diverse mathematical configurations on standard recurrent (LSTM) and transformer (BERT) models. The faithfulness of each method was measured by its precision in ranking true shortcut tokens at the top and the average ranking depth required to capture all shortcut tokens.

The findings reveal that explainability performance depends heavily on the model architecture and configuration choices, rather than the general algorithmic category alone. For BERT models, simple raw gradient norm methods achieved near-perfect precision (0.99 or higher in seven of nine configurations) and average ranks of 1 to 2, outperforming far more complex techniques. Conversely, Gradient times Input performed strongly on LSTM models (achieving up to 1.0 precision on single tokens) but failed severely on BERT (falling to 0.29–0.59 precision). Furthermore, configuration details proved critical: reducing gradient vectors via mean averaging rather than vector norms caused precision to drop sharply from nearly 1.0 to roughly 0.4 on BERT. Integrated Gradients gained little benefit from increasing interpolation steps from 100 to 1,000, and its performance with a zero baseline collapsed to match single-step Gradient times Input.

These results demonstrate that practitioner assumptions about explainability methods do not generalize across neural architectures or shortcut types. Relying on complex, computationally expensive techniques like Integrated Gradients does not guarantee faithful debugging, and evaluation on simple single-token shortcuts fails to predict performance on multi-token or contextual rules. For practitioners auditing BERT-based classifiers for lexical shortcuts, standard raw gradient norm methods should be prioritized as both a reliable and computationally inexpensive choice. When applying perturbation-based methods like LIME, increasing perturbation counts up to 1,000 and using unknown-token masking provides substantial accuracy gains over token erasure or mask tokens.

While the study provides strong, validated conclusions for English binary text classification on LSTM and BERT models, its scope is limited to lexical shortcuts and specific attribution techniques. Input salience represents an inherently local view of importance that cannot fully capture complex multi-feature interactions or broader internal model logic. Future work should expand the benchmarking protocol to larger foundation models, non-lexical linguistic shortcuts, and advanced feature-interaction attribution methods before establishing definitive debugging standards across broader production systems.

arXiv: 2111.07367
Cover for "Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification

Abstract

A common approach for explaining predictions made by neural networks

is to identify salient features using gradients or attention.

In natural language processing specifically,

there have been several approaches proposed

to compute gradient-based token importance scores;

yet their faithfulness remains unclear:

e.g., do they agree where important tokens occur?

different formulations may yield completely different explanations,

making comparisons across papers ambiguous.

Existing evaluation practice often relies solely on intuition—both when designing new explanation techniques

and reporting results.We propose experiments called “shortcut” tasks specially crafted such that we know which spurious patterns could exist in data and how easily models would pick them up instead of meaningful signals.Empirically, neither commonly used variants of input salience methods nor attention pass our tests consistently.This calls into question prior work based entirely on automatic metrics.Finally, we suggest simple practical recommendations relating design choices in evaluating model behavior to obtainmore faithful attributions.The code base supporting all analyses reported here hasbeen released alongwith an interactive online tool accompanyingthis paper.Both links accessibleat https://github.com/google-research/ \allowbreak google-research/tree/master/saliency_faith fulness .Thus,in summary,the contributionsofourpaperare( )a setoffive shortcuttasksincontrolledsettingsofdifferentcomplexities;(ii)

asystematicstudyofgradientbasedsaliencymethodsacrossthese five settings,w.r.t.multipletextclassifiersincludingstateoftheartmodelsfine tunedonfourlanguage datasets,and(iii)a concrete proposaltoaugmentexistingbenchmarksingle-taskmetricsbyadditionalmulti taskediagnosticsdrivenbyeitheroftheshortcutdatasetsintroducedorrealisticdomainspecificpatternsaswe exemplifyfortoxicitydetection.Wenotethatexplanationsamplersemustbecausally relatedtothedecisionprocess,butalsomusto beclearly communicatedtothestakeholder,i.e.acontentprovideroranend-user-inthatsenseouremphasisisonanalyzingthemethodratherthanjustperformance.Hence,intheabsenceofaformalframeworkforevaluatingex planation qualityweproposeaprotocolthatisolatesbehavioraldifferenceswhile encouragingtheadoptionofcommonscoresfordistinctclasses ofexplanation methods

Table of Contents

  • 1 Introduction
  • 2 Methodology
  • 2.1 Protocol
  • 2.2 Shortcut Types
  • 2.3 Creating (Partially) Synthetic Data
  • 2.4 Verification Steps
  • 3 Experimental Setup
  • 3.1 Models
  • 3.2 Salience Methods
  • 3.2.1 Gradient
  • 3.2.2 Gradient times Input
  • 3.2.3 Integrated Gradients
  • 3.2.4 LIME
  • 4 Results
  • 5 Related Work
  • 6 Conclusions
  • 7 Limitations
  • References
  • A Appendix
  • A.1 Architecture and training details
  • A.2 Budget
  • A.3 Methods implementation details
  • A.4 Rank results
  • A.5 Full Results

Knowls

  1. Knowl 1 — Partially synthetic protocol for testing salience faithfulness

    algorithm

    The protocol evaluates whether an input-salience method helps a developer discover a specified shortcut in a text classifier. It proceeds as follows: (1) define a shortcut type and its decision rule; (2) augment a real training dataset with synthetic examples governed by that rule, and create a fully synthetic test set in which every example contains the shortcut; (3) train the same model architecture on the original and augmented training sets, checking that their performance remains comparable on the unmodified source test set; (4) verify that the augmented-data model uses the shortcut by testing it on the fully synthetic set and checking that the original-data model performs at chance there; (5) obtain token-salience rankings for examples from the augmented-data model; and (6) compare the highest-ranked tokens against the known shortcut tokens using precision and rank measures. The known shortcut is treated as ground-truth importance only after the verification tests support that interpretation.

  2. Knowl 2 — Precision-at-k and mean-rank measures

    equation

    Let DD be a set of examples, mm a trained classifier, and ss a salience method that ranks the input tokens for a prediction. For each example x∈Dx\in D, let G(x)G(x) be the known set of shortcut tokens, and let top⁡r(s,m,x)\operatorname{top}_r(s,m,x) denote the first rr tokens in the ranking. If the shortcut has size kk, precision-at-kk is the average fraction of those kk positions occupied by ground-truth shortcut tokens:

    P@k(s)=1k∣D∣∑x∈D∣top⁡k(s,m,x)∩G(x)∣.P@k(s)=\frac{1}{k|D|}\sum_{x\in D}\left|\operatorname{top}_k(s,m,x)\cap G(x)\right|.

    Mean rank is the average number of positions that must be examined to include every ground-truth shortcut token:

    R(s)=1∣D∣∑x∈Dmin⁡{r: G(x)⊆top⁡r(s,m,x)}.R(s)=\frac{1}{|D|}\sum_{x\in D}\min\left\{r:\ G(x)\subseteq\operatorname{top}_r(s,m,x)\right\}.

    Higher P@kP@k and lower RR are preferable. Precision measures how many important tokens appear immediately, whereas rank measures how far down the list a reader must look to find all of them. For example, on the Toxicity ordered-pair task, GRAD-L2 and GRAD-mean have similar precision (0.560.56 and 0.600.60), but their mean ranks are 1313 and 2121, respectively. For the experiments, k=1k=1 for single-token shortcuts and k=2k=2 for the two-token shortcuts. Rankings were adjusted where needed so positive salience indicated contribution toward the prediction.

  3. Knowl 3 — Three lexical shortcut rules

    definition

    The experiments define three lexical shortcuts using special indicator tokens absent from the original datasets. In the single-token case (st), the presence of an indicator such as #0 or #1 determines the binary label. In the token-in-context case (tic), an indicator determines the label only when a separate context token is also present; the indicator alone is not predictive. In the ordered-pair case (op), the order of #0 and #1 determines the label: #0 before #1 indicates class 0, and #1 before #0 indicates class 1; neither indicator alone predicts the label. The tic and op implementations constrain the relevant special tokens to be no more than 50 tokens apart.

  4. Knowl 4 — Construction and verification of shortcut data

    model/method

    Shortcut tokens are chosen to be absent from the source data, making their label association unambiguous. To create a synthetic training example, the procedure samples a source example, inserts the shortcut tokens at randomly selected positions (using a random order where order defines the rule), and assigns the label specified by the shortcut. The resulting mixed training set is 20% larger than its source version. To make multi-token shortcut examples less isolated from ordinary inputs, the tic and op settings also insert one of the rule’s indicator tokens into source examples with probability 0.250.25, leaving those examples’ original labels unchanged. A fully synthetic test set contains a shortcut in every example and labels determined by the shortcut. Shortcut use is checked by requiring the mixed-data model to reach close to 100% accuracy on this synthetic test set, while the model trained only on source data should be at chance; the latter check indicates that the shortcut is needed to predict labels in the synthetic test set.

  5. Knowl 5 — Datasets, classifiers, and shortcut-use checks

    experimental setup

    The evaluation combines three binary classification datasets with a bidirectional LSTM and BERT. SST2 is balanced sentiment data with inputs averaging about 20 tokens; IMDB is balanced sentiment data with inputs about ten times longer than SST2; Toxicity contains variable-length Wikipedia comments, with 9% labeled toxic. The LSTM uses pretrained GloVe embeddings of size 300, a bidirectional LSTM with hidden size 256, and a keys-only attention classifier followed by a one-layer MLP. Its word-dropout rate is 0.1 and dropout is 0.5; it is trained with SGD (learning rate 0.03, momentum 0.9, weight decay 5×10−65\times10^{-6}), batches of 64 for SST2/IMDB and 32 for Toxicity, and a maximum of 35,000 steps. BERT is the uncased Base model (12 layers, 12 heads, hidden size 768), fine-tuned with dropout 0.5, Adam, weight decay 5×10−65\times10^{-6}, batches of 16, and maximum sequence lengths of 100 for SST2 and 500 for IMDB and Toxicity; the learning rate is 2×10−52\times10^{-5} except for Toxicity, where it is 10−510^{-5}. Both setups use early stopping after 10,000 steps without validation improvement.

    On the unmodified source test sets, LSTM accuracy is 87.8% on SST2, 91.9% on IMDB, and 92.5% on Toxicity; BERT accuracy is 93.1%, 93.5%, and 93.2%, respectively. Training on shortcut-augmented data causes mean source-test accuracy drops of 0.6, 0.0, and 0.0 percentage points for LSTM, and 0.6, 0.2, and 0.2 points for BERT, in the same dataset order. Across the nine dataset–shortcut combinations, mixed-data models achieve minimum/mean synthetic-test accuracy of 99.8%/99.95% for LSTM and 99.7%/99.91% for BERT. Source-only models achieve 50% on those synthetic tests.

  6. Knowl 6 — Salience methods and configuration choices evaluated

    model/method

    The comparison covers four salience-method classes and a random baseline. For an input of nn token embeddings x1,…,xn∈Rdx_1,\ldots,x_n\in\mathbb{R}^d, a target class cc, logit fc(x1:n)f_c(x_{1:n}), and probability p(c∣x1:n)=σ(fc(x1:n))p(c\mid x_{1:n})=\sigma(f_c(x_{1:n})), gradient salience computes ∇xig(x1:n)\nabla_{x_i}g(x_{1:n}) for each token, where gg is either the logit or probability. The vector is reduced to a scalar using its L1L_1 norm, L2L_2 norm, or componentwise mean. Gradient-times-input (GxI) instead scores token ii by ∇xig(x1:n)⋅xi\nabla_{x_i}g(x_{1:n})\cdot x_i, again using either logit or probability.

    Integrated Gradients (IG) averages gradients along a straight-line interpolation from a baseline embedding sequence b1:nb_{1:n} to the input and takes the dot product with the embedding difference. For mm interpolation steps, the score for token ii is

    IG⁡i=1m∑q=1m∇xig ⁣(b1:n+qm(x1:n−b1:n))⋅(xi−bi).\operatorname{IG}_i=\frac{1}{m}\sum_{q=1}^{m}\nabla_{x_i}g\!\left(b_{1:n}+\frac{q}{m}(x_{1:n}-b_{1:n})\right)\cdot(x_i-b_i).

    Here qq indexes the interpolation steps; tested baselines include zero, UNK, PAD, and, for BERT, [MASK], and tested step counts are 100 and 1000. The target gg can be a logit or probability. LIME fits a locally weighted linear surrogate to the classifier’s predictions on randomly masked or erased versions of an input; the perturbation masks are surrogate features. It uses cosine distance with an exponential kernel of width 25, leaves beginning- and end-of-sequence tokens unchanged, and tests 100, 1000, or 3000 perturbations. Perturbations use UNK for LSTM and BERT, [MASK] for BERT, or token erasure for both architectures.

  7. Knowl 7 — Gradient-method performance depends strongly on architecture

    empirical result

    For the precision results below, each triplet is ordered as single token (st), token in context (tic), and ordered pair (op); the three triplets are ordered as SST2, IMDB, and Toxicity. BERT’s GRAD-L2 configuration performs strongly across all nine settings: precision is (0.99,0.99,1.00)(0.99,0.99,1.00) on SST2, (0.99,0.87,0.96)(0.99,0.87,0.96) on IMDB, and (0.99,0.99,1.00)(0.99,0.99,1.00) on Toxicity. Its corresponding mean ranks are (1,2,2)(1,2,2) for each dataset. Using logits rather than probabilities, or L1 rather than L2 reduction, does not materially change these BERT results.

    LSTM instead favors GxI: its precision is (1.00,0.76,0.92)(1.00,0.76,0.92) on SST2, (1.00,0.35,0.81)(1.00,0.35,0.81) on IMDB, and (1.00,0.68,0.88)(1.00,0.68,0.88) on Toxicity. LSTM GRAD-L2 precision is (0.95,0.50,0.51)(0.95,0.50,0.51), (1.00,0.50,0.52)(1.00,0.50,0.52), and (0.37,0.54,0.56)(0.37,0.54,0.56) across those datasets. Thus the best-performing gradient-based method depends on architecture, and success on a single-token shortcut does not establish success on multi-token shortcuts. The authors suggest that BERT residual connections may make its gradients less noisy than LSTM gradients, but present this as a possible explanation rather than a demonstrated cause. GRAD-mean is weak for BERT, with precision ranging from 0.34 to 0.46.

  8. Knowl 8 — Integrated Gradients is sensitive to baseline, not usually step count

    empirical result

    Across the tested models and datasets, increasing Integrated Gradients (IG) from 100 to 1000 interpolation steps generally changes precision by no more than 3%. For BERT, the zero-baseline IG scores match the GxI scores in the reported comparisons, so increasing the integration steps does not improve on that baseline’s results. Baseline choice has a larger effect: the best tested BERT IG configuration uses logits with a [MASK] baseline. With 1000 steps, its precision is (0.71,0.58,0.71)(0.71,0.58,0.71) on SST2, (0.99,0.64,0.61)(0.99,0.64,0.61) on IMDB, and (0.70,0.50,0.47)(0.70,0.50,0.47) on Toxicity, where each triplet is ordered st, tic, op. Even this relatively strong IG configuration remains below BERT GRAD-L2, whose precision is at least 0.87 in every setting. The experiments therefore show that IG performance is configuration-dependent and that more interpolation steps alone do not generally resolve weak performance.

  9. Knowl 9 — LIME benefits from more perturbations up to a plateau

    empirical result

    For LIME with UNK masking on BERT, increasing the perturbation budget from 100 to 1000 improves precision in examples that include longer inputs or more complex shortcuts: IMDB ordered-pair precision rises from 0.44 to 0.70, and Toxicity ordered-pair precision rises from 0.51 to 0.75. Raising the budget again from 1000 to 3000 produces smaller gains in those cases, to 0.71 and 0.78, respectively; SST2 token-in-context precision changes from 0.80 to 0.87 to 0.88 across the same budgets. UNK masking also outperforms [MASK] in almost all tested configurations. For example, at 3000 perturbations on BERT, UNK versus [MASK] precision is 0.88 versus 0.62 on SST2 tic, 0.71 versus 0.67 on IMDB op, and 0.78 versus 0.70 on Toxicity op. Erasing tokens gives worse precision on average than masking.

  10. Knowl 10 — Scope and interpretability limitations

    limitation

    The experiments cover English binary text classification, three lexical shortcut types, two model families (LSTM and BERT), and a selected set of popular salience methods. Results may differ for other shortcut structures, tasks, languages, architectures, or model sizes; the protocol is intended to be reusable in those settings, but the reported measurements do not establish how methods will perform there. The evaluation also does not cover newer methods designed to capture feature interactions. More generally, token salience highlights do not reveal the classifier’s full decision logic or interactions between input features, so they cannot by themselves provide a complete account of a nonlinear model’s prediction.

Coverage note — The exhaustive per-configuration precision and rank matrices are not reproduced; representative exact scores and the main cross-configuration patterns are included instead, since the full matrices are repetitive. No other substantial contributed material is deliberately omitted.

References

  1. 1.Julius Adebayo, Michael Muelly, Harold Abelson, and Been Kim. 2022. Post hoc explanations may be ineffective for detecting unknown spurious correlation. In International Conference on Learning Representations.
  2. 2.Julius Adebayo, Michael Muelly, Ilaria Liccardi, and Been Kim. 2020. Debugging tests for model explanations. In NeurIPS.
  3. 3.Leila Arras, Ahmed Osman, Klaus-Robert Müller, and Wojciech Samek. 2019. Evaluating recurrent neural network explanations. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 113–126, Florence, Italy. Association for Computational Linguistics.
  4. 4.Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2020. A diagnostic study of explainability techniques for text classification. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3256–3274, Online. Association for Computational Linguistics.
  5. 5.Sebastian Bach, A. Binder, G. Montavon, F. Klauschen, Klaus-Robert Müller, and W. Samek. 2015. On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. PLoS ONE, 10(7).
  6. 6.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473.
  7. 7.Jasmijn Bastings and Katja Filippova. 2020. The elephant in the interpretability room: Why use attention as explanation when we have saliency methods? In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 149–155, Online. Association for Computational Linguistics.
  8. 8.Yonatan Belinkov, Adam Poliak, Stuart Shieber, Benjamin Van Durme, and Alexander Rush. 2019. Don’t take the premise for granted: Mitigating artifacts in natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 877–891, Florence, Italy. Association for Computational Linguistics.
  9. 9.Oana-Maria Camburu, Eleonora Giunchiglia, Jakob Foerster, Thomas Lukasiewicz, and Phil Blunsom. 2019. Can I trust the explainer? verifying post-hoc explanatory methods. In NeurIPS 2019 Workshop on Safety and Robustness in Decision Making, Vancouver, Canada.
  10. 10.Oana-Maria Camburu, Eleonora Giunchiglia, Jakob Foerster, Thomas Lukasiewicz, and Phil Blunsom. 2020. The struggles of feature-based explanations: Shapley values vs. minimal sufficient subsets.
  11. 11.Zhihong Chen, Yan Song, Tsung-Hui Chang, and Xiang Wan. 2020. Generating radiology reports via memory-driven transformer. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1439–1449, Online. Association for Computational Linguistics.
  12. 12.Noel Codella, Veronica Rotemberg, Philipp Tschandl, M. Emre Celebi, Stephen Dusza, David Gutman, Brian Helba, Aadi Kalloo, Konstantinos Liopyris, Michael Marchetti, Harald Kittler, and Allan Halpern. 2019. Skin lesion analysis toward melanoma detection 2018: A challenge hosted by the international skin imaging collaboration (isic).
  13. 13.Misha Denil, Alban Demiraj, and Nando de Freitas. 2015. Extraction of salient sentences from labelled documents.
  14. 14.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  15. 15.Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2020. ERASER: A benchmark to evaluate rationalized NLP models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4443–4458, Online. Association for Computational Linguistics.
  16. 16.Shuoyang Ding and Philipp Koehn. 2021. Evaluating saliency methods for neural language models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5034–5052, Online. Association for Computational Linguistics.
  17. 17.Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2018. Measuring and mitigating unintended bias in text classification. In AAAI/ACM Conference on AI, Ethics, and Society.
  18. 18.Mengnan Du, Varun Manjunatha, Rajiv Jain, Ruchi Deshpande, Franck Dernoncourt, Jiuxiang Gu, Tong Sun, and Xia Hu. 2021. Towards interpreting and mitigating shortcut learning behavior of NLU models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 915–929, Online. Association for Computational Linguistics.
  19. 19.Shi Feng and Jordan Boyd-Graber. 2019. What can ai do for me? evaluating machine learning interpretations in cooperative play. In Proceedings of the 24th International Conference on Intelligent User Interfaces, IUI ’19, page 229–239, New York, NY, USA. Association for Computing Machinery.
  20. 20.R. Geirhos, J.-H. Jacobsen, C. Michaelis, R. Zemel, W. Brendel, M. Bethge, and F. A. Wichmann. 2020. Shortcut learning in deep neural networks. Nature Machine Intelligence, 2:665––673.
  21. 21.Mor Geva, Yoav Goldberg, and Jonathan Berant. 2019. Are we modeling the task or the annotator? an investigation of annotator bias in natural language understanding datasets. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1161–1166, Hong Kong, China. Association for Computational Linguistics.
  22. 22.Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018. Annotation artifacts in natural language inference data. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 107–112, New Orleans, Louisiana. Association for Computational Linguistics.
  23. 23.Xiaochuang Han, Byron C. Wallace, and Yulia Tsvetkov. 2020. Explaining black box predictions and unveiling data artifacts through influence functions. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5553–5563, Online. Association for Computational Linguistics.
  24. 24.Yiding Hao. 2020. Evaluating attribution methods using white-box LSTMs. In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 300–313, Online. Association for Computational Linguistics.
  25. 25.Sara Hooker, Dumitru Erhan, Pieter jan Kindermans, and Been Kim. 2018. Evaluating feature importance estimates. arXiv.
  26. 26.Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim. 2019. A benchmark for interpretability methods in deep neural networks. In Advances in Neural Information Processing Systems, volume 32, pages 9737–9748. Curran Associates, Inc.
  27. 27.Maximilian Idahl, Lijun Lyu, Ujwal Gadiraju, and Avishek Anand. 2021. Towards benchmarking the utility of explanations for model debugging. In Proceedings of the First Workshop on Trustworthy Natural Language Processing, pages 68–73, Online. Association for Computational Linguistics.
  28. 28.Alon Jacovi and Yoav Goldberg. 2020. Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4198–4205, Online. Association for Computational Linguistics.
  29. 29.Sarthak Jain and Byron C. Wallace. 2019. Attention is not Explanation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3543–3556, Minneapolis, Minnesota. Association for Computational Linguistics.
  30. 30.Divyansh Kaushik, Eduard H. Hovy, and Zachary Chase Lipton. 2020. Learning the difference that makes a difference with counterfactually-augmented data. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  31. 31.Siwon Kim, Jihun Yi, Eunji Kim, and Sungroh Yoon. 2020. Interpretation of NLP models through input marginalization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3154–3167, Online. Association for Computational Linguistics.
  32. 32.Pieter-Jan Kindermans, Sara Hooker, Julius Adebayo, Maximilian Alber, Kristof T. Schütt, Sven Dähne, Dumitru Erhan, and Been Kim. 2017. The (un)reliability of saliency methods.
  33. 33.Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky. 2016. Visualizing and understanding neural models in NLP. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 681–691, San Diego, California. Association for Computational Linguistics.
  34. 34.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 142–150, Portland, Oregon, USA. Association for Computational Linguistics.
  35. 35.Andreas Madsen, Nicholas Meade, Vaibhav Adlakha, and Siva Reddy. 2021. Evaluating the faithfulness of importance measures in nlp by recursively masking allegedly important tokens and retraining.
  36. 36.Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019. Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3428–3448, Florence, Italy. Association for Computational Linguistics.
  37. 37.Pramod Kaushik Mudrakarta, Ankur Taly, Mukund Sundararajan, and Kedar Dhamdhere. 2018. Did the model understand the question? In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1896–1906, Melbourne, Australia. Association for Computational Linguistics.
  38. 38.John Pavlopoulos, Prodromos Malakasiotis, and Ion Androutsopoulos. 2017. Deeper attention to abusive user content moderation. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 1125–1135, Copenhagen, Denmark. Association for Computational Linguistics.
  39. 39.Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. Glove: Global vectors for word representation. In Empirical Methods in Natural Language Processing (EMNLP), pages 1532–1543.
  40. 40.Pouya Pezeshkpour, Sarthak Jain, Sameer Singh, and Byron C. Wallace. 2021. Combining feature and instance attribution to detect artifacts.
  41. 41.Nina Poerner, Hinrich Schütze, and Benjamin Roth. 2018. Evaluating neural network explanation methods using hybrid documents and morphosyntactic agreement. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 340–350, Melbourne, Australia. Association for Computational Linguistics.
  42. 42.Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018. Hypothesis only baselines in natural language inference. In Proceedings of the Seventh Joint Conference on Lexical and Computational Semantics, pages 180–191, New Orleans, Louisiana. Association for Computational Linguistics.
  43. 43.Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016. "why should I trust you?": Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, August 13-17, 2016, pages 1135–1144.
  44. 44.Shachar Rosenman, Alon Jacovi, and Yoav Goldberg. 2020. Exposing Shallow Heuristics of Relation Extraction Models with Challenge Data. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3702–3710, Online. Association for Computational Linguistics.
  45. 45.Alexis Ross, Ana Marasovic, and Matthew Peters. 2021. Explaining NLP models via minimal contrastive editing (MiCE). In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3840–3852, Online. Association for Computational Linguistics.
  46. 46.Wojciech Samek and Klaus-Robert Müller. 2019. Towards Explainable Artificial Intelligence, pages 5–22. Springer International Publishing, Cham.
  47. 47.Philipp Schmidt and Felix Biessmann. 2019. Quantifying interpretability and trust in machine learning systems.
  48. 48.Mike Schuster, Kuldip K. Paliwal, and A. General. 1997. Bidirectional recurrent neural networks. IEEE Transactions on Signal Processing.
  49. 49.Sandipan Sikdar, Parantapa Bhattacharya, and Kieran Heese. 2021. Integrated directional gradients: Feature interaction attribution for neural NLP models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 865–878, Online. Association for Computational Linguistics.
  50. 50.Jacob Sippy, Gagan Bansal, and Daniel S. Weld. 2020. Data staining: A method for comparing faithfulness of explainers. In The Fifth Annual Workshop on Human Interpretability in Machine Learning (WHI 2020).
  51. 51.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 1631–1642, Seattle, Washington, USA. Association for Computational Linguistics.
  52. 52.Julia Strout, Ye Zhang, and Raymond Mooney. 2019. Do human rationales improve machine explanations? In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 56–62, Florence, Italy. Association for Computational Linguistics.
  53. 53.M. Sundararajan, Jinhua Xu, Ankur Taly, R. Sayres, and A. Najmi. 2019. Exploring principled visualizations for deep network attributions. In IUI Workshops.
  54. 54.Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017a. Axiomatic attribution for deep networks. In Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017, volume 70 of Proceedings of Machine Learning Research, pages 3319–3328. PMLR.
  55. 55.Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017b. Axiomatic attribution for deep networks. In Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pages 3319–3328, International Convention Centre, Sydney, Australia. PMLR.
  56. 56.Ian Tenney, James Wexler, Jasmijn Bastings, Tolga Bolukbasi, Andy Coenen, Sebastian Gehrmann, Ellen Jiang, Mahima Pushkarna, Carey Radebaugh, Emily Reif, and Ann Yuan. 2020. The language interpretability tool: Extensible, interactive visualizations and analysis for NLP models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 107–118, Online. Association for Computational Linguistics.
  57. 57.Julia K. Winkler, Christine Fink, Ferdinand Toberer, Alexander Enk, Teresa Deinlein, Rainer Hofmann-Wellenhof, Luc Thomas, Aimilios Lallas, Andreas Blum, Wilhelm Stolz, and Holger A. Haenssle. 2019. Association Between Surgical Skin Markings in Dermoscopic Images and Diagnostic Performance of a Deep Learning Convolutional Neural Network for Melanoma Recognition. JAMA Dermatology, 155(10):1135–1141.
  58. 58.Ellery Wulczyn, Nithum Thain, and Lucas Dixon. 2017. Ex machina: Personal attacks seen at scale. In Proceedings of the 26th International Conference on World Wide Web, WWW ’17, pages 1391–1399, Republic and Canton of Geneva, CHE. International World Wide Web Conferences Steering Committee.
  59. 59.Mengjiao Yang and Been Kim. 2019. Benchmarking Attribution Methods with Relative Feature Importance. CoRR, abs/1907.09701.
  60. 60.Yinchong Yang, Volker Tresp, Marius Wunderle, and Peter A. Fasching. 2018. Explaining therapy predictions with layer-wise relevance propagation in neural networks. In 2018 IEEE International Conference on Healthcare Informatics (ICHI), pages 152–162.
  61. 61.Chih-Kuan Yeh, Cheng-Yu Hsieh, Arun Suggala, David I Inouye, and Pradeep K Ravikumar. 2019. On the (in)fidelity and sensitivity of explanations. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  62. 62.Fan Yin, Zhouxing Shi, Cho-Jui Hsieh, and Kai-Wei Chang. 2022. On the sensitivity and stability of model interpretations in NLP. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2631–2647, Dublin, Ireland. Association for Computational Linguistics.
  63. 63.Yilun Zhou, Serena Booth, Marco Tílio Ribeiro, and Julie Shah. 2021. Do feature attribution methods correctly attribute features? CoRR, abs/2104.14403.

Citation

MLA
Bastings, J., et al. ““Will You Find These Shortcuts?” A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 976–91, https://doi.org/10.18653/v1/2022.emnlp-main.64.
APA
Bastings, J., Ebert, S., Zablotskaia, P., Sandholm, A., & Filippova, K. (2022). “Will You Find These Shortcuts?” A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 976–991. https://doi.org/10.18653/v1/2022.emnlp-main.64
Chicago
Bastings, J., S. Ebert, P. Zablotskaia, A. Sandholm, and K. Filippova. 2022. ““Will You Find These Shortcuts?” A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 976–91. https://doi.org/10.18653/v1/2022.emnlp-main.64.
Harvard
Bastings, J. et al. (2022) ““Will You Find These Shortcuts?” A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 976–991. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.64.
Vancouver
1. Bastings J, Ebert S, Zablotskaia P, Sandholm A, Filippova K (2022) “Will You Find These Shortcuts?” A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 976–991

BibTeX

@inproceedings{bastings-etal-2022-will,
    title = "``Will You Find These Shortcuts?'' A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification",
    author = "Bastings, Jasmijn  and
      Ebert, Sebastian  and
      Zablotskaia, Polina  and
      Sandholm, Anders  and
      Filippova, Katja",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.64/",
    doi = "10.18653/v1/2022.emnlp-main.64",
    pages = "976--991"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/