"Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification
Jasmijn BastingsSebastian EbertPolina ZablotskaiaAnders SandholmKatja Filippova
Establishes a rigorous evaluation protocol using synthetic shortcut injection to benchmark the faithfulness of input salience methods, revealing that popular explanation techniques often fail to detect even simple lexical patterns used by text classifiers.
Modern natural language processing models often achieve high test accuracy by learning superficial shortcuts or spurious correlations rather than true linguistic patterns, leading to severe failures when deployed in real-world scenarios. While input salience methods—techniques that highlight the most important words influencing a model's prediction—are widely used for model debugging, different methods frequently yield contradictory explanations for the exact same input. The article establishes an objective, ground-truth evaluation protocol to determine how faithfully common salience methods identify known lexical shortcuts across various text classification architectures and datasets.
To establish an unambiguous ground truth, the authors augmented real text classification datasets (SST-2, IMDB, and Wikipedia Toxicity) with synthetic lexical shortcuts. These shortcuts spanned three complexity levels: single predictive tokens, contextual token combinations, and ordered token pairs. The evaluation benchmarked four primary classes of explainability techniques—raw Gradient, Gradient times Input, Integrated Gradients, and LIME—across diverse mathematical configurations on standard recurrent (LSTM) and transformer (BERT) models. The faithfulness of each method was measured by its precision in ranking true shortcut tokens at the top and the average ranking depth required to capture all shortcut tokens.
The findings reveal that explainability performance depends heavily on the model architecture and configuration choices, rather than the general algorithmic category alone. For BERT models, simple raw gradient norm methods achieved near-perfect precision (0.99 or higher in seven of nine configurations) and average ranks of 1 to 2, outperforming far more complex techniques. Conversely, Gradient times Input performed strongly on LSTM models (achieving up to 1.0 precision on single tokens) but failed severely on BERT (falling to 0.29–0.59 precision). Furthermore, configuration details proved critical: reducing gradient vectors via mean averaging rather than vector norms caused precision to drop sharply from nearly 1.0 to roughly 0.4 on BERT. Integrated Gradients gained little benefit from increasing interpolation steps from 100 to 1,000, and its performance with a zero baseline collapsed to match single-step Gradient times Input.
These results demonstrate that practitioner assumptions about explainability methods do not generalize across neural architectures or shortcut types. Relying on complex, computationally expensive techniques like Integrated Gradients does not guarantee faithful debugging, and evaluation on simple single-token shortcuts fails to predict performance on multi-token or contextual rules. For practitioners auditing BERT-based classifiers for lexical shortcuts, standard raw gradient norm methods should be prioritized as both a reliable and computationally inexpensive choice. When applying perturbation-based methods like LIME, increasing perturbation counts up to 1,000 and using unknown-token masking provides substantial accuracy gains over token erasure or mask tokens.
While the study provides strong, validated conclusions for English binary text classification on LSTM and BERT models, its scope is limited to lexical shortcuts and specific attribution techniques. Input salience represents an inherently local view of importance that cannot fully capture complex multi-feature interactions or broader internal model logic. Future work should expand the benchmarking protocol to larger foundation models, non-lexical linguistic shortcuts, and advanced feature-interaction attribution methods before establishing definitive debugging standards across broader production systems.
- Paper: Sanity Checks for Saliency Maps, Julius Adebayo et al. (2018). Its sanity checks establish why saliency maps must be tested for dependence on learned parameters and data before they can be trusted for model debugging.
- Paper: “Why Should I Trust You?”: Explaining the Predictions of Any Classifier, Marco Tulio Ribeiro et al. (2016). Its introduction of LIME provides the methodological foundation for understanding the perturbation-based attribution method benchmarked in the source.
- Paper: Axiomatic Attribution for Deep Networks, Mukund Sundararajan et al. (2017). Its axioms and definition of Integrated Gradients clarify the attribution method whose configurations and faithfulness the source evaluates.
- Paper: Shortcut learning in deep neural networks, Robert Geirhos et al. (2020). Its account of shortcut learning explains the spurious-correlation problem that motivates the source’s controlled evaluation of shortcut detection.
- Paper: Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference, R. Thomas McCoy et al. (2019). Its controlled demonstrations of linguistic heuristics in NLI provide concrete context for the source’s focus on shortcut reliance in text classifiers.
- Paper: Attention is not Explanation, Sarthak Jain et al. (2019). Its tests of whether attention reflects word importance establish a key faithfulness concern that motivates evaluating salience methods against ground truth.
- Paper: AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers, Reduan Achtibat et al. (2024). It extends attribution faithfulness evaluation to large transformers with attention-aware relevance propagation, testing broader models and internal representations than the source.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). It carries the source’s concern with explanation faithfulness into language-model rationales, testing whether generated reasoning reveals the cues that actually drive predictions.
- Paper: Faithfulness Tests for Natural Language Explanations, Pepa Atanasova et al. (2023). It extends faithfulness testing from token salience to natural-language explanations through counterfactual and input-reconstruction diagnostics.
- Paper: Not All Neuro-Symbolic Concepts Are Created Equal: Analysis and Mitigation of Reasoning Shortcuts, Emanuele Marconato et al. (2023). It broadens shortcut analysis from lexical cues to reasoning shortcuts in neuro-symbolic predictors and evaluates strategies for mitigating them.
