Built independently by an author, for readers. Read the story and support ChapterPal

keyword

single-token shortcut

A single-token shortcut is a type of spurious correlation in natural language processing where a machine learning model relies heavily on the presence of an individual word or subword token to make a prediction rather than learning the broader context or true semantic meaning of the text. This phenomenon occurs when a single token is strongly associated with a specific target class in the training data, leading the model to adopt this simple lexical cue as a heuristic for decision-making. In model evaluation and interpretability research, single-token shortcuts are often used as minimal, controlled test cases to assess whether feature attribution and input salience methods can faithfully identify the precise features influencing a model decisions.

1 item

"Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification

"Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification

Jasmijn Bastings, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm, Katja Filippova

Why you should read this

Establishes a rigorous evaluation protocol using synthetic shortcut injection to benchmark the faithfulness of input salience methods, revealing that popular explanation techniques often fail to detect even simple lexical patterns used by text classifiers.

A common approach for explaining predictions made by neural networks is to identify salient features using gradients or attention. In natural language processing specifically, there have been several approaches proposed to compute gradient-based token importance scores; yet their faithfulness remains unclear: e.g., do they agree where important tokens occur? different formulations may yield completely different explanations, making comparisons across papers ambiguous. Existing evaluation practice often relies solely on intuition—both when designing new explanation techniques and reporting results.We propose experiments called “shortcut” tasks specially crafted such that we know which spurious patterns could exist in data and how easily models would pick them up instead of meaningful signals.Empirically, neither commonly used variants of input salience methods nor attention pass our tests consistently.This calls into question prior work based entirely on automatic metrics.Finally, we suggest simple practical recommendations relating design choices in evaluating model behavior to obtainmore faithful attributions.The code base supporting all analyses reported here hasbeen released alongwith an interactive online tool accompanyingthis paper.Both links accessibleat https://github.com/google-research/ \allowbreak google-research/tree/master/saliency\_faith fulness .Thus,in summary,the contributionsofourpaperare( )a setoffive shortcuttasksincontrolledsettingsofdifferentcomplexities;(ii) asystematicstudyofgradientbasedsaliencymethodsacrossthese five settings,w.r.t.multipletextclassifiersincludingstateoftheartmodelsfine tunedonfourlanguage datasets,and(iii)a concrete proposaltoaugmentexistingbenchmarksingle-taskmetricsbyadditionalmulti taskediagnosticsdrivenbyeitheroftheshortcutdatasetsintroducedorrealisticdomainspecificpatternsaswe exemplifyfortoxicitydetection.Wenotethatexplanationsamplersemustbecausally relatedtothedecisionprocess,butalsomusto beclearly communicatedtothestakeholder,i.e.acontentprovideroranend-user-inthatsenseouremphasisisonanalyzingthemethodratherthanjustperformance.Hence,intheabsenceofaformalframeworkforevaluatingex planation qualityweproposeaprotocolthatisolatesbehavioraldifferenceswhile encouragingtheadoptionofcommonscoresfordistinctclasses ofexplanation methods

Added

2026-10-02