XAI for Transformers: Better Explanations through Conservative Propagation

Ameen AliThomas SchnakeOliver EberleGrégoire MontavonKlaus-Robert MüllerLior Wolf

article2022ICML170 citations

Extends Layer-wise Relevance Propagation to Transformers by identifying and fixing conservation failures in attention heads and LayerNorm layers, producing more reliable feature attributions across text, vision, and graph benchmarks.

Listen

Modern artificial intelligence increasingly relies on transformer models across language processing, computer vision, and scientific data analysis. Because these models contain up to billions of parameters, their internal decision-making is opaque and difficult to verify. In sensitive settings like automated hiring or risk assessment, stakeholders require reliable explainable artificial intelligence (XAI) to verify that decisions are trustworthy, fair, and free of bias. However, existing interpretability tools often struggle with the unique internal mechanics of transformers.

The article evaluates why traditional attribution methods fail on transformers and demonstrates a theoretically sound, conservative propagation technique that significantly improves explanation accuracy across diverse datasets and architectures.

To address this challenge, the authors analyzed standard gradient-based interpretability through the framework of Layer-wise Relevance Propagation (LRP), specifically evaluating whether the axiom of conservation—the requirement that input importance scores sum to the model's total output score—holds across internal layers. After identifying where standard gradient calculations fail mathematically, they introduced modified propagation rules for attention heads and normalization layers. The method was benchmarked across nine datasets covering natural language sentiment and emotion detection, graph-based digit recognition, and molecular property prediction, using input perturbation tests that measure output stability when removing or adding key features.

The investigation produced several key findings. First, standard gradient-based attribution severely violates conservation in transformers, primarily due to attention gating and variance rescaling in layer normalization; on image-based graph benchmarks, naive gradient attributions were almost anticorrelated with model outputs. Second, the proposed method, which treats these gating and normalization scaling factors as locally constant during backpropagation, restored conservation and achieved the highest explanation performance across all tested datasets. For example, on the Stanford Sentiment Treebank dataset, the area under the error curve during feature removal dropped from 2.10 with naive gradient attribution to 1.56 with the combined approach. Third, the technique improved computational efficiency, running in 0.012 seconds per sample on benchmark language data compared to 0.017 seconds for standard gradient methods and 0.024 seconds for the original model prediction. Fourth, applied to bias auditing in sentiment analysis models, the method successfully surfaced specific entity biases (such as disparate sentiment shifts tied to particular names) without requiring synthetic or out-of-distribution test inputs.

These findings indicate that organizations relying on standard gradient or attention visualization tools may be making governance and safety decisions based on flawed or misleading explanations. By ensuring mathematical conservation throughout the network, the proposed approach lowers operational risk and compliance overhead in high-stakes deployments, providing leaders with high-fidelity insights into why a model produces a specific prediction without increasing compute costs.

Organizations deploying transformer models should adopt conservative propagation rules in place of raw attention or standard gradient attribution when auditing model decisions and monitoring for demographic bias. Implementation is straightforward, as the adjustment only requires strategically detaching specific internal terms in the code during the backward attribution pass rather than retraining the underlying models.

While confidence in these findings is reinforced by consistent mathematical derivations and comprehensive empirical validation across language, vision, and molecular tasks, minor conservation gaps remain due to unhandled bias terms in linear layers. Stakeholders should consider this method a robust tool for feature-level attribution and bias detection, while continuing to pair it with broader system-level validation practices.

Cover for XAI for Transformers: Better Explanations through Conservative Propagation

Abstract

Transformers have become an important workhorse of machine learning, with numerous applications. This necessitates the development of reliable methods for increasing their transparency. Multiple interpretability methods, often based on gradient information, have been proposed. We show that the gradient in a Transformer reflects the function only locally, and thus fails to reliably identify the contribution of input features to the prediction. We identify Attention Heads and LayerNorm as the main reasons for such unreliable explanations and propose a more stable way for propagation through these layers. Our proposal, which can be seen as a proper extension of the well-established LRP method to Transformers, is shown both theoretically and empirically to overcome the deficiency of a simple gradient-based approach, and achieves state-of-the-art explanation performance on a broad range of Transformer models and datasets.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. A Theoretical View on Explaining Transformers
  • 3.1. Propagation in Attention Heads
  • 3.2. Propagation in LayerNorm
  • 4. Better LRP Rules for Transformers
  • 5. Experimental Setup
  • 5.1. Datasets
  • 5.2. Benchmark Methods
  • 6. Results
  • 6.1. Conservation
  • 6.2. Quantitative Evaluation
  • 6.3. Qualitative Results
  • 6.4. Use Case: Analyzing Bias in Transformers
  • 7. Conclusion
  • Acknowledgements
  • References
  • A. Derivations for Attention Heads
  • A.1. Gradient of Softmax
  • A.2. Gradient Propagation Rule for Attention Heads
  • A.3. Relevance Conservation
  • B. Derivations for LayerNorm
  • B.1. Centering step
  • B.2. Rescaling step
  • C. Implementation Details
  • D. Additional Conservation Experiments
  • E. Runtime Analysis
  • F. Additional Qualitative Results on SST-2
  • G. Additional Qualitative Results on MNIST Superpixels
  • H. Perturbation Experiments

Knowls

  1. Knowl 1 — Conservative propagation rules for attention and LayerNorm

    model/method

    The proposed Transformer explanation method modifies Gradient × Input (GI) propagation at attention heads and LayerNorm by locally treating their input-dependent gating or scaling factors as constants. In an attention head, let xix_i be the value vector at input position ii, yjy_j the output vector at position jj, and pijp_{ij} the softmax weight on input position ii for output position jj. For each embedding coordinate dd, the attention-head rule redistributes output relevance using the locally linear weights pijp_{ij}:

    R(xi,d)=∑jxi,dpij∑i′xi′,dpi′jR(yj,d).R(x_{i,d})=\sum_j\frac{x_{i,d}p_{ij}}{\sum_{i'}x_{i',d}p_{i'j}}R(y_{j,d}).

    For LayerNorm, let xix_i and yjy_j denote input and output activations at positions ii and jj of a channel, and let NN be the number of normalized positions. The method treats the multiplicative factor α=(ϵ+Var⁡[x])−1/2\alpha=(\epsilon+\operatorname{Var}[x])^{-1/2} as fixed and applies the linear rule associated with centering:

    R(xi,d)=∑jxi,d(δij−1/N)∑i′xi′,d(δi′j−1/N)R(yj,d),R(x_{i,d})=\sum_j\frac{x_{i,d}(\delta_{ij}-1/N)}{\sum_{i'}x_{i',d}(\delta_{i'j}-1/N)}R(y_{j,d}),

    where δij\delta_{ij} is the Kronecker delta and ϵ\epsilon is LayerNorm’s positive stabilizing constant. In both rules, relevance is propagated coordinate-wise; token relevance can be obtained by summing over embedding coordinates. At explanation time, the rules are implemented by detaching the attention probabilities in yj=∑ixipijy_j=\sum_i x_i p_{ij} and the LayerNorm scale factor in yi=(xi−E[x])/ϵ+Var⁡[x]y_i=(x_i-\mathbb{E}[x])/\sqrt{\epsilon+\operatorname{Var}[x]}, then applying standard GI to the modified computation. Detaching attention weights cuts relevance through the query/key gating path while retaining propagation through the value inputs. The paper evaluates attention-only (LRP (AH)), LayerNorm-only (LRP (LN)), and combined (LRP (AH+LN)) variants; other layers use GI-equivalent propagation.

  2. Knowl 2 — Attention heads can violate GI conservation

    theoretical result

    Consider an attention head with input value vectors xix_i, query-side vectors xj′x'_j, outputs yj=∑ixipijy_j=\sum_i x_i p_{ij}, and softmax weights pij=exp⁡(qij)/∑i′exp⁡(qi′j)p_{ij}=\exp(q_{ij})/\sum_{i'}\exp(q_{i'j}), where qijq_{ij} is the matching score between xix_i and xj′x'_j. For a scalar model output ff, define token relevance by R(xi)=xi⊤∂f/∂xiR(x_i)=x_i^\top\partial f/\partial x_i, R(xj′)=xj′⊤∂f/∂xj′R(x'_j)=x_j'^\top\partial f/\partial x'_j, and R(yj)=yj⊤∂f/∂yjR(y_j)=y_j^\top\partial f/\partial y_j. Write Ej\mathbb{E}_j and Cov⁡j\operatorname{Cov}_j for expectation and covariance over input positions ii under the distribution pijp_{ij} for fixed jj; let q:j=(qij)iq_{:j}=(q_{ij})_i. Under the paper’s centering assumptions Ej[q:j]=0\mathbb{E}_j[q_{:j}]=0 and Ej[x]=0\mathbb{E}_j[x]=0 for every jj, GI relevance obeys

    ∑iR(xi)+∑jR(xj′)=∑jR(yj)+2∑jCov⁡j(q:j,x)⊤∂f∂yj.\sum_i R(x_i)+\sum_j R(x'_j)=\sum_j R(y_j)+2\sum_j\operatorname{Cov}_j(q_{:j},x)^\top\frac{\partial f}{\partial y_j}.

    Thus GI is not conservative across this attention-head component when the covariance correction is nonzero. The result identifies input-dependent attention gating as a source of relevance imbalance; the centering assumptions are used for this simplified form.

  3. Knowl 3 — LayerNorm causes GI relevance collapse

    theoretical result

    For the core centering-and-standardization part of LayerNorm, excluding its subsequent affine transformation, let xix_i be the input activations over the NN positions of a channel and let yi=(xi−E[x])/ϵ+Var⁡[x]y_i=(x_i-\mathbb{E}[x])/\sqrt{\epsilon+\operatorname{Var}[x]}. Here E[x]\mathbb{E}[x] and Var⁡[x]\operatorname{Var}[x] are the mean and variance across those activations, and ϵ>0\epsilon>0 is the stabilizing constant. With GI relevance R(xi)=xi ∂f/∂xiR(x_i)=x_i\,\partial f/\partial x_i and R(yi)=yi ∂f/∂yiR(y_i)=y_i\,\partial f/\partial y_i for scalar output ff, the relevance sums satisfy

    ∑iR(xi)=(1−Var⁡[x]ϵ+Var⁡[x])∑iR(yi).\sum_i R(x_i)=\left(1-\frac{\operatorname{Var}[x]}{\epsilon+\operatorname{Var}[x]}\right)\sum_i R(y_i).

    For positive variance, the factor is less than one, so GI relevance is not conserved through this operation; the discrepancy is especially strong when ϵ\epsilon is small relative to the variance. The paper calls this loss of relevance a “relevance collapse.”

  4. Knowl 4 — GI as a layer-wise relevance propagation rule

    equation

    For a neural-network component with scalar input activations xix_i, scalar output activations yjy_j, and scalar network output ff, Gradient × Input assigns R(xi)=xi ∂f/∂xiR(x_i)=x_i\,\partial f/\partial x_i and R(yj)=yj ∂f/∂yjR(y_j)=y_j\,\partial f/\partial y_j. Using the chain rule, the same attribution can be written as a relevance propagation rule:

    R(xi)=∑j∂yj∂xixiyjR(yj),R(x_i)=\sum_j\frac{\partial y_j}{\partial x_i}\frac{x_i}{y_j}R(y_j),

    with the convention 0/0=00/0=0. This representation lets GI be analyzed as an LRP-style propagation method: a component conserves relevance when ∑iR(xi)=∑jR(yj)\sum_iR(x_i)=\sum_jR(y_j), and conservation at every component implies global conservation.

  5. Knowl 5 — Models, datasets, and perturbation evaluation protocol

    experimental setup

    The experiments cover text classification, image classification represented as graphs, and molecular graphs. Text tasks use Transformer models on SST-2 (11,844 movie reviews) and IMDB (50,000) for binary sentiment; TweetEval for sentiment (59,899 examples), hate detection (12,970), and emotion classification (5,052); and SILICONE for Semaine emotion detection (13,708) and Meld-S utterance sentiment analysis (5,627). The graph experiments use Graphormer on MNIST superpixels (70,000 digit examples, with image patches represented as connected graph nodes) and BACE (1,522 labeled compounds). For SST-2 and IMDB, the models initialize embeddings and tokenizers from pretrained BERT checkpoints and are trained with batch size 32, AdamW, learning rate 2×10−52\times10^{-5}, and at most 20 epochs with early stopping; the other NLP tasks use the same settings with pretrained bert-base-uncased. The two-layer Graphormer uses batch size 64 and AdamW at learning rate 2×10−42\times10^{-4}, trained for 1,000 epochs on BACE and 10 on MNIST.

    Explanations are evaluated by adding or removing input elements in relevance order. For text activation, an empty sequence of “UNK” tokens is progressively replaced by original tokens from highest to lowest relevance; for graph activation, the most relevant nodes are added first. The metric AUAC is the area under the correct-class probability curve, with higher values indicating that relevant inputs activate the correct prediction earlier and more strongly. For pruning, the least-relevant elements are removed first (replaced by “UNK” for text); AU-MSE is the area under the squared difference between the original model logits y0y_0 and logits ymty_{m_t} after the step-tt mask mtm_t. Lower AU-MSE indicates that removing low-relevance inputs changes the prediction less.

  6. Knowl 6 — Perturbation benchmarks favor combined conservative propagation

    data/table

    The table reports activation performance (AUAC; higher is better) and pruning performance (AU-MSE; lower is better) for nine datasets. Activation adds relevant inputs first, while pruning removes least-relevant inputs first; the perturbation procedures and metric definitions are given here so the scores are interpretable. LRP (AH+LN), which applies both proposed rules, has the highest AUAC and lowest AU-MSE among the reported methods on every dataset. Raw-attention baselines and GI are generally weaker, though individual methods can perform well on particular metrics. A dash means the paper did not report that result.

    Method IMDB SST-2 BACE MNIST T-Emotions T-Hate T-Sentiment Meld-S Semaine
    AUAC (higher is better)
    Random .673 .664 .624 .324 .516 .640 .484 .460 .432
    A-Last .708 .712 .620 .862 .542 .663 .515 .483 .451
    A-Flow – .711 .637 – – – – – –
    Rollout .738 .713 .653 .358 .554 .659 .520 .489 .441
    GAE .872 .821 .675 .426 .675 .762 .611 .548 .532
    GI .920 .847 .646 .942 .652 .772 .651 .591 .529
    LRP (AH) .911 .855 .645 .942 .675 .797 .668 .594 .544
    LRP (LN) .935 .907 .702 .947 .735 .829 .710 .632 .593
    LRP (AH+LN) .939 .908 .707 .948 .750 .838 .713 .635 .606
    AU-MSE (lower is better)
    Random 2.16 3.97 1.95 69.82 4.25 9.12 2.87 2.54 1.92
    A-Last 1.65 2.56 1.99 45.82 3.73 7.77 1.90 1.74 1.42
    A-Flow – 2.52 1.87 – – – – – –
    Rollout 1.04 2.43 1.77 115.2 2.85 6.55 1.71 1.53 1.40
    GAE 1.63 2.26 1.66 59.81 2.21 7.40 1.61 1.56 1.37
    GI 0.87 2.10 2.06 18.06 2.09 6.69 1.41 1.57 1.43
    LRP (AH) 0.77 2.02 2.08 18.03 1.83 6.43 1.43 1.69 1.38
    LRP (LN) 0.69 1.78 1.65 17.55 1.55 5.02 1.25 1.50 1.13
    LRP (AH+LN) 0.65 1.56 1.61 17.49 1.47 4.88 1.23 1.48 1.08
  7. Knowl 7 — Proposed explanations track model outputs more conservatively

    empirical result

    The authors compared GI and combined LRP (AH+LN) on Transformer models trained for SST-2 and MNIST, plotting each example’s model output score against the sum of its input-feature relevance scores. The combined method’s points lay much closer to the equality diagonal than GI’s, although some conservation error remained; the authors suggest non-attributable biases in linear layers as a possible source. On MNIST, GI relevance sums were almost anticorrelated with output scores. Additional SST-2 tests found that GAE, Rollout, and Attention Flow also did not satisfy conservation. These observations empirically support the theoretical concern that common explanation methods need not preserve output attribution in Transformer models.

  8. Knowl 8 — Token-level explanations expose entity-specific sentiment effects

    empirical result

    For a use case, the authors applied the proposed attribution method to a publicly available DistilBERT checkpoint fine-tuned for SST-2 sentiment classification. They examined token relevance for the difference between positive and negative output scores, rather than inferring effects from name occurrence alone. The attributed relevance distributions showed no consistent difference between female and male names, despite more male names occurring in the dataset. They did show variation across entity categories: common Western male names such as “lee,” “barry,” and “coen” were among those most strongly associated with positive sentiment, while “chan,” “saddam hussein,” and “castro” were among names with strong negative impact. “Sally jesse raphael” ranked highly in part because of the surname “raphael.” The analysis demonstrates how token-wise attribution can separate a word’s modeled contribution from other sentence content; it is an analysis of this model and dataset, not a causal finding about the named groups.

  9. Knowl 9 — Qualitative explanations focus on task-relevant words and pixels

    empirical result

    In SST-2 sentence examples, the proposed LRP explanations highlighted sentiment-bearing words such as “best” and “virtues” and assigned less relevance to the entity token “eastwood” than the last-layer attention baseline, which emphasized that name more strongly. In Graphormer explanations for MNIST superpixels, LRP (AH) and LRP (AH+LN) more reliably highlighted pixels forming the digit than the attention-based methods, which often assigned relevance to background superpixels. The paper’s visual comparisons also show that naive GI can produce distorted relevance patterns relative to the proposed combined rule. These are qualitative illustrations of the same relevance-specificity differences measured in the perturbation experiments.

  10. Knowl 10 — Explanation runtime is competitive with GI

    empirical result

    On SST-2, the authors measured average wall-clock time in seconds per example for explanation generation. Random attribution took 6×10−56\times10^{-5} seconds, Attention-last 0.0002, Rollout 0.0004, GAE 0.016, GI 0.017, LRP (AH) 0.014, and LRP (AH+LN) 0.012; model prediction took 0.024 seconds. Thus, the combined method was faster than GI and GAE in this measurement. Attention Flow and LRP (LN) are not included in the reported timing values.

Coverage note — Omitted proof-only derivative calculations and supplementary repeated visualization panels; they support or illustrate the stated results without adding a distinct contributed result.

References

  1. 1.Abnar, S. and Zuidema, W. H. Quantifying attention flow in transformers. In Jurafsky, D., Chai, J., Schluter, N., and Tetreault, J. R. (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, pp. 4190–4197, 2020.
  2. 2.Aleisa, M. A., Beloff, N., and White, M. Airm: a new ai recruiting model for the saudi arabia labor market. In Arai, K. (ed.), Intelligent Systems Conference (IntelliSys) 2021, volume 3: 296 of Lecture Notes in Networks and Systems, pp. 105–124, Cham, September 2021. Springer.
  3. 3.Ancona, M., Ceolini, E., Oztireli, C., and Gross, M. H. Gradient-based attribution methods. In Explainable AI: Interpreting, Explaining and Visualizing Deep Learning, volume 11700 of Lecture Notes in Computer Science, pp. 169–191. Springer, 2019.
  4. 4.Arras, L., Arjona-Medina, J., Widrich, M., Montavon, G., Gillhofer, M., Muller, K.-R., Hochreiter, S., and Samek, W. Explaining and Interpreting LSTMs. In Samek, W., Montavon, G., Vedaldi, A., Hansen, L. K., and Muller, K.-R. (eds.), Explainable AI: Interpreting, Explaining and Visualizing Deep Learning, volume 11700 of Lecture Notes in Computer Science, pp. 211–238. Springer, Cham, 2019.
  5. 5.Arrieta, A. B., Rodríguez, N. D., Ser, J. D., Bennetot, A., Tabik, S., Barbado, A., García, S., Gil-Lopez, S., Molina, D., Benjamins, R., Chatila, R., and Herrera, F. Explainable artificial intelligence (XAI): concepts, taxonomies, opportunities and challenges toward responsible AI. Information Fusion, 58:82–115, 2020.
  6. 6.Atanasova, P., Simonsen, J. G., Lioma, C., and Augenstein, I. A diagnostic study of explainability techniques for text classification. In Webber, B., Cohn, T., He, Y., and Liu, Y. (eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pp. 3256–3274. Association for Computational Linguistics, 2020.
  7. 7.Bach, S., Binder, A., Montavon, G., Klauschen, F., Muller, K.-R., and Samek, W. On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. PLoS ONE, 10(7):e0130140, 2015.
  8. 8.Bahdanau, D., Cho, K., and Bengio, Y. Neural machine translation by jointly learning to align and translate. In Bengio, Y. and LeCun, Y. (eds.), 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, 2015.
  9. 9.Bolukbasi, T., Chang, K.-W., Zou, J., Saligrama, V., and Kalai, A. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. In Proceedings of the 30th International Conference on Neural Information Processing Systems, pp. 4356–4364, Red Hook, NY, USA, 2016. Curran Associates Inc.
  10. 10.Chapuis, E., Colombo, P., Manica, M., Labeau, M., and Clavel, C. Hierarchical pre-training for sequence labelling in spoken dialog. In Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 2636–2648, 2020.
  11. 11.Chefer, H., Gur, S., and Wolf, L. Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 397–406, 2021a.
  12. 12.Chefer, H., Gur, S., and Wolf, L. Transformer interpretability beyond attention visualization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 782–791, 2021b.
  13. 13.Danilevsky, M., Qian, K., Aharonov, R., Katsis, Y., Kawas, B., and Sen, P. A survey of the state of explainable AI for natural language processing. In AACL/IJCNLP, pp. 447–459. Association for Computational Linguistics, 2020.
  14. 14.De-Arteaga, M., Romanov, A., Wallach, H., Chayes, J., Borgs, C., Chouldechova, A., Geyik, S., Kenthapadi, K., and Kalai, A. T. Bias in bios: A case study of semantic representation bias in a high-stakes setting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, pp. 120–128, New York, NY, USA, 2019. Association for Computing Machinery.
  15. 15.Denil, M., Demiraj, A., and de Freitas, N. Extraction of salient sentences from labelled documents. ArXiv, abs/1412.6815, 2014.
  16. 16.Devlin, J., Chang, M., Lee, K., and Toutanova, K. BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, pp. 4171–4186, 2019.
  17. 17.Ding, Y., Liu, Y., Luan, H., and Sun, M. Visualizing and understanding neural machine translation. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, pp. 1150–1159, 2017.
  18. 18.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. In 9th International Conference on Learning Representations, ICLR 2021, 2021.
  19. 19.Feng, S., Wallace, E., Grissom II, A., Iyyer, M., Rodriguez, P., and Boyd-Graber, J. Pathologies of neural models make interpretations difficult. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 3719–3728, Brussels, Belgium, October-November 2018. Association for Computational Linguistics.
  20. 20.Gonen, H. and Goldberg, Y. Lipstick on a pig: Debiasing methods cover up systematic gender biases in word embeddings but do not remove them. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 609–614, 2019.
  21. 21.Guidotti, R., Monreale, A., Ruggieri, S., Turini, F., Giannotti, F., and Pedreschi, D. A survey of methods for explaining black box models. ACM Comput. Surv., 51(5):93:1–93:42, 2019.
  22. 22.Hesse, R., Schaub-Meyer, S., and Roth, S. Fast axiomatic attribution for neural networks. In Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, 2021.
  23. 23.Hollenstein, N. and Beinborn, L. Relative importance in sentence processing. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, pp. 141–150, 2021.
  24. 24.Huang, Q., Yamada, M., Tian, Y., Singh, D., Yin, D., and Chang, Y. Graphlime: Local interpretable model explanations for graph neural networks. arXiv preprint arXiv:2001.06216, 2020.
  25. 25.Jain, S. and Wallace, B. C. Attention is not Explanation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 3543–3556, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics.
  26. 26.Kiritchenko, S. and Mohammad, S. Examining gender and race bias in two hundred sentiment analysis systems. In Proceedings of the Seventh Joint Conference on Lexical and Computational Semantics, pp. 43–53, 2018.
  27. 27.Lundberg, S. M. and Lee, S. A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, pp. 4765–4774, 2017.
  28. 28.Luo, D., Cheng, W., Xu, D., Yu, W., Zong, B., Chen, H., and Zhang, X. Parameterized explainer for graph neural network. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, volume 33, 2020.
  29. 29.Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pp. 142–150, 2011.
  30. 30.Maziarka, Ł., Danel, T., Mucha, S., Rataj, K., Tabor, J., and Jastrzkebski, S. Molecule attention transformer. arXiv preprint arXiv:2002.08264, 2020.
  31. 31.Montavon, G. Gradient-based vs. propagation-based explanations: An axiomatic comparison. In Explainable AI: Interpreting, Explaining and Visualizing Deep Learning, pp. 253–265. Springer International Publishing, 2019.
  32. 32.Montavon, G., Samek, W., and Muller, K. Methods for interpreting and understanding deep neural networks. Digit. Signal Process., 73:1–15, 2018.
  33. 33.Monti, F., Boscaini, D., Masci, J., Rodola, E., Svoboda, J., and Bronstein, M. M. Geometric deep learning on graphs and manifolds using mixture model cnns. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, pp. 5115–5124, 2017.
  34. 34.Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V., Vainbrand, D., Kashinkunti, P., Bernauer, J., Catanzaro, B., Phanishayee, A., and Zaharia, M. Efficient large-scale language model training on GPU clusters using megatron-lm. In SC ’21: The International Conference for High Performance Computing, Networking, Storage and Analysis, 2021.
  35. 35.Ousidhoum, N., Zhao, X., Fang, T., Song, Y., and Yeung, D.-Y. Probing toxic content in large pre-trained language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, pp. 4262–4274. Association for Computational Linguistics, 2021.
  36. 36.Pope, P. E., Kolouri, S., Rostami, M., Martin, C. E., and Hoffmann, H. Explainability methods for graph convolutional neural networks. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10772–10781, 2019.
  37. 37.Prabhakaran, V., Hutchinson, B., and Mitchell, M. Perturbation sensitivity analysis to detect unintended model biases. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP, pp. 5739–5744, 2019.
  38. 38.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI blog, 2019.
  39. 39.Ribeiro, M. T., Wu, T. S., Guestrin, C., and Singh, S. Beyond accuracy: Behavioral testing of nlp models with checklist. In Proc. Association for Computational Linguistics (ACL), pp. 4902–4912, 2020.
  40. 40.Samek, W., Montavon, G., Vedaldi, A., Hansen, L. K., and Muller, K.-R. (eds.). Explainable AI: Interpreting, Explaining and Visualizing Deep Learning, volume 11700 of Lecture Notes in Computer Science. Springer, 2019.
  41. 41.Samek, W., Montavon, G., Lapuschkin, S., Anders, C. J., and Muller, K.-R. Explaining deep neural networks and beyond: A review of methods and applications. Proceedings of the IEEE, 109(3):247–278, 2021.
  42. 42.Sanh, V., Debut, L., Chaumond, J., and Wolf, T. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. ArXiv, abs/1910.01108, 2019.
  43. 43.Schnake, T., Eberle, O., Lederer, J., Nakajima, S., Schutt, K. T., Muller, K.-R., and Montavon, G. Higher-order explanations of graph neural networks via relevant walks. IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 10.1109/TPAMI.2021.3115452, 2021.
  44. 44.Shapley, L. S. A Value for n-Person Games, pp. 307–318. Princeton University Press, 1953.
  45. 45.Shrikumar, A., Greenside, P., and Kundaje, A. Learning important features through propagating activation differences. In Proceedings of the 34th International Conference on Machine Learning - Volume 70, ICML’17, pp. 3145–3153, 2017.
  46. 46.Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pp. 1631–1642, 2013.
  47. 47.Srinivas, S. and Fleuret, F. Full-gradient representation for neural network visualization. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, pp. 4126–4135, 2019.
  48. 48.Strumbelj, E. and Kononenko, I. An efficient explanation of individual classifications using game theory. J. Mach. Learn. Res., 11:1–18, 2010.
  49. 49.Subramanian, G., Ramsundar, B., Pande, V., and Denny, R. A. Computational modeling of β-secretase 1 (bace-1) inhibitors using ligand based approaches. Journal of chemical information and modeling, 56(10):1936–1949, 2016.
  50. 50.Sundararajan, M., Taly, A., and Yan, Q. Axiomatic attribution for deep networks. In ICML, volume 70 of Proceedings of Machine Learning Research, pp. 3319–3328. PMLR, 2017.
  51. 51.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. In Advances in neural information processing systems, pp. 5998–6008, 2017.
  52. 52.Voita, E., Talbot, D., Moiseev, F., Sennrich, R., and Titov, I. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 5797–5808, 2019.
  53. 53.Wallace, E., Tuyls, J., Wang, J., Subramanian, S., Gardner, M., and Singh, S. AllenNLP interpret: A framework for explaining predictions of NLP models. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP): System Demonstrations, pp. 7–12, 2019.
  54. 54.Wu, Z. and Ong, D. C. On explaining your explanations of bert: An empirical study with sequence classification. CoRR, abs/2101.00196, 2021.
  55. 55.Wu, Z., Ramsundar, B., Feinberg, E. N., Gomes, J., Geniesse, C., Pappu, A. S., Leswing, K., and Pande, V. Moleculenet: a benchmark for molecular machine learning. Chemical science, 9(2):513–530, 2018.
  56. 56.Xiong, W., Wu, J., Wang, H., Kulkarni, V., Yu, M., Guo, X., Chang, S., and Wang, W. Y. Tweetqa: A social media focused question answering dataset. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 5020–5031, 2019.
  57. 57.Ying, C., Cai, T., Luo, S., Zheng, S., Ke, G., He, D., Shen, Y., and Liu, T.-Y. Do transformers really perform badly for graph representation? In Thirty-Fifth Conference on Neural Information Processing Systems, 2021.
  58. 58.Ying, R., Bourgeois, D., You, J., Zitnik, M., and Leskovec, J. Gnnexplainer: Generating explanations for graph neural networks. Advances in neural information processing systems, pp. 9240–9251, 2019.
  59. 59.Yoo, S.-Y., Kim, Y.-S., Lee, K., Jeong, K., Choi, J., Lee, H., and Choi, Y. S. Graph-aware transformer: Is attention all graphs need? ArXiv, abs/2006.05213, 2020.
  60. 60.Yun, S., Jeong, M., Kim, R., Kang, J., and Kim, H. J. Graph transformer networks. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, pp. 11960–11970, 2019.
  61. 61.Zhao, J., Li, C., Wen, Q., Wang, Y., Liu, Y., Sun, H., Xie, X., and Ye, Y. Gophormer: Ego-graph transformer for node classification. arXiv preprint arXiv:2110.13094, 2021.

Citation

MLA
Ali, A., et al. “XAI for Transformers: Better Explanations Through Conservative Propagation”. International Conference on Machine Learning, vol. 162, 2022, pp. 435–51, https://proceedings.mlr.press/v162/ali22a.html.
APA
Ali, A., Schnake, T., Eberle, O., Montavon, G., Müller, K.-R., & Wolf, L. (2022). XAI for Transformers: Better Explanations through Conservative Propagation. International Conference on Machine Learning, 162, 435–451. https://proceedings.mlr.press/v162/ali22a.html
Chicago
Ali, A., T. Schnake, O. Eberle, G. Montavon, K.-R. Müller, and L. Wolf. 2022. “XAI for Transformers: Better Explanations Through Conservative Propagation”. International Conference on Machine Learning 162: 435–51. https://proceedings.mlr.press/v162/ali22a.html.
Harvard
Ali, A. et al. (2022) “XAI for Transformers: Better Explanations through Conservative Propagation”, International Conference on Machine Learning. PMLR, pp. 435–451. Available at: https://proceedings.mlr.press/v162/ali22a.html.
Vancouver
1. Ali A, Schnake T, Eberle O, Montavon G, Müller K-R, Wolf L (2022) XAI for Transformers: Better Explanations through Conservative Propagation. In: International Conference on Machine Learning. PMLR, pp 435–451

BibTeX

@InProceedings{pmlr-v162-ali22a,
  title = 	 {{XAI} for Transformers: Better Explanations through Conservative Propagation},
  author =       {Ali, Ameen and Schnake, Thomas and Eberle, Oliver and Montavon, Gr{\'e}goire and M{\"u}ller, Klaus-Robert and Wolf, Lior},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {435--451},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/ali22a/ali22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/ali22a.html},
  abstract = 	 {Transformers have become an important workhorse of machine learning, with numerous applications. This necessitates the development of reliable methods for increasing their transparency. Multiple interpretability methods, often based on gradient information, have been proposed. We show that the gradient in a Transformer reflects the function only locally, and thus fails to reliably identify the contribution of input features to the prediction. We identify Attention Heads and LayerNorm as main reasons for such unreliable explanations and propose a more stable way for propagation through these layers. Our proposal, which can be seen as a proper extension of the well-established LRP method to Transformers, is shown both theoretically and empirically to overcome the deficiency of a simple gradient-based approach, and achieves state-of-the-art explanation performance on a broad range of Transformer models and datasets.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/