AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers

Reduan AchtibatSayed Mohammad Vakilzadeh HatefiMaximilian DreyerAakriti JainThomas WiegandSebastian LapuschkinWojciech Samek

article2024ICML163 citations

Develops an efficient Layer-wise Relevance Propagation framework tailored for attention mechanisms and non-linear transformer components, enabling faithful attribution of both input features and latent representations in a single backward pass.

Listen

Large transformer architectures, such as Large Language Models and Vision Transformers, are widely used across industries but remain vulnerable to bias and hallucinations. Because these models operate as opaque "black boxes," organizations face compliance, safety, and operational risks when deploying them in high-stakes environments. Existing interpretability methods are either too computationally expensive—requiring massive compute resources for iterative perturbations—or produce unfaithful, noisy explanations that fail to handle the complex non-linear components inside transformers.

The article introduces and evaluates AttnLRP, an attention-aware extension of Layer-wise Relevance Propagation designed to provide highly faithful, holistic explanations for transformer models. Its primary objective is to accurately attribute model predictions back to both input tokens and internal latent representations within the computational efficiency of a single backward pass.

To achieve this, the authors mathematically formulated relevance propagation rules within the Deep Taylor Decomposition framework specifically tailored to non-linear operations, such as softmax, matrix multiplication, and normalization layers. They evaluated AttnLRP across multiple architectures, including LLaMa 2, Mixtral 8x7b, Flan-T5-XL, and Vision Transformers, utilizing standard datasets such as ImageNet, IMDB reviews, Wikipedia, and SQuAD v2. Faithfulness was rigorously assessed by measuring performance drops when perturbing model inputs based on assigned importance scores, and computational efficiency was benchmarked against existing perturbation and gradient-based approaches.

The key findings demonstrate that AttnLRP consistently outperforms competing explainability techniques across text and vision domains. In language models, AttnLRP achieved superior faithfulness scores, such as an area metric of 10.93 on Wikipedia next-word prediction compared to 7.85 for conservative propagation baselines and negative or negligible scores for standard gradient methods. In complex architectures with routing and non-linear feed-forward layers, like Mixtral 8x7b, AttnLRP improved top-1 attribution accuracy to 0.96, representing a 46% gain over prior conservative propagation methods. Furthermore, the approach requires only constant computational complexity relative to token length, avoiding the exponential time and energy costs associated with linear perturbation methods. Finally, the authors demonstrated that AttnLRP can pinpoint specific internal "knowledge neurons," enabling targeted model editing to manipulate outputs systematically without full retraining.

These findings indicate that organizations can explain and audit large-scale transformer models faithfully with minimal computational and financial overhead. High-efficiency attribution enables real-time interpretability, systematic debugging of hallucinations, and safer model deployment in regulated sectors such as healthcare and finance. The ability to identify and edit latent concepts also provides a viable path for steering model behavior and mitigating biases directly within internal layers.

Moving forward, practitioners and technical leaders seeking deeper transparency should adopt backpropagation-based frameworks like AttnLRP over noisy gradient or costly perturbation methods. Future implementation work should focus on optimizing custom graphics processing unit kernels, analyzing the impact of low-bit quantization on attribution quality, and establishing automated procedures for latent model editing. Decision-makers should note that applying AttnLRP to Vision Transformers requires tuning a specific noise-dampening parameter to maintain high explanation quality, and attribution in extremely large models requires adequate hardware memory configurations.

No sufficiently relevant recommendations were found.

Cover for AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers

Abstract

Large Language Models are prone to biased predictions and hallucinations, underlining the paramount importance of understanding their model-internal reasoning process. However, achieving faithful attributions for the entirety of a black-box transformer model and maintaining computational efficiency is an unsolved challenge. By extending the Layer-wise Relevance Propagation attribution method to handle attention layers, we address these challenges effectively. While partial solutions exist, our method is the first to faithfully and holistically attribute not only input but also latent representations of transformer models with the computational efficiency similar to a single backward pass. Through extensive evaluations against existing methods on LLaMa 2, Mixtral 8x7b, Flan-T5 and vision transformer architectures, we demonstrate that our proposed approach surpasses alternative methods in terms of faithfulness and enables the understanding of latent representations, opening up the door for concept-based explanations. We provide an LRP library at https://github.com/rachtibat/LRP-eXplains-Transformers.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Perturbation & Local Surrogates
  • 2.2. Attention-based
  • 2.3. Backpropagation-based
  • 3. Attention-Aware LRP for Transformers
  • 3.1. Layer-wise Relevance Propagation
  • 3.1.1. DECOMPOSITION THROUGH LINEARIZATION
  • 3.2. Attributing the Multilayer Perceptron
  • 3.2.1. THE ε- AND γ-LRP RULE
  • 3.2.2. HANDLING ELEMENT-WISE NON-LINEARITIES
  • 3.3. Attributing Non-linear Attention
  • 3.3.1. HANDLING THE SOFTMAX NON-LINEARITY
  • 3.3.2. HANDLING MATRIX-MULTIPLICATION
  • 3.3.3. HANDLING NORMALIZATION LAYERS
  • 3.4. Understanding Latent Features
  • 4. Experiments
  • 4.1. Evaluating Explanations (Q1)
  • 4.1.1. BASELINES
  • 4.1.2. DISCUSSION
  • 4.2. Computational Complexity and Memory Consumption (Q2)
  • 4.3. Understanding & Manipulating Neurons (Q3)
  • 5. Conclusion
  • Limitations & Open Problems
  • Acknowledgements
  • Impact Statement
  • References
  • Appendix
  • A. Appendix I: Methodological Details
  • A.1. Details on Baseline Methods
  • A.1.1. INPUT × GRADIENT
  • A.1.2. INTEGRATED GRADIENTS
  • A.1.3. SMOOTHGRAD
  • A.1.4. ATTENTION ROLLOUT
  • A.1.6. KERNELSHAP
  • A.2. Details on AttnLRP
  • A.2.1. CONSERVATION & NUMERICAL STABILITY OF BIAS TERMS
  • A.2.2. HIGHLIGHTING THE DIFFERENCE BETWEEN VARIOUS LRP METHODS
  • A.2.3. TACKLING NOISE IN VISION TRANSFORMERS
  • A.2.4. IMPACT OF TEMPERATURE SCALING ON THE SOFTMAX RULE
  • A.3. Proofs
  • A.3.1. PROPOSITION 3.1: DECOMPOSING SOFTMAX
  • A.3.2. PROPOSITION 3.2: DECOMPOSING MULTIPLICATION
  • A.3.3. PROPOSITION 3.3: DECOMPOSING BI-LINEAR MATRIX MULTIPLICATION
  • A.3.4. PROPOSITION 3.4: LAYER NORMALIZATION
  • A.3.5. VIOLATION OF THE CONSERVATION PROPERTY IN BI-LINEAR MATRIX MULTIPLICATION
  • B. Appendix II: Experimental Details
  • B.1. Models and Datasets
  • B.2. Input Perturbation Metrics
  • B.3. Hyperparameter search for Baselines
  • B.4. Impact of Model Architectural Choices on AttnLRP Performance
  • B.5. LRP Composites for ViT
  • B.6. Additional Perturbation Evaluations on Vision Transformers
  • B.7. Attributions on SQuAD v2
  • B.8. Benchmarking Cost, Time and Memory Consumption
  • B.9. Attributions of Knowledge Neurons

Knowls

  1. Knowl 1 — AttnLRP propagates target-specific relevance through the full transformer

    model/method

    AttnLRP extends Layer-wise Relevance Propagation (LRP) to transformer operations so that a selected model output can be traced backward through the network, yielding relevance values for input features and latent neurons, including representations in the attention module. In LRP, the relevance RjR_j of an output is apportioned among its inputs as contributions Ri←jR_{i\leftarrow j}; the contributions to an input connected to multiple outputs are summed. For a function fjf_j of NN inputs xix_i, the additive explanation is expressed as

    fj(x)≈Rj=∑i=1NRi←j.f_j(x) \approx R_j = \sum_{i=1}^{N} R_{i\leftarrow j}.

    For adjacent layers l−1l-1 and ll, relevance is propagated in reverse order, with the layer total conserved when bias or other absorbing terms are accounted for:

    Rl−1=∑iRil−1=∑i,jRi←j(l−1,l)=∑jRjl=Rl.R^{l-1}=\sum_i R_i^{l-1}=\sum_{i,j}R_{i\leftarrow j}^{(l-1,l)}=\sum_jR_j^l=R^l.

    The method initializes relevance at the selected output and applies operation-specific rules during one backward traversal. Its Taylor-decomposition-based rules can let bias terms absorb some relevance; AttnLRP treats this as relevance assigned to a bias unit rather than as relevance propagated to the input. Unlike approaches that stop relevance at attention softmax, AttnLRP propagates through the attention computation, making query- and key-side latent attributions available as well as value-side attributions.

  2. Knowl 2 — Taylor decomposition gives a dedicated relevance rule for softmax

    equation

    For a softmax over nn logits xix_i, let si=exp⁡(xi)/∑k=1nexp⁡(xk)s_i=\exp(x_i)/\sum_{k=1}^{n}\exp(x_k) be its iith output. AttnLRP linearizes softmax at the current logits and uses the resulting relevance rule

    Riin=xi(Riout−si∑j=1nRjout),R_i^{\mathrm{in}}=x_i\left(R_i^{\mathrm{out}}-s_i\sum_{j=1}^{n}R_j^{\mathrm{out}}\right),

    where RioutR_i^{\mathrm{out}} is the relevance assigned to softmax output ii, and RiinR_i^{\mathrm{in}} is the relevance sent to input logit ii. The Taylor expansion includes a hidden bias term, which absorbs part of the relevance. This handles softmax without distributing that bias across inputs that may be zero, a practice the paper identifies as a source of numerical instability in successive propagation. The rule allows relevance to flow through the attention softmax rather than treating attention weights as fixed.

  3. Knowl 3 — Conservative rule for bilinear attention matrix multiplication

    equation

    For an attention-weight matrix AA and value matrix VV, let O=AVO=AV. For each query index jj, key index ii, and value-feature index pp, AjiA_{ji} is an attention weight, VipV_{ip} is a value, OjpO_{jp} is an output, and RjpOR^O_{jp} is the relevance of that output. AttnLRP splits the relevance of each product AjiVipA_{ji}V_{ip} equally between its two factors, then sums contributions over the other indices. With a small stabilizer ϵ\epsilon in the denominator, the rules are

    RjiA=∑pAjiVipRjpO2Ojp+ϵ,RipV=∑jAjiVipRjpO2Ojp+ϵ.R^A_{ji}=\sum_p\frac{A_{ji}V_{ip}R^O_{jp}}{2O_{jp}+\epsilon}, \qquad R^V_{ip}=\sum_j\frac{A_{ji}V_{ip}R^O_{jp}}{2O_{jp}+\epsilon}.

    The factor of two reflects the equal split between the two inputs to each product. Applying the ordinary linear relevance rule to both inputs of a bilinear product would count the output relevance twice and violate conservation. The uniform split avoids that overcounting; the stabilizer absorbs a negligible amount for nonzero outputs.

  4. Knowl 4 — AttnLRP uses identity propagation for the normalization core

    equation

    The nonlinear normalization core can be written as hj(x)=xj/g(x)h_j(x)=x_j/g(x). For LayerNorm, g(x)=Var⁡(x)+ϵng(x)=\sqrt{\operatorname{Var}(x)+\epsilon_n}; for RMSNorm, g(x)=1d∑k=1dxk2+ϵng(x)=\sqrt{\frac{1}{d}\sum_{k=1}^{d}x_k^2+\epsilon_n}, where dd is the feature dimension and ϵn\epsilon_n is the normalization stabilizer. AttnLRP applies an identity relevance rule to this core:

    Riin=Riout.R_i^{\mathrm{in}}=R_i^{\mathrm{out}}.

    The rule is obtained by Taylor decomposition at the zero reference point and avoids a linearization at the current input whose bias could absorb most of the relevance. Learnable affine components of normalization layers, such as scaling and shifting, are handled separately as linear operations. The identity rule also lets the normalization core be excluded from the relevance computation graph.

  5. Knowl 5 — Faithfulness and answer localization across vision and language models

    data/table

    The table reports the paper's faithfulness score, defined as the area between least-relevant-first and most-relevant-first perturbation curves (higher is better), for ViT-B-16 on ImageNet and LLaMa 2-7B on IMDB and Wikipedia next-token prediction. For SQuAD v2, the entries instead report top-1 accuracy of the most relevant token, followed in parentheses by IoU between positive relevance and the answer mask; only correctly answered questions were evaluated. The vision result uses 3,200 ImageNet samples, and the LLaMa results use 4,000 randomly selected samples per dataset. AttnLRP has the highest reported faithfulness score in each of the three perturbation evaluations and the highest top-1 accuracy on Mixtral; it ties Grad×AttnRoll on Flan-T5 top-1 accuracy while reporting higher IoU. ViT-L-16 and ViT-L-32 results further support the comparison: AttnLRP scores 7.17±0.047.17\pm0.04 and 6.06±0.046.06\pm0.04, respectively, versus 6.97±0.046.97\pm0.04 and 5.99±0.045.99\pm0.04 for the authors' γ\gamma-CP-LRP variant.

    Method ViT-B-16 ImageNet LLaMa 2-7B IMDB LLaMa 2-7B Wikipedia Mixtral 8x7B SQuAD v2 Flan-T5-XL SQuAD v2
    Random 0.01 ±\pm 0.01 -0.01 ±\pm 0.05 -0.07 ±\pm 0.13 0.03 (0.09) 0.03 (0.08)
    Input x Grad 0.80 ±\pm 0.03 0.12 ±\pm 0.05 0.18 ±\pm 0.13 0.56 (0.35) 0.60 (0.39)
    IG 1.54 ±\pm 0.03 1.23 ±\pm 0.05 4.05 ±\pm 0.13 0.68 (0.44) 0.10 (0.16)
    SmoothGrad -0.04 ±\pm 0.03 0.25 ±\pm 0.05 -2.22 ±\pm 0.14 0.47 (0.24) 0.05 (0.09)
    GradCAM 0.27 ±\pm 0.04 -0.82 ±\pm 0.05 2.01 ±\pm 0.15 0.82 (0.72) 0.81 (0.70)
    AttnRoll 1.31 ±\pm 0.03 -0.64 ±\pm 0.05 -3.49 ±\pm 0.15 0.05 (0.10) 0.02 (0.08)
    Grad x AttnRoll 2.60 ±\pm 0.03 1.61 ±\pm 0.05 9.79 ±\pm 0.14 0.91 (0.40) 0.94 (0.53)
    AtMan 0.70 ±\pm 0.02 -0.20 ±\pm 0.05 3.31 ±\pm 0.15 0.86 (0.83) 0.88 (0.80)
    KernelSHAP 4.71 ±\pm 0.03 – – – –
    CP-LRP, ϵ\epsilon-rule 2.53 ±\pm 0.02 1.72 ±\pm 0.04 7.85 ±\pm 0.12 0.50 (0.40) 0.91 (0.83)
    CP-LRP, γ\gamma-rule for ViT 6.06 ±\pm 0.02 – – – –
    AttnLRP 6.19 ±\pm 0.02 2.50 ±\pm 0.05 10.93 ±\pm 0.13 0.96 (0.72) 0.94 (0.84)

    For the SQuAD columns, each value is top-1 accuracy (IoU), not the perturbation faithfulness score. Mixtral 8x7B linear weights were quantized to 4-bit for the reported experiments, with computation in bfloat16; other computations also used bfloat16.

  6. Knowl 6 — Linear relevance rules and the Vision Transformer noise treatment

    model/method

    For a linear layer with inputs xix_i, weights WjiW_{ji}, bias bjb_j, preactivation zj=∑iWjixi+bjz_j=\sum_i W_{ji}x_i+b_j, and downstream relevance RjR_j, AttnLRP uses the ϵ\epsilon-rule unless a different rule is specified:

    Ri=∑jWjixiRjzj+ϵ sign⁡(zj).R_i=\sum_j\frac{W_{ji}x_iR_j}{z_j+\epsilon\,\operatorname{sign}(z_j)}.

    Here ϵ\epsilon is a small stabilizer; the bias and stabilizer can absorb relevance. The paper uses this rule on linear layers in language models, where it reports no visible attribution noise. In Vision Transformers, where it observes noisy attributions associated with gradient shattering, it instead uses the γ\gamma-rule on convolutional and ordinary linear layers outside the attention module. The rule strengthens contributions according to their sign: writing zij=Wjixiz_{ij}=W_{ji}x_i, zij+=max⁡(zij,0)z_{ij}^{+}=\max(z_{ij},0), zij−=min⁡(zij,0)z_{ij}^{-}=\min(z_{ij},0), and zjz_j for the layer output activation, it propagates

    Ri←j={zij+γzij+zj+γ∑kzkj+Rj,zj>0,zij+γzij−zj+γ∑kzkj−Rj,zj≤0,R_{i\leftarrow j}=\begin{cases} \dfrac{z_{ij}+\gamma z_{ij}^{+}}{z_j+\gamma\sum_k z_{kj}^{+}}R_j, & z_j>0,\\[6pt] \dfrac{z_{ij}+\gamma z_{ij}^{-}}{z_j+\gamma\sum_k z_{kj}^{-}}R_j, & z_j\leq 0, \end{cases}

    where γ>0\gamma>0 controls the strengthening. Their selected ViT composite used γ=0.25\gamma=0.25 for convolutional layers and γ=0.05\gamma=0.05 for ordinary linear layers, while retaining the ϵ\epsilon-rule for attention input projections (query, key, and value) and the attention output projection. These values were selected through a composite search on the evaluated Vision Transformer; the paper cautions that the same values are not guaranteed to suit other models.

  7. Knowl 7 — Layer-wise ablation shows gains beyond the attention mechanism

    data/table

    This cumulative ablation begins with CP-LRP rules on every layer, then replaces rules for attention, FFN nonlinear weighting, and, where present, expert routing with AttnLRP rules. Faithfulness is measured by the area between least- and most-relevant perturbation curves on LLaMa 2-7B (IMDB and Wikipedia); SQuAD v2 entries give top-1 accuracy with IoU in parentheses for Mixtral 8x7B and Flan-T5-XL. The gains on LLaMa increase after adding AttnLRP's handling of FFN nonlinear weighting, while adding the routing rule in Mixtral raises accuracy from 0.78 to 0.96. The results show that the improvement is not confined to attention, particularly in architectures with more nonlinear operations.

    Cumulative rule replacement LLaMa 2-7B IMDB LLaMa 2-7B Wikipedia Mixtral 8x7B SQuAD v2 Flan-T5-XL SQuAD v2
    CP-LRP on all layers 1.72 7.85 0.50 (0.40) 0.90 (0.83)
    Add AttnLRP attention rules 2.09 9.49 0.70 (0.53) 0.94 (0.84)
    Also add FFN nonlinear-weighting rules 2.50 10.93 0.78 (0.57) –
    Also add routing rule – – 0.96 (0.72) –

    The final routing step applies to Mixtral's expert-routing layer; the dashes indicate that the corresponding model or ablation result is not reported.

  8. Knowl 8 — ActMax-guided relevance identifies and enables intervention on knowledge neurons

    model/method

    AttnLRP supports a two-part analysis of latent knowledge neurons in FFN blocks. First, collect reference prompts that maximize a neuron's activation (Activation Maximization, or ActMax), then propagate relevance from that neuron to identify the input tokens most responsible for its activation. For knowledge neurons at the last nonlinearity of an FFN, the corresponding row of the second FFN weight matrix can also be projected onto the vocabulary to inspect which tokens the neuron promotes. The paper demonstrates this analysis on Phi-1.5 using highly activating sentences from Wikipedia summaries: neuron 3948 in layer 17 is associated with cold temperatures and cold-region concepts; neuron 5687 in layer 18 with candy and sweetness; and neuron 4104 in layer 17 with dryness and desert-related concepts. In the prompt “The ice bear lives in the,” the predicted continuation is “Arctic.” Disabling neuron 3948 while strongly amplifying neuron 5687 changed the continuation toward “sweet, sugary treats of the candy store”; increasing neuron 4104 shifted it toward “desert.” This is a case study showing that relevance-guided neuron identification can support targeted changes to generation, not evidence that the same interventions will generalize to arbitrary prompts or models.

  9. Knowl 9 — Reported cost scales better than token-by-token perturbation

    data/table

    The paper compares one LRP-based attribution using checkpointing with linear-time perturbation, such as AtMan or a linear-time Shapley-based method. Relative to a single forward pass, the reported complexity and memory costs are:

    Method Computational complexity Memory consumption
    LRP with checkpointing O(1)O(1) O(NL)O(\sqrt{N_L})
    Linear-time perturbation O(NT)O(N_T) O(1)O(1)

    Here NLN_L is the number of model layers and NTN_T is the number of input tokens. Checkpointed LRP uses two forward passes and one backward pass; perturbation requires NTN_T forward passes but has constant memory use. The paper's benchmark also reports that checkpointed attribution of LLaMa 2-70B at a context length of 4096 exceeded 160 GB of memory, so it did not fit on the tested four-A100 node, whereas perturbation-based methods were sufficiently memory-efficient for that setting.

  10. Knowl 10 — AttnLRP limitations and open questions

    limitation

    The paper identifies tuning the γ\gamma parameter as important for accurate Vision Transformer attributions; its selected values are based on evaluated models and are not established as universal. It leaves the effect of quantized number formats on attributions and the design of custom GPU kernels for LRP rules for future investigation. It also does not systematically study how disentangled knowledge-neuron representations are, or the effects of temperature scaling when attributing a classification softmax output. For a saturated softmax, the paper notes that vanishing derivatives can distort relevance flow and suggests increasing temperature when explaining that output, while explicitly leaving this setting uninvestigated.

Coverage note — Detailed cost, runtime, and energy curves; additional ViT-L and qualitative attribution examples; and supplementary ActMax neuron examples are omitted because their main findings are represented by the reported benchmark, complexity comparison, and representative neuron case.

References

  1. 1.Abnar, S. and Zuidema, W. H. (2020). Quantifying attention flow in transformers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4190–4197.
  2. 2.Achanta, R., Shaji, A., Smith, K., Lucchi, A., Fua, P., and Süsstrunk, S. (2012). Slic superpixels compared to state-of-the-art superpixel methods. IEEE transactions on pattern analysis and machine intelligence, 34(11):2274–2282.
  3. 3.Achtibat, R., Dreyer, M., Eisenbraun, I., Bosse, S., Wiegand, T., Samek, W., and Lapuschkin, S. (2023). From attribution maps to human-understandable explanations through concept relevance propagation. Nature Machine Intelligence, 5(9):1006–1019.
  4. 4.Ali, A., Schnake, T., Eberle, O., Montavon, G., Müller, K.-R., and Wolf, L. (2022). Xai for transformers: Better explanations through conservative propagation. In International Conference on Machine Learning, pages 435–451. PMLR.
  5. 5.Anders, C. J., Neumann, D., Samek, W., Müller, K.-R., and Lapuschkin, S. (2021). Software for dataset-wide xai: from local explanations to global insights with zennit, corelay, and virelay. arXiv preprint arXiv:2106.13200.
  6. 6.Arras, L., Osman, A., and Samek, W. (2022). Clevr-xai: A benchmark dataset for the ground truth evaluation of neural network explanations. Information Fusion, 81:14–40.
  7. 7.Ba, J. L., Kiros, J. R., and Hinton, G. E. (2016). Layer normalization. arXiv preprint arXiv:1607.06450.
  8. 8.Bach, S., Binder, A., Montavon, G., Klauschen, F., Müller, K.-R., and Samek, W. (2015). On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. PLoS ONE, 10(7):e0130140.
  9. 9.Balduzzi, D., Frean, M., Leary, L., Lewis, J., Ma, K. W.-D., and McWilliams, B. (2017). The shattered gradients problem: If resnets are the answer, then what is the question? In International Conference on Machine Learning, pages 342–350. PMLR.
  10. 10.Binder, A., Montavon, G., Lapuschkin, S., Müller, K.-R., and Samek, W. (2016). Layer-wise relevance propagation for neural networks with local renormalization layers. In Artificial Neural Networks and Machine Learning–ICANN 2016: 25th International Conference on Artificial Neural Networks, Barcelona, Spain, September 6-9, 2016, Proceedings, Part II 25, pages 63–71. Springer.
  11. 11.Blücher, S., Vielhaben, J., and Strodthoff, N. (2024). Decoupling pixel flipping and occlusion strategy for consistent xai benchmarks. arXiv preprint arXiv:2401.06654.
  12. 12.Brocki, L. and Chung, N. C. (2023). Feature perturbation augmentation for reliable evaluation of importance estimators in neural networks. Pattern Recognition Letters, 176:131–139.
  13. 13.Caron, M., Touvron, H., Misra, I., Jegou, H., Mairal, J., Bojanowski, P., and Joulin, A. (2021). Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9650–9660.
  14. 14.Chang, C.-H., Creager, E., Goldenberg, A., and Duvenaud, D. (2018). Explaining image classifiers by counterfactual generation. arXiv preprint arXiv:1807.08024.
  15. 15.Chefer, H., Gur, S., and Wolf, L. (2021a). Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 397–406.
  16. 16.Chefer, H., Gur, S., and Wolf, L. (2021b). Transformer interpretability beyond attention visualization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 782–791.
  17. 17.Chen, T., Xu, B., Zhang, C., and Guestrin, C. (2016). Training deep nets with sublinear memory cost. arXiv preprint arXiv:1604.06174.
  18. 18.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al. (2022). Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  19. 19.Clark, K., Khandelwal, U., Levy, O., and Manning, C. D. (2019). What does bert look at? an analysis of bert’s attention. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 276–286.
  20. 20.Dai, D., Dong, L., Hao, Y., Sui, Z., Chang, B., and Wei, F. (2022). Knowledge neurons in pretrained transformers. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8493–8502.
  21. 21.Dao, T., Fu, D., Ermon, S., Rudra, A., and Re, C. (2022). Flashattention: Fast and memory-efficient exact attention with io-awareness. Advances in Neural Information Processing Systems, 35:16344–16359.
  22. 22.Deb, M., Deiseroth, B., Weinbach, S., Schramowski, P., and Kersting, K. (2023). Atman: Understanding transformer predictions through memory efficient attention manipulation. arXiv preprint arXiv:2301.08110.
  23. 23.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009). Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee.
  24. 24.Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. (2024). Qlora: Efficient finetuning of quantized llms. Advances in Neural Information Processing Systems, 36.
  25. 25.Ding, Y., Liu, Y., Luan, H., and Sun, M. (2017). Visualizing and understanding neural machine translation. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1150–1159.
  26. 26.Dombrowski, A.-K., Anders, C. J., Müller, K.-R., and Kessel, P. (2022). Towards robust explanations for deep neural networks. Pattern Recognition, 121:108194.
  27. 27.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., et al. (2021). An image is worth 16x16 words: Transformers for image recognition at scale. In 9th International Conference on Learning Representations, ICLR.
  28. 28.Fatima, S. S., Wooldridge, M., and Jennings, N. R. (2008). A linear approximation method for the shapley value. Artificial Intelligence, 172(14):1673–1699.
  29. 29.Fedus, W., Dean, J., and Zoph, B. (2022). A review of sparse expert models in deep learning. arXiv preprint arXiv:2209.01667.
  30. 30.Fong, R. C. and Vedaldi, A. (2017). Interpretable explanations of black boxes by meaningful perturbation. In IEEE International Conference on Computer Vision (ICCV), pages 3449–3457.
  31. 31.Fryer, D., Strumke, I., and Nguyen, H. (2021). Shapley values for feature selection: The good, the bad, and the axioms. IEEE Access, 9:144352–144360.
  32. 32.Geva, M., Caciularu, A., Wang, K., and Goldberg, Y. (2022). Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 30–45.
  33. 33.Geva, M., Schuster, R., Berant, J., and Levy, O. (2021). Transformer feed-forward layers are key-value memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5484–5495.
  34. 34.Gildenblat, J. (2020. Accessed on Dec 01, 2023). Exploring explainability for vision transformers. https://jacobgil.github.io/deeplearning/vision-transformer-explainability.
  35. 35.Guidotti, R., Monreale, A., Ruggieri, S., Pedreschi, D., Turini, F., and Giannotti, F. (2018). Local rule-based explanations of black box decision systems. arXiv preprint arXiv:1805.10820.
  36. 36.Hedström, A., Weber, L., Krakowczyk, D., Bareeva, D., Motzkus, F., Samek, W., Lapuschkin, S., and Höhne, M. M. M. (2023). Quantus: An explainable ai toolkit for responsible evaluation of neural network explanations and beyond. Journal of Machine Learning Research, 24(34):1–11.
  37. 37.Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., et al. (2023). A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. arXiv preprint arXiv:2311.05232.
  38. 38.Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al. (2024). Mixtral of experts. arXiv preprint arXiv:2401.04088.
  39. 39.Kokhlikyan, N., Miglani, V., Martin, M., Wang, E., Alsallakh, B., Reynolds, J., Melnikov, A., Kliushkina, N., Araya, C., Yan, S., et al. (2020). Captum: A unified and generic model interpretability library for pytorch. arXiv preprint arXiv:2009.07896.
  40. 40.Li, Y., Bubeck, S., Eldan, R., Del Giorno, A., Gunasekar, S., and Lee, Y. T. (2023). Textbooks are all you need ii: phi-1.5 technical report. arXiv preprint arXiv:2309.05463.
  41. 41.Lundberg, S. M. and Lee, S. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems 30, pages 4765–4774.
  42. 42.Maas, A., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C. (2011). Learning word vectors for sentiment analysis. In Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies, pages 142–150.
  43. 43.Mao, C., Jiang, L., Dehghani, M., Vondrick, C., Sukthankar, R., and Essa, I. (2021). Discrete representations strengthen vision transformer robustness. In International Conference on Learning Representations.
  44. 44.Miglani, V., Yang, A., Markosyan, A., Garcia-Olano, D., and Kokhlikyan, N. (2023). Using captum to explain generative language models. In Proceedings of the 3rd Workshop for Natural Language Processing Open Source Software (NLP-OSS 2023), pages 165–173.
  45. 45.Montavon, G., Binder, A., Lapuschkin, S., Samek, W., and Müller, K.-R. (2019). Layer-wise relevance propagation: an overview. Explainable AI: interpreting, explaining and visualizing deep learning, pages 193–209.
  46. 46.Montavon, G., Lapuschkin, S., Binder, A., Samek, W., and Müller, K.-R. (2017). Explaining nonlinear classification decisions with deep taylor decomposition. Pattern recognition, 65:211–222.
  47. 47.Nguyen, A., Dosovitskiy, A., Yosinski, J., Brox, T., and Clune, J. (2016). Synthesizing the preferred inputs for neurons in neural networks via deep generator networks. Advances in neural information processing systems, 29.
  48. 48.Pahde, F., Yolcu, G. U., Binder, A., Samek, W., and Lapuschkin, S. (2023). Optimizing explanations by network canonization and hyperparameter search. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3818–3827.
  49. 49.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. (2019). Pytorch: An imperative style, high-performance deep learning library. Advances in Neural Information Processing Systems, 32.
  50. 50.Rajpurkar, P., Jia, R., and Liang, P. (2018). Know what you don’t know: Unanswerable questions for squad. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 784–789.
  51. 51.Ribeiro, M. T., Singh, S., and Guestrin, C. (2016). ”why should I trust you?”: Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 1135–1144. ACM.
  52. 52.Samek, W., Binder, A., Montavon, G., Lapuschkin, S., and Müller, K.-R. (2017). Evaluating the visualization of what a deep neural network has learned. IEEE Transactions on Neural Networks and Learning Systems, 28(11):2660–2673.
  53. 53.Scheepers, T. (2017). Improving the compositionality of word embeddings. Master’s thesis, Universiteit van Amsterdam, Science Park 904, Amsterdam, Netherlands.
  54. 54.Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D. (2017). Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), pages 618–626.
  55. 55.Shaham, U., Ivgi, M., Efrat, A., Berant, J., and Levy, O. (2023). Zeroscrolls: A zero-shot benchmark for long text understanding. arXiv preprint arXiv:2305.14196.
  56. 56.Shrikumar, A., Greenside, P., and Kundaje, A. (2017). Learning important features through propagating activation differences. In International Conference on Machine Learning, pages 3145–3153. PMLR.
  57. 57.Simonyan, K., Vedaldi, A., and Zisserman, A. (2014). Deep inside convolutional networks: visualising image classification models and saliency maps. In Proceedings of the International Conference on Learning Representations (ICLR). ICLR.
  58. 58.Smilkov, D., Thorat, N., Kim, B., Viegas, F., and Wattenberg, M. (2017). Smoothgrad: removing noise by adding noise. arXiv preprint arXiv:1706.03825.
  59. 59.Sundararajan, M., Taly, A., and Yan, Q. (2017). Axiomatic attribution for deep networks. In International Conference on Machine Learning, pages 3319–3328. PMLR.
  60. 60.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  61. 61.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.
  62. 62.Voita, E., Ferrando, J., and Nalmpantis, C. (2023). Neurons in large language models: Dead, n-gram, positional. arXiv preprint arXiv:2309.04827.
  63. 63.Voita, E., Sennrich, R., and Titov, I. (2021). Analyzing the source and target contributions to predictions in neural machine translation. In 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP, pages 1126–1140.
  64. 64.Wiegreffe, S. and Pinter, Y. (2019). Attention is not not explanation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 11–20.
  65. 65.Wikimedia Foundation (2023. Accessed on Dec 01, 2023). Wikimedia downloads. https://dumps.wikimedia.org.
  66. 66.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al. (2019). Huggingface’s transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771.
  67. 67.Zeiler, M. D. and Fergus, R. (2014). Visualizing and understanding convolutional networks. In European Conference Computer Vision - ECCV 2014, pages 818–833.
  68. 68.Zhang, B. and Sennrich, R. (2019). Root mean square layer normalization. Advances in Neural Information Processing Systems, 32.

Citation

MLA
Achtibat, R., et al. “AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers”. arXiv, 2024, http://arxiv.org/abs/2402.05602v2.
APA
Achtibat, R., Hatefi, S. M. V., Dreyer, M., Jain, A., Wiegand, T., Lapuschkin, S., & Samek, W. (2024). AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers. arXiv. http://arxiv.org/abs/2402.05602v2
Chicago
Achtibat, R., S. M. V. Hatefi, M. Dreyer, et al. 2024. “AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers”. arXiv. http://arxiv.org/abs/2402.05602v2.
Harvard
Achtibat, R. et al. (2024) “AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.05602v2.
Vancouver
1. Achtibat R, Hatefi SMV, Dreyer M, Jain A, Wiegand T, Lapuschkin S, Samek W (2024) AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers. arXiv

BibTeX

@article{achtibat2024attnlrp,
  title = {AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers},
  author = {Achtibat, Reduan and Hatefi, Sayed Mohammad Vakilzadeh and Dreyer, Maximilian and Jain, Aakriti and Wiegand, Thomas and Lapuschkin, Sebastian and Samek, Wojciech},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.05602v2},
  eprint = {2402.05602}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/