AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
Reduan AchtibatSayed Mohammad Vakilzadeh HatefiMaximilian DreyerAakriti JainThomas WiegandSebastian LapuschkinWojciech Samek
Develops an efficient Layer-wise Relevance Propagation framework tailored for attention mechanisms and non-linear transformer components, enabling faithful attribution of both input features and latent representations in a single backward pass.
Large transformer architectures, such as Large Language Models and Vision Transformers, are widely used across industries but remain vulnerable to bias and hallucinations. Because these models operate as opaque "black boxes," organizations face compliance, safety, and operational risks when deploying them in high-stakes environments. Existing interpretability methods are either too computationally expensive—requiring massive compute resources for iterative perturbations—or produce unfaithful, noisy explanations that fail to handle the complex non-linear components inside transformers.
The article introduces and evaluates AttnLRP, an attention-aware extension of Layer-wise Relevance Propagation designed to provide highly faithful, holistic explanations for transformer models. Its primary objective is to accurately attribute model predictions back to both input tokens and internal latent representations within the computational efficiency of a single backward pass.
To achieve this, the authors mathematically formulated relevance propagation rules within the Deep Taylor Decomposition framework specifically tailored to non-linear operations, such as softmax, matrix multiplication, and normalization layers. They evaluated AttnLRP across multiple architectures, including LLaMa 2, Mixtral 8x7b, Flan-T5-XL, and Vision Transformers, utilizing standard datasets such as ImageNet, IMDB reviews, Wikipedia, and SQuAD v2. Faithfulness was rigorously assessed by measuring performance drops when perturbing model inputs based on assigned importance scores, and computational efficiency was benchmarked against existing perturbation and gradient-based approaches.
The key findings demonstrate that AttnLRP consistently outperforms competing explainability techniques across text and vision domains. In language models, AttnLRP achieved superior faithfulness scores, such as an area metric of 10.93 on Wikipedia next-word prediction compared to 7.85 for conservative propagation baselines and negative or negligible scores for standard gradient methods. In complex architectures with routing and non-linear feed-forward layers, like Mixtral 8x7b, AttnLRP improved top-1 attribution accuracy to 0.96, representing a 46% gain over prior conservative propagation methods. Furthermore, the approach requires only constant computational complexity relative to token length, avoiding the exponential time and energy costs associated with linear perturbation methods. Finally, the authors demonstrated that AttnLRP can pinpoint specific internal "knowledge neurons," enabling targeted model editing to manipulate outputs systematically without full retraining.
These findings indicate that organizations can explain and audit large-scale transformer models faithfully with minimal computational and financial overhead. High-efficiency attribution enables real-time interpretability, systematic debugging of hallucinations, and safer model deployment in regulated sectors such as healthcare and finance. The ability to identify and edit latent concepts also provides a viable path for steering model behavior and mitigating biases directly within internal layers.
Moving forward, practitioners and technical leaders seeking deeper transparency should adopt backpropagation-based frameworks like AttnLRP over noisy gradient or costly perturbation methods. Future implementation work should focus on optimizing custom graphics processing unit kernels, analyzing the impact of low-bit quantization on attribution quality, and establishing automated procedures for latent model editing. Decision-makers should note that applying AttnLRP to Vision Transformers requires tuning a specific noise-dampening parameter to maintain high explanation quality, and attribution in extremely large models requires adequate hardware memory configurations.
- Paper: XAI for Transformers: Better Explanations through Conservative Propagation, Ameen Ali et al. (2022). Its conservative propagation rules for attention and normalization provide the direct methodological groundwork that AttnLRP extends to attribute relevance throughout transformers.
- Paper: Methods for interpreting and understanding deep neural networks, Grégoire Montavon et al. (2018). This overview introduces LRP’s relevance-decomposition principles and propagation rules, which underpin AttnLRP’s transformer-specific attribution method.
- Paper: Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models, Wojciech Samek et al. (2017). Its account of LRP’s conservation-based redistribution gives the core attribution framework that AttnLRP adapts to attention layers.
- Paper: Evaluating the Visualization of What a Deep Neural Network Has Learned, Wojciech Samek et al. (2015). Its evaluation of LRP establishes how relevance maps are tested for faithfulness, helping contextualize AttnLRP’s attribution benchmarks.
- Paper: Quantifying Attention Flow in Transformers, Samira Abnar et al. (2020). Its analysis shows why raw attention weights fail to track token importance, motivating methods such as AttnLRP that propagate relevance through transformer internals.
No sufficiently relevant recommendations were found.
